bpf: Allow selecting numa node during map creation
The current map creation API does not allow to provide the numa-node
preference. The memory usually comes from where the map-creation-process
is running. The performance is not ideal if the bpf_prog is known to
always run in a numa node different from the map-creation-process.
One of the use case is sharding on CPU to different LRU maps (i.e.
an array of LRU maps). Here is the test result of map_perf_test on
the INNER_LRU_HASH_PREALLOC test if we force the lru map used by
CPU0 to be allocated from a remote numa node:
[ The machine has 20 cores. CPU0-9 at node 0. CPU10-19 at node 1 ]
># taskset -c 10 ./map_perf_test 512 8
1260000 8000000
5:inner_lru_hash_map_perf pre-alloc
1628380 events per sec
4:inner_lru_hash_map_perf pre-alloc
1626396 events per sec
3:inner_lru_hash_map_perf pre-alloc
1626144 events per sec
6:inner_lru_hash_map_perf pre-alloc
1621657 events per sec
2:inner_lru_hash_map_perf pre-alloc
1621534 events per sec
1:inner_lru_hash_map_perf pre-alloc
1620292 events per sec
7:inner_lru_hash_map_perf pre-alloc
1613305 events per sec
0:inner_lru_hash_map_perf pre-alloc
1239150 events per sec #<<<
After specifying numa node:
># taskset -c 10 ./map_perf_test 512 8
1260000 8000000
5:inner_lru_hash_map_perf pre-alloc
1629627 events per sec
3:inner_lru_hash_map_perf pre-alloc
1628057 events per sec
1:inner_lru_hash_map_perf pre-alloc
1623054 events per sec
6:inner_lru_hash_map_perf pre-alloc
1616033 events per sec
2:inner_lru_hash_map_perf pre-alloc
1614630 events per sec
4:inner_lru_hash_map_perf pre-alloc
1612651 events per sec
7:inner_lru_hash_map_perf pre-alloc
1609337 events per sec
0:inner_lru_hash_map_perf pre-alloc
1619340 events per sec #<<<
This patch adds one field, numa_node, to the bpf_attr. Since numa node 0
is a valid node, a new flag BPF_F_NUMA_NODE is also added. The numa_node
field is honored if and only if the BPF_F_NUMA_NODE flag is set.
Numa node selection is not supported for percpu map.
This patch does not change all the kmalloc. F.e.
'htab = kzalloc()' is not changed since the object
is small enough to stay in the cache.
Signed-off-by: Martin KaFai Lau <kafai@fb.com>
Acked-by: Daniel Borkmann <daniel@iogearbox.net>
Acked-by: Alexei Starovoitov <ast@fb.com>
Signed-off-by: David S. Miller <davem@davemloft.net>