Linux Kernel NUMA Balancing & Memory Tiering: Maximizing Dual-Socket AMD EPYC Dedicated Servers in Pakistan

Eliminate cross-socket memory latency and avoid CPU soft lockups on dual-socket AMD EPYC and Intel Xeon platforms. Deep-dive into automatic NUMA balancing, memory pinning, and sysctl tuning for Pakistani enterprise workloads.

Linux Kernel NUMA Balancing & Memory Tiering: Maximizing Dual-Socket AMD EPYC Dedicated Servers in Pakistan

Modern bare-metal enterprise servers routinely sport 64, 128, or 256 physical CPU cores paired with 512GB to 2TB of high-speed DDR5 memory. On multi-socket and multi-chiplet server platforms—such as dual-socket AMD EPYC (Zen 4 Bergamo/Genoa) and Intel Xeon Scalable processors—this hardware is fundamentally structured as a Non-Uniform Memory Access (NUMA) architecture.

In a NUMA system, physical memory is partitioned across multiple distinct nodes, with each CPU socket or Core Complex Die (CCD) directly wired to its own local memory channels. While a processor can access remote memory attached to neighboring sockets across the AMD Infinity Fabric or Intel UPI (Ultra Path Interconnect), doing so incurs a significant latency penalty—typically 80ns to 100ns for local DRAM versus 160ns to 220ns for cross-socket remote DRAM.

For enterprise workloads deployed on Dedicated Servers, including massive MariaDB/PostgreSQL databases, Redis caching clusters, and KVM virtualization hypervisors, unoptimized NUMA memory placement leads to mysterious CPU soft lockups, erratic p99 tail latency, and severe memory thrashing.


1. Deconstructing the NUMA Topology

Before tuning the Linux kernel, you must map the physical NUMA nodes, socket interconnect distances, and core assignments:

# Inspect hardware NUMA topology via numactl
numactl --hardware

On a dual-socket AMD EPYC 9654 server (192 physical cores, 384 threads, 1TB RAM), output typically reveals:

available: 2 nodes (0-1)
node 0 cpus: 0-95 192-287
node 0 size: 515840 MB
node 0 free: 489210 MB
node 1 cpus: 96-191 288-383
node 1 size: 516096 MB
node 1 free: 494110 MB
node distances:
node   0   1 
  0:  10  32 
  1:  32  10 

The node distances matrix shows the relative latency factor. A distance of 10 represents local memory access, while 32 represents a 3.2x latency multiple when Node 0 fetches cachelines from Node 1’s physical memory bus across the interconnect.

+------------------------------------+        +------------------------------------+
|            NUMA NODE 0             |        |            NUMA NODE 1             |
|  +------------------------------+  |        |  +------------------------------+  |
|  |     Socket 0 (Cores 0-95)    |  |        |  |    Socket 1 (Cores 96-191)   |  |
|  +--------------+---------------+  |        |  +--------------+---------------+  |
|                 |                  |        |                 |                  |
|          [Local Bus: ~85ns]        |        |          [Local Bus: ~85ns]        |
|                 |                  |        |                 |                  |
|  +--------------v---------------+  |        |  +--------------v---------------+  |
|  |    Local DDR5 Memory Pool    |  |        |  |    Local DDR5 Memory Pool    |  |
|  |           (512 GB)           |  |        |  |           (512 GB)           |  |
|  +------------------------------+  |        |  +------------------------------+  |
+-----------------|------------------+        +------------------|-----------------+
                  |                                              |
                  +=====[ AMD Infinity Fabric / Intel UPI ]======+
                            (Cross-Socket: ~200ns)

2. The Danger of Automatic Kernel NUMA Balancing

By default, modern enterprise Linux distributions (CentOS/AlmaLinux 9, Ubuntu 24.04 LTS, Debian 12) enable Automatic NUMA Balancing (kernel.numa_balancing = 1).

The kernel periodically unmaps memory pages to provoke minor page faults (NUMA hinting faults). When a thread running on Node 1 faults on a page located on Node 0, the kernel tracks access frequency and attempts to migrate the physical memory page across the interconnect to Node 1.

Why This Destroys Database & In-Memory Performance

For large-memory monolithic applications:

  1. Massive Page Migration Overhead: If MariaDB’s InnoDB buffer pool is 384GB, and query worker threads execute across both sockets, the kernel spends massive CPU cycles scanning pages and copying memory across the interconnect, consuming memory bandwidth and triggering TLB shootdowns.
  2. CPU Lockups and Jitter: The kcompactd and numa kernel worker threads lock memory zones during migration, causing execution stalls that manifest as 500ms+ database latency spikes.

3. Disabling or Restricting Automatic NUMA Balancing

For dedicated database servers and latency-critical backends, automatic NUMA balancing should be explicitly disabled:

# Disable automatic NUMA balancing immediately in runtime
sysctl -w kernel.numa_balancing=0

# Persist in sysctl.conf
cat << 'EOF' >> /etc/sysctl.d/99-numa-tuning.conf
# Disable automatic NUMA page migration overhead
kernel.numa_balancing = 0

# Prevent aggressive local zone reclamation thrashing
vm.zone_reclaim_mode = 0

# Enforce strict contiguous memory compaction thresholds
vm.extfrag_threshold = 500
EOF

sysctl -p /etc/sysctl.d/99-numa-tuning.conf

Understanding vm.zone_reclaim_mode = 0

When vm.zone_reclaim_mode is set to 1 or 2, the Linux kernel aggressively attempts to reclaim local page caches or swap clean pages when a local NUMA node runs out of memory, even if neighboring NUMA nodes have hundreds of gigabytes of free RAM. Setting vm.zone_reclaim_mode = 0 instructs the kernel to allocate from the remote node instead of pausing the application to reclaim local pages.


4. Workload Isolation: Binding Workloads with numactl and cgroups

When architecting production services on Dedicated Servers in Pakistan, the most performant strategy is NUMA-Aware Workload Partitioning: dedicating specific NUMA nodes to specific workloads.

Strategy A: Memory Interleaving for Large Relational Databases

For monolithic databases that must span multiple sockets and access a unified 1TB buffer pool, memory should be interleaved round-robin across all nodes to distribute memory bandwidth equally:

# Launch MariaDB with interleaved memory allocation across all NUMA nodes
numactl --interleave=all /usr/sbin/mariadbd --defaults-file=/etc/my.cnf.d/server.cnf

In systemd, configure this directly in the service unit override:

# /etc/systemd/system/mariadb.service.d/override.conf
[Service]
ExecStart=
ExecStart=/usr/bin/numactl --interleave=all /usr/sbin/mariadbd $MYSQLD_OPTS $_OPTS

Strategy B: Strict Node Pinning for In-Memory Redis Caching

For Redis instances, which are single-threaded event loops, allocating cross-node memory causes catastrophic cacheline misses. Bind each Redis instance strictly to a single NUMA node and its local CPU cores:

# /etc/systemd/system/redis-instance-01.service.d/override.conf
[Service]
ExecStart=
ExecStart=/usr/bin/numactl --cpunodebind=0 --membind=0 /usr/bin/redis-server /etc/redis/instance-01.conf

# /etc/systemd/system/redis-instance-02.service.d/override.conf
[Service]
ExecStart=
ExecStart=/usr/bin/numactl --cpunodebind=1 --membind=1 /usr/bin/redis-server /etc/redis/instance-02.conf

This guarantees 100% local DRAM access (~85ns) with zero interconnect traffic across the Infinity Fabric.


5. NUMA Performance Monitoring & Diagnostics

To verify whether your system is experiencing remote memory stalls, monitor NUMA hit/miss metrics using numastat:

# Monitor NUMA allocation statistics per node
numastat -c

Output inspection:

                           Node 0          Node 1           Total
Numa_Hit             125849201948    124982140281    250831342229
Numa_Miss                 4820192         5120491         9940683
Numa_Foreign              5120491         4820192         9940683
Interleave_Hit       894819283018    894819102491   1789638385509
Local_Node           125841002914    124978921820    250819924734
Other_Node                8199034         3218461        11417495
  • Numa_Hit: Memory intended for this node that was successfully allocated locally. High is ideal.
  • Numa_Miss: Memory intended for this node that had to be allocated on another node due to local exhaustion.
  • Numa_Foreign: Memory intended for another node that was forced into this node.
  • Goal: Maintain Numa_Miss / Numa_Hit below 0.5% for all latency-sensitive services.

6. Summary Comparison: NUMA Tuning Modes

Parameter / Configuration Default Out-of-the-Box Enterprise Optimized Target Benefit
kernel.numa_balancing 1 (Enabled) 0 (Disabled) Eliminates unmapped page fault spikes & CPU migration thrashing
vm.zone_reclaim_mode 0 or 1 (Distro dependent) 0 Eliminates synchronous I/O pauses when local zone is full
Relational DB Allocation OS Default (Node 0 bias) numactl --interleave=all Evenly distributes DRAM bandwidth across multi-socket buses
Microservices / Redis Unconstrained scheduler numactl --cpunodebind=X --membind=X Guarantees 100% sub-100ns local DRAM cache hits
Virtualization (KVM/QEMU) Floating vCPUs vCPU pinning + NUMA guest mapping Prevents cross-socket L3 cache pollution between noisy VMs

Carefully structuring NUMA boundaries unlocks the raw hardware potential of multi-socket bare metal, ensuring uninterrupted enterprise throughput across high-concurrency Dedicated Servers in Pakistan.

High-Density AMD EPYC & Intel Xeon Bare-Metal Infrastructure

Need dedicated computational muscle without virtualization overhead? NextGen offers high-core-count, multi-channel DDR5 dedicated bare-metal servers engineered for low-latency databases and mission-critical enterprise workloads.

Explore Enterprise Dedicated Servers