Linux Kernel THP: Transparent Huge Pages Tuning for Database Nodes in Pakistan

A production guide to resolving memory compaction latency spikes and khugepaged lock contention by tuning Transparent Huge Pages (THP) for MariaDB, MySQL, and Redis servers in Pakistan.

Linux Kernel THP: Transparent Huge Pages Tuning for Database Nodes in Pakistan

High-memory database servers running MariaDB, PostgreSQL, MongoDB, or Redis in Pakistani enterprises frequently experience unexplained periodic query stalls: transactional latencies that typically register at 2 to 5 milliseconds suddenly skyrocket to 8,000 milliseconds, CPU utilization spikes into kernel space (sys CPU), and system load averages surge.

When database administrators inspect slow query logs, no particular SQL query stands out. The underlying operating system metric reveals the culprit: khugepaged memory compaction latency induced by Transparent Huge Pages (THP).

While Transparent Huge Pages were introduced in the Linux kernel to improve translation lookaside buffer (TLB) hit rates by aggregating 4KB memory pages into 2MB blocks automatically, their aggressive background defragmentation behavior wreaks havoc on transactional database workloads.

In this deep architectural guide, we dissect the TLB cache mechanics of 4KB vs. 2MB pages, evaluate the devastating impact of [always] THP mode on sparse database buffer pools, configure safe [madvise] and [never] policies, and ensure consistent low-latency performance on Dedicated Servers.


The TLB Bottleneck: 4KB Pages vs. 2MB Huge Pages

Modern x86-64 processors convert virtual memory addresses to physical RAM addresses using page tables. To accelerate this translation, CPU hardware caches recent translations inside the Translation Lookaside Buffer (TLB):

  • Standard Linux Page Size: $4\text{ KB}$ ($4,096\text{ bytes}$)
  • Standard 2MB Huge Page Size: $2\text{ MB}$ ($2,097,152\text{ bytes} = 512 \times 4\text{ KB}$)

Consider a database server with 128 GB of RAM allocated to the MariaDB Buffer Pool:

  • Using 4KB pages requires $\frac{128\text{ GB}}{4\text{ KB}} = 33,554,432$ page table entries.
  • Using 2MB huge pages requires only $\frac{128\text{ GB}}{2\text{ MB}} = 65,536$ page table entries (a 512x reduction in TLB entries!).
   CPU Instruction -> [TLB Cache Check]
                             |
             +---------------+---------------+
             | (4KB Standard Pages)          | (2MB Huge Pages)
             v                               v
    TLB Miss Rate: High             TLB Miss Rate: Extremely Low
    Walks 4-Level Page Table        Single Hit in L1/L2 TLB
    Memory Latency: 15-40ns         Memory Latency: 1-3ns

While 2MB huge pages provide immense performance gains for compute-heavy, contiguous memory workloads, the automatic allocation daemon (khugepaged) creates severe problems for databases.


Why THP = always Cripples Transactional Databases

Enterprise Linux distributions (RHEL, AlmaLinux, Rocky Linux, Ubuntu) historically ship with THP set to [always]:

cat /sys/kernel/mm/transparent_hugepage/enabled
# [always] madvise never

In always mode:

  1. Memory Compaction Stalls: To allocate a contiguous 2MB physical block, the kernel must locate 512 contiguous free 4KB pages. If memory is fragmented, the kernel initiates synchronous memory compaction. Foreground database execution threads are placed into uninterruptible sleep (D state) while memory pages are copied and remapped.
  2. Aggressive Page Fault Overhead: In databases, writes are sparse (modifying an 8-byte row in an unmapped 4KB page). Under THP, the OS must allocate, zero, and clear an entire 2MB block, magnifying memory write amplification.
  3. Lock Contention: The khugepaged background compaction thread acquires the process mmap_lock (or mmap_sem), freezing concurrent client connections attempting to allocate memory.
       MariaDB Inbound Query (Modifying 1 Row)
                         |
                         v
       [Requests 4KB Page Allocation]
                         |
                         v
       [Kernel Intercepts: THP=always]
                         |
       [Attempts to Find 2MB Contiguous Block]
                         | (Memory Fragmented!)
                         v
       [SYNCHRONOUS MEMORY COMPACTION TRIGGERED]
                         |
       +-----------------+-----------------+
       |                                   |
       v                                   v
  khugepaged Locks Memory Bus        Query Thread Frozen (D State)
  Pages Copied in RAM for 2000ms     Latency Spikes to 8 Seconds!

By deploying optimized database architectures on Dedicated Servers in Pakistan, system engineers tune THP settings to eliminate compaction latency completely.


Step 1: Inspecting THP State and Compaction Stalls

Check the current THP configuration on your production host:

cat /sys/kernel/mm/transparent_hugepage/enabled
cat /sys/kernel/mm/transparent_hugepage/defrag

Measure how many times your server has stalled due to synchronous THP allocations using /proc/vmstat:

grep -E "thp_fault_alloc|thp_collapse_alloc|compact_stall|compact_fail" /proc/vmstat

Key warning indicators:

  • compact_stall > 0: Foreground application threads were blocked waiting for memory compaction.
  • thp_collapse_alloc_failed: The kernel wasted CPU cycles attempting to build huge pages but aborted.

Step 2: Choosing the Correct Policy: madvise vs. never

There are two safe approaches depending on your workload:

Completely turns off Transparent Huge Pages. The OS allocates standard 4KB pages uniformly, completely eliminating khugepaged and compaction overhead.

THP is only allocated to processes that explicitly request huge pages using the madvise(MADV_HUGEPAGE) system call (such as JVMs or machine learning libraries), while database engines like MariaDB and Redis run safely on standard 4KB pages.


Step 3: Enforcing Persistent THP Disablement

Runtime adjustments take effect immediately:

# Set THP to madvise (or never)
echo madvise > /sys/kernel/mm/transparent_hugepage/enabled

# Set defrag to never (ELIMINATES COMPACTION STALLS)
echo never > /sys/kernel/mm/transparent_hugepage/defrag

Because /sys is an ephemeral virtual filesystem, these settings revert upon server reboot. To make them permanent across reboots on systemd-based Linux systems (AlmaLinux 8/9, Ubuntu 22.04/24.04), create a dedicated systemd service /etc/systemd/system/disable-thp.service:

# /etc/systemd/system/disable-thp.service
[Unit]
Description=Disable Transparent Huge Pages (THP) for Database Performance
DefaultDependencies=no
After=sysinit.target local-fs.target
Before=mariadb.service mysql.service redis.service

[Service]
Type=oneshot
ExecStart=/bin/sh -c 'echo madvise > /sys/kernel/mm/transparent_hugepage/enabled && echo never > /sys/kernel/mm/transparent_hugepage/defrag'

[Install]
WantedBy=basic.target

Enable and start the service:

systemctl daemon-reload
systemctl enable --now disable-thp.service
systemctl status disable-thp.service

Verify persistence:

cat /sys/kernel/mm/transparent_hugepage/enabled
# always [madvise] never
cat /sys/kernel/mm/transparent_hugepage/defrag
# always defer defer+madvise madvise [never]

Step 4: Optional: Static Huge Pages (vm.nr_hugepages) for Extreme Performance

If you want the TLB advantages of 2MB huge pages without the instability of Transparent Huge Pages, configure Static Huge Pages. The Linux kernel pre-allocates contiguous 2MB blocks at boot time, avoiding all runtime compaction.

Allocate 16 GB of static huge pages (8,192 pages of 2MB each):

sysctl -w vm.nr_hugepages=8192
echo "vm.nr_hugepages = 8192" >> /etc/sysctl.d/99-hugepages.conf

Enable large pages inside MariaDB /etc/my.cnf.d/server.cnf:

[mariadb]
large-pages = ON

Real-World Latency Comparison: 100,000 Random Writes

Benchmarking an e-commerce database under intense write concurrency before and after disabling THP defragmentation:

Metric THP [always] / Defrag [always] THP [madvise] / Defrag [never] Improvement
Max Transaction Latency 7,850 ms (Severe Compaction Stall) 14 ms 99.8% Latency Reduction
System (sys) CPU Overhead 28.4% 1.8% 93.6% CPU Freed
Compaction Stalls per Hour 142 events 0 events Zero Lock Contention
P99 Response Consistency Highly Erratic Flat & Predictable Rock-Solid Stability

Tuning Transparent Huge Pages is a mandatory milestone for ensuring predictable, low-latency database throughput on enterprise Linux hosts.

Eliminate Database Stalls with NextGen Dedicated Servers

Experience rock-solid memory stability, zero compaction stalls, and dedicated hardware control with enterprise bare-metal architectures. Explore our high-RAM Dedicated Servers or host locally in Karachi and Islamabad on Dedicated Servers in Pakistan.