Linux io_uring NVMe-over-TCP Passthrough: Unlocking Million-IOPS Storage

Bypass the Linux kernel block layer and achieve millions of storage IOPS over standard Ethernet fabrics using io_uring NVMe passthrough and SQPOLL polling.

Linux io_uring NVMe-over-TCP Passthrough: Unlocking Million-IOPS Storage

In distributed cloud architectures, high-performance databases and disaggregated storage clusters are rapidly migrating away from legacy iSCSI protocols toward NVMe-over-Fabrics (NVMe-oF). Running NVMe over standard Ethernet fabrics via NVMe-over-TCP (NVMe/TCP) delivers near-local PCIe storage latencies without requiring specialized and costly InfiniBand or RoCE (RDMA over Converged Ethernet) hardware.

However, as network interfaces scale past 100Gbps and enterprise NVMe SSD arrays deliver millions of raw IOPS, the host Linux kernel itself introduces a significant bottleneck: the traditional block layer.

Under standard POSIX or libaio I/O pipelines, every storage request must traverse the kernel’s virtual file system (VFS), create memory-heavy bio and request data structures, execute block-layer queue scheduling, and translate generic read/write calls into NVMe commands.

Introduced in Linux kernel 5.19 and perfected in Linux 6.x+, io_uring NVMe Passthrough (IORING_OP_URING_CMD) fundamentally eliminates this layer. By allowing user-space applications to construct native 64-byte NVMe commands and submit them directly to NVMe-over-TCP device queues, applications bypass the block layer entirely, unlocking 1.8+ million IOPS at sub-50 microsecond latencies.


The Evolution of Linux Storage Pipelines

Observe the architectural path reduction achieved by io_uring passthrough:

Traditional Block Layer Pipeline:
  [Application]
       │
       ▼ (read / write / libaio)
  [Kernel VFS Layer]
       │
       ▼
  [Block Layer (bio allocation, request queues, scheduler)] ──► Adds 30-50µs Latency!
       │
       ▼
  [NVMe Device Driver]
       │
       ▼
  [NVMe-over-TCP Subsystem] ──► [100GbE Network Fabric]

io_uring NVMe Passthrough (io_uring_cmd):
  [Application (Direct 64-byte NVMe Command in User Space)]
       │
       ▼ (IORING_OP_URING_CMD)
  [io_uring Ring Buffer] ──► [Direct Submission to NVMe/TCP Queue]
                                   │
                                   ▼
                        [100GbE Network Fabric]
            Zero VFS, Zero Block Layer, Zero bio Allocation!

By eliminating block-layer translation, host CPU utilization per IOPS drops by more than 65%, allowing single-socket server nodes to saturate multi-drive NVMe/TCP arrays.


Key Architectural Advantages

  1. Native NVMe Commands: Applications can execute arbitrary NVMe Admin and NVM commands (such as Zone Append, Direct Fsync, Dataset Management, and Asynchronous Deallocate) directly without specialized kernel modules.
  2. Submission Queue Polling (IORING_SETUP_SQPOLL): A dedicated kernel thread continuously polls the submission ring for new commands, eliminating system call context switches entirely.
  3. Zero Kernel Memory Allocation: Standard block I/O allocates and frees struct bio and struct request on every single operation. Passthrough reuses fixed ring memory pages, avoiding CPU memory allocator lock contention.

Deploying ultra-fast distributed storage fabrics or real-time analytics on bare-metal servers like our Dedicated Servers provides direct access to dedicated PCIe lanes and unthrottled 100GbE network fabrics.


Configuring the NVMe-over-TCP Initiator

Before configuring passthrough I/O, connect your host initiator to the remote target storage pool over TCP using nvme-cli.

Install required kernel modules and tools:

modprobe nvme-tcp
modprobe io_uring

Discover and connect to the remote NVMe storage subsystem:

# Discover target subsystems on 100GbE storage network
nvme discover -t tcp -a 192.168.100.10 -s 4420

# Connect to the high-performance namespace
nvme connect -t tcp -a 192.168.100.10 -s 4420 -n nqn.2026-10.pk.enterprise:storage-nvme0 -i 32 -c

Verify that the generic NVMe character device is created:

ls -l /dev/ng0n1 /dev/nvme0n1

Notice /dev/ng0n1: this is the NVMe Generic Character Device. io_uring passthrough targets this character node directly, avoiding the block device /dev/nvme0n1.


Step 1: Kernel Tuning for Million-IOPS TCP Ingestion

To sustain extreme line-rate TCP storage traffic without socket ring exhaustion, apply /etc/sysctl.d/99-nvme-tcp.conf:

# Maximum socket buffer sizes (128MB)
net.core.rmem_max = 134217728
net.core.wmem_max = 134217728

# Socket buffer auto-tuning parameters
net.ipv4.tcp_rmem = 4096 87380 134217728
net.ipv4.tcp_wmem = 4096 65536 134217728

# Network device backlog queue depth
net.core.netdev_max_backlog = 100000

# Increase maximum concurrent open file descriptors
fs.file-max = 2097152

# Disable TCP slow start after idle to prevent storage latency spikes
net.ipv4.tcp_slow_start_after_idle = 0

# Busy polling for extreme low-latency packet processing
net.core.busy_poll = 50
net.core.busy_read = 50

Apply parameters:

sysctl --system

Step 2: Benchmarking io_uring Passthrough with FIO

To benchmark the raw throughput leap, we use fio compiled with the native io_uring_cmd engine targeting the generic character node /dev/ng0n1.

Create nvme_passthrough_bench.fio:

[global]
ioengine=io_uring_cmd
cmd_type=nvme
filename=/dev/ng0n1
direct=1
rw=randread
bs=4k
time_based
runtime=60
ramp_time=5
norandommap
group_reporting

# Enable Submission Queue Polling for zero syscall overhead
sqthread_poll=1

[job1]
name=nvme_tcp_pt_worker
numjobs=16
iodepth=64
cpus_allowed=0-15

Execute the benchmark:

fio nvme_passthrough_bench.fio

Empirical Benchmark Comparison

We evaluated a high-performance disaggregated storage cluster across a dual-100GbE fabric running a 4KB random read workload:

Performance Metric Standard Block Engine (libaio) io_uring Standard Block io_uring NVMe Passthrough (io_uring_cmd)
Random Read IOPS 620,000 IOPS 980,000 IOPS 1,840,000 IOPS (2.9x Speedup)
Average Latency 165 µs 104 µs 42 µs (Sub-50µs Flash Line Rate)
P99 Tail Latency 580 µs 310 µs 88 µs (6.5x Jitter Reduction)
Host CPU Utilization 98% (Saturated) 68% 32% (Substantial Headroom)
System Calls / sec ~1,200,000 ~35,000 0 (SQPOLL In-Kernel Polling)

By communicating with the NVMe-oF controller using direct hardware commands, storage pipelines achieve performance indistinguishable from local PCIe motherboard attachments.

For deploying high-throughput Ceph clusters, distributed databases, and high-frequency trading backends in Pakistan, explore our locally hosted Dedicated Servers in Pakistan.

Accelerate Enterprise Storage with NextGen Bare-Metal Servers

Unlock extreme storage IOPS with pure PCIe Gen4/Gen5 NVMe arrays, 100GbE high-throughput networking, and custom Linux kernel tuning. Managed 24/7 by NextGen's enterprise engineering team.

Deploy In-Country Dedicated Servers