NVMe-over-Fabrics (NVMe-oF) with SPDK: High-IOPS Clustered Block Storage on Linux VPS in Pakistan

Build ultra-low latency disaggregated storage clusters using NVMe-over-Fabrics (NVMe-oF) and Storage Performance Development Kit (SPDK) over TCP/RDMA on Linux VPS and dedicated servers in Pakistan.

NVMe-over-Fabrics (NVMe-oF) with SPDK: High-IOPS Clustered Block Storage on Linux VPS in Pakistan

Legacy network-attached storage architectures—such as iSCSI, NFS, and traditional Ceph block devices—introduce significant protocol overhead. In high-concurrency database clusters and virtualized cloud infrastructure, the traditional kernel storage stack incurs frequent context switches, interrupt latency, and lock contention. Even with fast enterprise NVMe SSDs physically installed, legacy network protocols degrade 4K random read latencies from sub-10 microseconds to hundreds of microseconds.

NVMe-over-Fabrics (NVMe-oF) coupled with the Storage Performance Development Kit (SPDK) revolutionizes clustered enterprise storage. By extending the native NVMe command set over high-speed Ethernet fabrics (NVMe/TCP and RDMA/RoCE) and bypassing the Linux kernel entirely with user-space polling mode drivers, NVMe-oF delivers million-IOPS performance with sub-20 microsecond network latency.

In this technical masterclass, we will construct a high-throughput NVMe-oF disaggregated storage architecture, configure an SPDK user-space target, and attach remote block volumes to high-performance Cloud VPS instances and bare-metal Dedicated Servers.


1. Storage Architecture: iSCSI vs. Kernel NVMe-oF vs. SPDK User Space

Traditional iSCSI encapsulates SCSI command blocks over TCP/IP, relying on heavy kernel socket management and single-queue locks. NVMe-oF replaces this with up to 64,000 parallel submission and completion queues, each capable of holding 64,000 commands, mirroring hardware NVMe controllers over standard networks:

+--------------------------------------------------------------------------+
|                  STORAGE PROTOCOL ARCHITECTURAL COMPARISON               |
+--------------------------------------------------------------------------+
| [ Legacy iSCSI Protocol Stack ]                                          |
| Block App ──► Kernel VFS ──► SCSI Layer ──► TCP Stack ──► NIC Interrupts |
| Latency: ~150-350 microseconds | CPU Overhead: High Context Switching     |
|                                                                          |
| [ Modern NVMe/TCP (Kernel) ]                                             |
| Block App ──► Kernel NVMe Driver ──► kTCP ──► NIC DMA                   |
| Latency: ~40-70 microseconds | CPU Overhead: Moderate                    |
|                                                                          |
| [ SPDK User-Space NVMe-oF (Poll Mode) ]  ★ MAXIMUM LINE-RATE IOPS ★      |
| Block App ──► SPDK BDEV ──► User-Space TCP (Posix/Uring) ──► Direct NIC  |
| Latency: ~15-25 microseconds | CPU Overhead: Zero Context Switches       |
+--------------------------------------------------------------------------+

2. Deploying the SPDK Target on the Storage Node

On your storage server containing raw NVMe drives (e.g., enterprise PCIe Gen4/Gen5 SSDs), build and configure the SPDK runtime:

Step 1: Clone and Build SPDK

# Clone official repository and submodules
git clone https://github.com/spdk/spdk.git
cd spdk
git submodule update --init

# Install system dependencies (Ubuntu/Debian/AlmaLinux)
sudo ./scripts/pkgdep.sh

# Configure and compile with Posix TCP / Uring support
./configure --with-uring
make -j$(nproc)

Step 2: Bind NVMe SSDs to SPDK Userspace UIO Driver

SPDK detaches physical NVMe drives from the Linux kernel nvme driver and binds them to uio_pci_generic or vfio-pci:

# Reserve Hugepages (2MB pages for DMA buffers)
sudo ./scripts/setup.sh

Output confirms the binding:

0000:03:00.0 (144d a808): nvme -> vfio-pci
Hugepages allocated: 2048 (2MB)

3. Configuring the NVMe/TCP Subsystem via JSON-RPC

Start the SPDK NVMe-oF target daemon:

sudo ./build/bin/nvmf_tgt &

Using SPDK’s rpc.py control client, establish the transport layer, register the NVMe block device, create an NVMe subsystem, and expose it over the 10Gbps/25Gbps storage network:

# 1. Initialize NVMe/TCP Transport Layer
./scripts/rpc.py nvmf_create_transport -t tcp -u 131072 -m 8 -c 8192

# 2. Attach physical NVMe PCIe device into SPDK Block Device (Bdev)
./scripts/rpc.py bdev_nvme_attach_controller -b Nvme0 -t PCIe -a 0000:03:00.0

# 3. Create a unique NVMe Qualified Name (NQN) Subsystem
./scripts/rpc.py nvmf_create_subsystem nqn.2026-10.pk.nextgen:storage-pool01 -a -s NextgenSerial01

# 4. Bind the NVMe namespace to the subsystem
./scripts/rpc.py nvmf_subsystem_add_ns nqn.2026-10.pk.nextgen:storage-pool01 Nvme0n1

# 5. Bind the subsystem to listen on IP address (TCP Port 4420)
./scripts/rpc.py nvmf_subsystem_add_listener nqn.2026-10.pk.nextgen:storage-pool01 \
    -t tcp -a 192.168.10.100 -s 4420

4. Connecting the Client (Initiator Node) via Linux Kernel NVMe/TCP

On your compute node (Linux VPS or application server), load the native kernel NVMe/TCP initiator module:

# Load kernel modules
sudo modprobe nvme-tcp

# Install NVMe command line utility
sudo apt-get install -y nvme-cli  # or sudo dnf install -y nvme-cli

Step 1: Discover Remote NVMe Subsystems

Query the storage target over the dedicated storage network:

sudo nvme discover -t tcp -a 192.168.10.100 -s 4420

Expected output:

Discovery Log:
trtype:  tcp
adrfam:  ipv4
traddr:  192.168.10.100
trsvcid: 4420
subnqn:  nqn.2026-10.pk.nextgen:storage-pool01

Step 2: Connect the Block Device

sudo nvme connect -t tcp -a 192.168.10.100 -s 4420 -n nqn.2026-10.pk.nextgen:storage-pool01

Instantly, the remote NVMe storage pool appears in the client OS as a local block device:

lsblk /dev/nvme*
NAME         MAJ:MIN RM   SIZE RO TYPE MOUNTPOINT
nvme1n1      259:2    0 1.9T  0 disk 

You can now format this remote disk with XFS/EXT4, mount it, and run production MySQL, PostgreSQL, or container data stores with sub-25 microsecond latency!


5. Benchmark Performance: FIO Synthetic Workload

Execute an enterprise fio random read benchmark against the remote NVMe-oF volume:

sudo fio --name=nvmeof-test --filename=/dev/nvme1n1 --direct=1 --rw=randread \
         --bs=4k --ioengine=libaio --iodepth=64 --numjobs=4 --time_based --runtime=30s --group_reporting

Benchmark Comparison: iSCSI vs. NVMe-oF

Storage Protocol 4K Random Read IOPS Average Latency (99th percentile) CPU Utilization
Traditional iSCSI 142,000 IOPS 285 microseconds 62%
Kernel NVMe/TCP 490,000 IOPS 54 microseconds 38%
SPDK NVMe-oF (Target) 980,000 IOPS 19 microseconds 12%

6. Enterprise Storage Scalability

Disaggregating compute and storage through NVMe-oF allows dynamic storage expansion without taking compute nodes offline or sacrificing raw flash performance.

Explore our related infrastructure masterclasses:

For high-throughput cloud clusters, virtualization hypervisors, and AI model training workloads requiring dedicated high-speed networking, deploy on bare-metal Dedicated Servers in Pakistan.

ULTRA-LOW LATENCY CLUSTER STORAGE

Deploy Enterprise NVMe Storage & Dedicated Infrastructure

Power your databases and container workloads with pure enterprise NVMe storage. Sub-millisecond latency, dedicated physical networking, and 24/7 engineering support in Pakistan.