Linux Kernel TCP Timestamps and PAWS: Preventing Wrapped Sequence Numbers Across 100GbE Pakistan Transit

A technical architectural breakdown of RFC 7323 TCP Timestamps and Protect Against Wrapped Sequences (PAWS) in the Linux kernel on high-bandwidth, high-latency 40GbE and 100GbE transit links in Pakistan.

Linux Kernel TCP Timestamps and PAWS: Preventing Wrapped Sequence Numbers Across 100GbE Pakistan Transit

Modern enterprise networks, financial trading systems, and content delivery backbones across Pakistan increasingly operate over 10GbE, 40GbE, and 100GbE fiber optics interconnecting PTCL, TransWorld, and the Pakistan Internet Exchange (PKIX). However, running gigabit and multi-gigabit flows over high Bandwidth-Delay Product (BDP) international routes (e.g., Karachi to Frankfurt or Singapore) uncovers an inherent mathematical boundary in classical TCP: sequence number wrapping.

TCP sequence numbers are 32-bit unsigned integers, providing an address space of $2^{32} \approx 4.29$ billion octets (4.29 GB). On a 10Gbps link, this 32-bit sequence space cycles and wraps around in less than 3.4 seconds; on a 100Gbps link, it wraps in just 340 milliseconds. If an old, delayed packet arrives after sequence number rollover, a receiver could mistake the old packet for fresh payload data—causing silent stream corruption.

To prevent this catastrophically subtle issue, the Linux networking stack implements RFC 7323 TCP Extensions for High Performance, specifically TCP Timestamps and Protect Against Wrapped Sequences (PAWS). In this deep dive, we explore the internal mechanics of PAWS, diagnose NAT-induced timestamp drops, tune net.ipv4.tcp_timestamps, and deploy high-speed network topologies on Dedicated Servers.


The Mathematics of Sequence Number Rollover

Consider a sustained connection running over high-speed transit. The time $T_{wrap}$ required for a TCP sequence number to wrap completely around the 32-bit space ($4,294,967,296$ bytes) is calculated as:

$$T_{wrap} = \frac{2^{32} \times 8}{\text{Throughput (bits/sec)}}$$

Interface Speed Theoretical Max Throughput Sequence Wrap Time ($T_{wrap}$) Risk Without PAWS
100 Mbps 12.5 MB/s ~343.6 seconds (~5.7 min) Low (old packets expire via MSL)
1 Gbps 125 MB/s ~34.3 seconds Moderate
10 Gbps 1,250 MB/s ~3.43 seconds Critical (well within 2-minute MSL)
40 Gbps 5,000 MB/s ~858 milliseconds Extreme
100 Gbps 12,500 MB/s ~343 milliseconds Instantaneous

Under the standard Internet Maximum Segment Lifetime (MSL) of 120 seconds (or Linux’s 60-second default TCP_TIMEWAIT_LEN), any packet delayed by multi-path routing or transient routing flaps will arrive within the current sequence window if $T_{wrap} < \text{MSL}$. Without a secondary monotonic clock discriminator, the receiver accepts the stale segment as valid data.

       32-bit Sequence Space Rollover on 100GbE (340ms)
       
  Seq 0 ------------> Seq 2^31 ------------> Seq 2^32 (Wrap back to 0)
    |                                                |
    +---- Delayed Packet 'P' from Epoch 1 ---------->+ Receiver Window (Epoch 2)
                                                     * MISTAKEN FOR FRESH DATA!
                                                     (SILENT CORRUPTION)

By deploying carrier-grade Dedicated Servers in Pakistan tuned with PAWS, sequence wrapping is disambiguated by monotonic timestamp headers.


Mechanics of PAWS (Protect Against Wrapped Sequences)

TCP Timestamps add a 10-byte Option field to the TCP header:

  • TSval (Timestamp Value): Generated by the sending kernel, incrementing monotonically at a local clock rate.
  • TSecr (Timestamp Echo Reply): Reflected by the receiver from the most recently received packet.

The receiver maintains a per-connection state variable: TS.Recent. When a segment arrives with sequence number $S$ inside the acceptable receive window, the kernel evaluates:

$$\text{PAWS Condition: } \text{SEG.TSval} \ge \text{TS.Recent}$$

If the arriving segment’s timestamp is older than TS.Recent (taking modulo $2^{32}$ clock arithmetic into account), the segment is definitively identified as a delayed ghost segment from an earlier wrap cycle and is silently dropped:

/* kernel/net/ipv4/tcp_input.c */
static bool tcp_paws_discard(const struct sock *sk, const struct sk_buff *skb)
{
    const struct tcp_sock *tp = tcp_sk(sk);
    return ((s32)(tp->rcv_tsval - tp->ts_recent) < 0);
}

Diagnosing the Linux Kernel Sysctl Options

Linux manages TCP Timestamps via the net.ipv4.tcp_timestamps sysctl parameter. Depending on the kernel version (Linux 4.10+ and 6.x), this integer supports three operational modes:

sysctl net.ipv4.tcp_timestamps
  • 0 (Disabled): RFC 7323 Timestamps and PAWS disabled. Danger on >1Gbps interfaces.
  • 1 (Enabled - Default): Enables timestamps using randomized offsets per connection for privacy and anti-fingerprinting.
  • 2 (Enabled with Fixed Ticks): RFC 7323 enabled with predictable 1ms/10ms ticks (useful for high-frequency internal cluster latency benchmarking).

Inspect active timestamp usage and PAWS drops on your network interfaces using nstat and netstat:

# Monitor PAWS discards in real time
nstat -az | grep -i paws

# Sample output:
# TcpExtPAWSEstab 1420   0.0
# TcpExtPAWSSparse 32    0.0

TcpExtPAWSEstab counts valid established connections where old packets were successfully quarantined and discarded by PAWS.


The Dreaded NAT/Carrier-Grade NAT (CGNAT) Conflict

In Pakistan, mobile operators (Jazz, Zong, Telenor, Ufone) and many regional broadband providers route residential users through large-scale Carrier-Grade NAT (CGNAT) gateways.

If multiple client devices behind the same public CGNAT IP initiate TCP connections to your Linux web server:

  1. Client A sends a packet with high TSval. Server records TS.Recent = 500,000.
  2. Client B (booted 10 seconds ago) sends a SYN or data packet with low TSval = 1,000 over a recycled 5-tuple.
  3. If tcp_tw_recycle was enabled (deprecated in Linux 4.12), the server drops Client B’s SYN packet because $1,000 < 500,000$.
  4. Result: Entire offices or mobile ISP subnets fail to load websites hosted on that server.

To guarantee zero drops while maintaining full PAWS protection for high-speed flows, ensure tcp_tw_reuse is configured safely and modern timestamp offsets are active:

# Catastrophic setting - NEVER use this (removed in recent kernels)
# sysctl -w net.ipv4.tcp_tw_recycle=0

# Safe socket recycling for outgoing proxy/FastCGI connections
sysctl -w net.ipv4.tcp_tw_reuse=1

# Maintain RFC 7323 Timestamps enabled
sysctl -w net.ipv4.tcp_timestamps=1

Production Sysctl Hardening for 40GbE/100GbE Pakistan Backbone

Persist the optimal network stack tuning for high-throughput transit by appending the following to /etc/sysctl.d/99-network-paws.conf:

# /etc/sysctl.d/99-network-paws.conf

# Enable RFC 7323 TCP Timestamps and PAWS
net.ipv4.tcp_timestamps = 1

# Enable Window Scaling (RFC 7323) for buffers > 64KB
net.ipv4.tcp_window_scaling = 1

# Enable Selective Acknowledgments (RFC 2018)
net.ipv4.tcp_sack = 1

# Maximize socket memory buffers for high BDP transit
net.core.rmem_max = 67108864
net.core.wmem_max = 67108864
net.ipv4.tcp_rmem = 4096 87380 67108864
net.ipv4.tcp_wmem = 4096 65536 67108864

# Fair Queueing packet scheduler for pacing
net.core.default_qdisc = fq

# Modern congestion control algorithm
net.ipv4.tcp_congestion_control = bbr

Apply the parameters immediately without rebooting:

sysctl -p /etc/sysctl.d/99-network-paws.conf

Verify interface packet statistics and verify offloading features on your Mellanox/Intel NICs:

ethtool -k eth0 | grep -E "tcp-segmentation-offload|generic-segmentation-offload"

Hardware TSO/GSO modules natively preserve TCP timestamp option headers across segmented frames, maintaining wire-speed line rate across 100GbE transit.

Deploy Multi-Gigabit Infrastructure with NextGen Dedicated Servers

Eliminate packet drops, sequence wrapping, and buffer bottlenecks on high-throughput workloads. Leverage unmetered 10GbE and 100GbE uplinks on enterprise-class bare-metal servers. Discover our high-performance Dedicated Servers or deploy locally with Dedicated Servers in Pakistan for ultra-low latency.