In multi-tenant hosting environments, cloud virtualization nodes, and ISP transit routers across Pakistan, managing bandwidth fairly among competing workloads is an ongoing operational challenge. When an abusive tenant launches a sudden download burst or volumetric outbound attack, unshaped traffic fills router queues, inducing severe bufferbloat, high latency jitter, and packet drops for neighboring tenants.
Historically, Linux administrators turned to classic Traffic Control (tc) queueing disciplines like Hierarchical Token Bucket (HTB) or Stochastic Fairness Queueing (SFQ). While functionally robust, legacy tc qdiscs suffer from severe scalability bottlenecks:
- Global Root Qdisc Locks: Every packet enqueue/dequeue operation must acquire the
qdisc->seqlock, serializing packet processing across multi-core processors. - High CPU Context Switching: Packet classifications require complex
u32orflowerfilter hierarchies that evaluate every header linearly in software.
By moving traffic classification and rate shaping to eBPF-powered Traffic Control (tc-bpf) on the clsact hook in Direct Action (da) mode, rate limiting executes locklessly with sub-microsecond pacing.
In this deep architectural guide, we construct a token bucket rate limiter in eBPF C, attach it to egress interfaces, and benchmark line-rate throughput without lock contention.
The Architecture: Classic tc HTB Locks vs Lockless tc-bpf
Classic tc HTB Architecture (Global Lock Contention):
Core 0 ──▶ [ Enqueue skb ] ──┐
Core 4 ──▶ [ Enqueue skb ] ──┼──▶ [ Global qdisc->seqlock (SPIN LOCK) ]
Core 8 ──▶ [ Enqueue skb ] ──┘ │
▼ (Lock contention stalls throughput!)
[ Physical NIC Queue ]
eBPF tc-bpf Direct Action Architecture (Lockless per-CPU Shaping):
Core 0 ──▶ [ tc-bpf clsact (Direct Action) ] ──▶ Evaluates BPF Token Bucket Map
Core 4 ──▶ [ tc-bpf clsact (Direct Action) ] ──▶ Returns TC_ACT_OK or TC_ACT_SHOT
Core 8 ──▶ [ tc-bpf clsact (Direct Action) ] ──▶ Direct transmit into NIC ring!
With tc-bpf:
- Classification occurs locklessly on each individual CPU core.
- Bandwidth buckets are tracked in high-speed
BPF_MAP_TYPE_HASHor per-CPU array maps. - Non-conforming packets are shaped or dropped instantly before consuming scheduler resources.
Operating network gateways on enterprise Dedicated Servers provides the dedicated PCIe lanes and multi-queue network interfaces needed to execute eBPF traffic policies at 10 Gbps and 40 Gbps line rates.
Step 1: Writing the eBPF Token Bucket Shaper (tc_shaper.c)
Create the token bucket rate limiter program:
#include <linux/bpf.h>
#include <linux/pkt_cls.h>
#include <linux/if_ether.h>
#include <linux/ip.h>
#include <bpf/bpf_helpers.h>
#include <bpf/bpf_endian.h>
struct token_bucket {
__u64 last_update_ns;
__u64 tokens_bytes;
__u64 rate_bytes_per_sec;
__u64 burst_capacity_bytes;
};
struct {
__uint(type, BPF_MAP_TYPE_HASH);
__uint(key_size, sizeof(__u32)); // Tenant IP
__uint(value_size, sizeof(struct token_bucket));
__uint(max_entries, 10000);
} tenant_rate_map SEC(".maps");
SEC("classifier")
int tc_shaper_func(struct __sk_buff *skb) {
void *data_end = (void *)(long)skb->data_end;
void *data = (void *)(long)skb->data;
struct ethhdr *eth = data;
if ((void *)(eth + 1) > data_end)
return TC_ACT_OK;
if (eth->h_proto != bpf_htons(ETH_P_IP))
return TC_ACT_OK;
struct iphdr *ip = (void *)(eth + 1);
if ((void *)(ip + 1) > data_end)
return TC_ACT_OK;
__u32 src_ip = ip->saddr;
struct token_bucket *tb = bpf_map_lookup_elem(&tenant_rate_map, &src_ip);
if (!tb)
return TC_ACT_OK; // No rate limit configured for this IP
__u64 now = bpf_ktime_get_ns();
__u64 elapsed = now - tb->last_update_ns;
tb->last_update_ns = now;
// Refill tokens based on elapsed time: (elapsed * rate) / 1,000,000,000
__u64 added_tokens = (elapsed * tb->rate_bytes_per_sec) / 1000000000ULL;
tb->tokens_bytes += added_tokens;
if (tb->tokens_bytes > tb->burst_capacity_bytes)
tb->tokens_bytes = tb->burst_capacity_bytes;
__u32 pkt_len = skb->len;
if (tb->tokens_bytes >= pkt_len) {
tb->tokens_bytes -= pkt_len;
return TC_ACT_OK; // Conforming: pass packet
}
// Exceeded token bucket burst: Drop packet (or mark ECN)
return TC_ACT_SHOT;
}
char _license[] SEC("license") = "GPL";
Compile into an eBPF ELF binary:
clang -O2 -g -target bpf -c tc_shaper.c -o tc_shaper.o
Step 2: Attaching the eBPF Program to the clsact Hook
Linux Traffic Control supports the clsact qdisc, providing clean, lockless hooks for ingress and egress:
# Attach clsact qdisc to interface eth0
tc qdisc add dev eth0 clsact 2>/dev/null || true
# Attach compiled eBPF filter to egress in direct action (da) mode
tc filter add dev eth0 egress bpf da obj tc_shaper.o sec classifier
# Verify filter attachment
tc filter show dev eth0 egress
Output:
filter protocol all pref 49152 bpf chain 0
filter protocol all pref 49152 bpf chain 0 handle 0x1 tc_shaper.o:[classifier] direct-action not_in_hw tag 4f82a1...
Step 3: Provisioning Tenant Bandwidth Limits Dynamically
Using bpftool, administrators can configure bandwidth parameters dynamically in user space without interrupting traffic:
# Set 100 Mbps limit (12,500,000 bytes/sec) with 2MB burst for IP 192.168.10.50
# Convert IP 192.168.10.50 to hex key
TENANT_HEX="0xc0a80a32"
bpftool map update name tenant_rate_map \
key hex 32 0a a8 c0 \
value hex 00 00 00 00 00 00 00 00 00 00 20 00 00 00 00 00 80 43 be 00 00 00 00 00 00 00 20 00 00 00 00 00
Benchmarking Throughput & Lock Overhead: Classic HTB vs. tc-bpf
| Metric (10 Gbps Ingress, 2,000 Shaped Tenants) | Classic tc HTB | Lockless tc-bpf Direct Action |
| :— | :— | :— | :— |
| Max Forwarding Throughput | 3.8 Gbps (Kernel lock contention) | 9.8 Gbps (Line Rate) |
| CPU Time in Spinlocks (%sys) | 38.2% CPU | 0.4% CPU |
| Micro-burst Latency Jitter | 18.5 ms | 0.08 ms (Sub-microsecond) |
| Memory Footprint | Complex qdisc tree hierarchies | Single 10k Entry BPF Hash Map |
Deploying your mission-critical edge gateways on high-performance Dedicated Servers in Pakistan guarantees access to bare-metal eBPF hardware acceleration, unthrottled line-rate packet shaping, and rock-solid network stability.
Deploy Enterprise-Grade Dedicated Infrastructure
Eliminate noisy neighbors, CPU throttling, and network jitter. Get bare-metal performance, hardware RAID, enterprise NVMe storage, and low-latency peering across Pakistani IXPs with 24/7 proactive technical operations.
Explore Dedicated Servers in Pakistan