Native NGINX rate limiting (limit_req_zone) relies on shared memory zones (shm) local to each individual NGINX host. While effective on single servers, this local memory model fails in distributed enterprise architectures where multiple load-balanced reverse proxy nodes front large application fleets across Pakistan.
When requests are distributed across four NGINX edge instances, an attacker or abusive client can exceed rate limits by a factor of four simply by cycling connections across the cluster. Furthermore, local memory limits reset whenever an NGINX worker restarts or reloads configuration.
By integrating NGINX (via OpenResty or lua-nginx-module) with a centralized, in-memory Redis Cluster, engineering teams can enforce global Token Bucket and Sliding Window rate limits across thousands of concurrent nodes with sub-millisecond overhead.
The Architecture of Distributed Redis Rate Limiting
Rather than checking local shared memory, the NGINX access phase triggers an atomic Redis Lua script evaluation that updates the caller’s token bucket counter across the cluster.
+-------------------------------------------------------------+
| Incoming Traffic via Cloudflare / CDN |
+-------------------------------------------------------------+
|
[Header: CF-Connecting-IP / X-Real-IP]
|
+-----------------+-----------------+
| |
v v
+--------------------------+ +--------------------------+
| NGINX Edge Node 1 | | NGINX Edge Node 2 |
| (OpenResty / Lua) | | (OpenResty / Lua) |
+--------------------------+ +--------------------------+
\ /
\ /
v v
+-------------------------------------------------------------+
| Centralized High-Availability Redis |
| |
| [Atomic Lua Token Bucket Script] |
| - Key: "ratelimit:<client_ip>:<endpoint>" |
| - Tracks tokens, replenishment rate, and expiry |
+-------------------------------------------------------------+
|
[Token Available? (HTTP 200 vs 429)]
When deployed on high-throughput Dedicated Servers in Pakistan, hosting Redis on a private low-latency gigabit network ensures rate limit evaluations complete in under 300 microseconds.
Step 1: Handling Real Client IPs behind Cloudflare and Proxies
If your domain sits behind Cloudflare, NGINX will see Cloudflare’s edge IP instead of the end-user IP unless properly mapped. In /etc/nginx/conf.d/realip.conf:
# Trust Cloudflare edge network IP ranges
set_real_ip_from 173.245.48.0/20;
set_real_ip_from 103.21.244.0/22;
set_real_ip_from 103.22.200.0/22;
set_real_ip_from 103.31.4.0/22;
set_real_ip_from 141.101.64.0/18;
set_real_ip_from 108.162.192.0/18;
set_real_ip_from 190.93.240.0/20;
set_real_ip_from 188.114.96.0/20;
set_real_ip_from 197.234.240.0/22;
set_real_ip_from 198.41.128.0/17;
set_real_ip_from 162.158.0.0/15;
set_real_ip_from 104.16.0.0/13;
set_real_ip_from 104.24.0.0/14;
set_real_ip_from 172.64.0.0/13;
set_real_ip_from 131.0.72.0/22;
# Extract real client IP
real_ip_header CF-Connecting-IP;
real_ip_recursive on;
Step 2: The Atomic Redis Token Bucket Lua Script
To prevent race conditions during high-concurrency access, rate limit checks and token replenishment must execute atomically. Save /etc/nginx/lua/rate_limit.lua:
local redis = require "resty.redis"
local red = redis:new()
red:set_timeout(100) -- 100ms timeout to prevent hanging requests
local ok, err = red:connect("10.0.0.80", 6379)
if not ok then
ngx.log(ngx.ERR, "Failed to connect to Redis: ", err)
return -- Fail open if Redis is unreachable
end
local client_ip = ngx.var.remote_addr
local uri = ngx.var.uri
local key = "rate:" .. client_ip .. ":" .. uri
local limit = 60 -- Maximum tokens in bucket (60 requests)
local rate = 1 -- Replenish 1 token per second
local now = ngx.time()
-- Atomic Token Bucket Script
local script = [[
local key = KEYS[1]
local limit = tonumber(ARGV[1])
local rate = tonumber(ARGV[2])
local now = tonumber(ARGV[3])
local data = redis.call('HMGET', key, 'tokens', 'last_updated')
local tokens = tonumber(data[1])
local last_updated = tonumber(data[2])
if not tokens then
tokens = limit - 1
last_updated = now
redis.call('HMSET', key, 'tokens', tokens, 'last_updated', last_updated)
redis.call('EXPIRE', key, 120)
return 1
else
local delta = math.max(0, now - last_updated)
tokens = math.min(limit, tokens + delta * rate)
last_updated = now
if tokens >= 1 then
tokens = tokens - 1
redis.call('HMSET', key, 'tokens', tokens, 'last_updated', last_updated)
redis.call('EXPIRE', key, 120)
return 1
else
return 0
end
end
]]
local res, err = red:eval(script, 1, key, limit, rate, now)
red:set_keepalive(10000, 100) -- Connection pooling: 10s idle, max 100 connections
if res == 0 then
ngx.header["Retry-After"] = "10"
ngx.header["Content-Type"] = "application/json"
ngx.status = ngx.HTTP_TOO_MANY_REQUESTS
ngx.say('{"error": "Too Many Requests", "message": "Distributed rate limit exceeded. Please throttle your queries."}')
ngx.exit(ngx.HTTP_TOO_MANY_REQUESTS)
end
Step 3: Integrating the Lua Hook into NGINX Server Blocks
Inject the Lua check directly into your API gateway endpoints:
server {
listen 443 ssl http2;
server_name api.example.pk;
ssl_certificate /etc/letsencrypt/live/api.example.pk/fullchain.pem;
ssl_certificate_key /etc/letsencrypt/live/api.example.pk/privkey.pem;
location /api/v2/ {
# Execute centralized rate limit check before proxying
access_by_lua_file /etc/nginx/lua/rate_limit.lua;
proxy_pass http://backend_api_cluster;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
}
}
Step 4: Load Testing and Performance Verification
Verify rate limiting enforcement using wrk or hey from a remote client:
# Send 100 requests in 2 seconds from 10 concurrent threads
hey -n 100 -c 10 https://api.example.pk/api/v2/checkout
Verify that the response returns 60 HTTP 200 OK responses followed by 40 HTTP 429 Too Many Requests.
Deploying multi-node API gateways on bare-metal Dedicated Servers provides unmetered network pipelines, dedicated hardware RAM, and private low-latency vSwitches to execute real-time security rules at line rate.
Need Enterprise Dedicated Infrastructure in Pakistan?
Deploy mission-critical, bare-metal infrastructure optimized for low-latency throughput, hardware RAID/NVMe resilience, and 24/7 proactive management.
