Enterprise financial applications, telecommunication billing systems, and mission-critical SaaS platforms in Pakistan require true active-active database high availability. To satisfy Disaster Recovery (DR) and business continuity mandates, organizations deploy two primary database nodes across geographically separated data centers—typically one in Karachi (near international submarine cable landings) and a secondary node in Lahore or Islamabad (for inland geographical redundancy).
However, MariaDB Galera Cluster relies on a strict majority quorum consensus model ((N/2) + 1). In a naive 2-node cluster spanning Karachi and Lahore:
- Total nodes:
N = 2. - Quorum requirement:
(2/2) + 1 = 2 votes.
If the inter-city WAN connection between Karachi and Lahore experiences an undersea cable fault, fiber cut, or transient routing drop, both nodes drop into non-primary state. Neither node can achieve the required 2 votes independently. To protect data integrity from catastrophic split-brain divergence, both nodes reject all write transactions, bringing down the entire corporate application.
Adding a full third database server solely to break quorum ties is cost-prohibitive, requiring expensive CPU licenses, massive RAM allocations, and multi-terabyte NVMe storage. The architectural solution is the MariaDB Galera Arbitrator (garbd). garbd is a lightweight, stateless witness daemon that participates in cluster quorum voting without storing any database tables or executing transactions.
Deploying primary database nodes on high-IOPS Dedicated Servers and localized Dedicated Servers in Pakistan coupled with an offsite garbd witness guarantees automatic failover and zero split-brain corruption across Pakistani data centers.
1. Architectural Anatomy: 2-Node Split-Brain vs 2-Node + garbd Witness
Comparing failure states under an inter-city WAN partition shows why the arbitrator is essential:
Naive 2-Node Cluster (Without garbd - WAN Partition Failure):
[Node 1: Karachi (1 vote)] <==== WAN FIBER CUT ====> [Node 2: Lahore (1 vote)]
Quorum Needed: 2 votes.
Result: Node 1 has 1/2 votes (NO QUORUM) -> Rejects Writes!
Node 2 has 1/2 votes (NO QUORUM) -> Rejects Writes!
TOTAL OUTAGE FOR ALL CORPORATE USERS!
Resilient 3-Vote Topology (With garbd Witness):
[Node 1: Karachi DB] <══════════════ WAN ══════════════> [Node 2: Lahore DB]
│ │
│ (Galera wsrep) │ (Galera wsrep)
▼ ▼
┌────────────────────────────────┐
│ Cloud / 3rd DC Witness (garbd) │
│ (Islamabad / Third Cloud) │
│ Stateless (0 MB DB) │
└────────────────────────────────┘
Total Votes: 3 (Karachi + Lahore + garbd)
Quorum Needed: (3/2) + 1 = 2 votes!
WAN Partition Scenario:
If Lahore is disconnected, Karachi + garbd = 2/3 votes (MAJORITY QUORUM ACHIEVED!)
Karachi automatically stays Primary, zero downtime, zero data corruption!
2. Resource Comparison: Full Third DB Node vs garbd Witness
| Infrastructure Requirement | Full MariaDB Replica Node | MariaDB Galera Arbitrator (garbd) |
Operational Savings |
|---|---|---|---|
| Server Hardware Spec | 64-core AMD EPYC, 256GB RAM, 4TB NVMe | Minimal Cloud Instance (1 vCPU, 512MB RAM) | 95% Hardware Cost Reduction |
| Disk Storage Needed | 4,000 GB (Full Database Copy) | 0 GB (Stateless, no DB storage) | 100% Storage Savings |
| Replication Network Bandwidth | Receives all row images & SST | Receives only small quorum heartbeats | 98% WAN Bandwidth Saved |
| Software Maintenance | Complex backups, schema sync | Simple systemd service | Zero Administrative Overhead |
| Quorum Voting Power | 1 Vote | 1 Full Vote (Equal to DB Node) | Complete Split-Brain Immunity |
3. Configuring MariaDB Galera on Karachi and Lahore Nodes
On Node 1 (Karachi) and Node 2 (Lahore), configure Galera replication in /etc/my.cnf.d/server.cnf:
# /etc/my.cnf.d/server.cnf
[galera]
wsrep_on = ON
wsrep_provider = /usr/lib64/galera-4/libgalera_smm.so
wsrep_cluster_name = "nextgen_pk_galera_cluster"
# Include both DB nodes and the garbd arbitrator in the address list
wsrep_cluster_address = "gcomm://10.10.1.10:4567,10.20.1.20:4567,10.30.1.30:4567"
# Local node identification (Karachi Node example)
wsrep_node_name = "db-node-karachi-01"
wsrep_node_address = "10.10.1.10"
# WAN Latency and Timeout Optimization for Inter-City Links
wsrep_provider_options = "evs.keepalive_period=PT1S;evs.suspect_timeout=PT5S;evs.inactive_timeout=PT15S;evs.install_timeout=PT15S"
# Storage Engine & Safety
default_storage_engine = InnoDB
innodb_autoinc_lock_mode = 2
innodb_flush_log_at_trx_commit = 2
4. Installing and Deploying garbd on the Third Witness Node
On the third isolated host (located in an independent cloud region or tertiary data center, e.g., Islamabad or a low-cost cloud node at 10.30.1.30), install the Galera Arbitrator package:
# On RHEL/AlmaLinux:
dnf install -y MariaDB-backup galera-4
# On Debian/Ubuntu:
apt-get install -y galera-arbitrator-4
Configure /etc/sysconfig/garb (or /etc/default/garbd on Ubuntu):
# /etc/sysconfig/garb
# Galera Arbitrator configuration
# Cluster name must match exactly
GALERA_CLUSTER="nextgen_pk_galera_cluster"
# Addresses of the two primary database nodes
GALERA_NODES="10.10.1.10:4567,10.20.1.20:4567"
# Optional Galera provider parameters matching inter-city timeouts
GALERA_OPTIONS="evs.keepalive_period=PT1S;evs.suspect_timeout=PT5S;evs.inactive_timeout=PT15S"
# Arbitrator listening port
GALERA_PORT="4567"
# Path to log file
LOG_FILE="/var/log/garbd.log"
Start and enable the garbd systemd service:
systemctl daemon-reload
systemctl enable --now garb
systemctl status garb
5. Live Quorum Verification & Partition Simulation
Check cluster health and active quorum size from either database node in MariaDB:
SHOW STATUS LIKE 'wsrep_cluster_size';
-- Expected Output:
-- | wsrep_cluster_size | 3 | (2 DB nodes + 1 garbd)
SHOW STATUS LIKE 'wsrep_cluster_status';
-- Expected Output:
-- | wsrep_cluster_status | Primary |
Simulating a Karachi-Lahore Network Partition:
Simulate an inter-city fiber cut by blocking traffic between Karachi and Lahore using iptables:
# On Lahore DB Node (10.20.1.20)
iptables -A INPUT -s 10.10.1.10 -j DROP
iptables -A OUTPUT -d 10.10.1.10 -j DROP
Now inspect Karachi (10.10.1.10):
- Karachi still sees
garbd(10.30.1.30). - Quorum votes: Karachi (1) + garbd (1) = 2/3 votes.
- Karachi maintains
wsrep_cluster_status: Primaryand continues processing eCommerce orders without hesitation!
Lahore sees only itself (1/3 votes), gracefully transitions to Non-Primary, and drops connections until connectivity is restored, completely eliminating data corruption.
Build Disaster-Resilient High Availability Database Clusters
Protect your enterprise applications against data center outages and network disruptions. Build your active-active MariaDB Galera clusters across Pakistan on NextGen's enterprise Dedicated Servers and low-latency Dedicated Servers in Pakistan featuring dedicated private VLANs, hardware RAID, and 24/7 database operations support.
