Traditional MySQL and MariaDB primary-replica setups suffer from asynchronous replication lag. When a database node unexpectedly crashes during high-traffic checkout events on Pakistani e-commerce platforms or enterprise banking portals, failover mechanisms inevitably risk silent data loss, duplicate transaction processing, or painful manual interventions.
MariaDB Galera Cluster solves this with synchronous, multi-master replication. Every node in a Galera cluster is an active read-and-write master. Transactions are validated synchronously across all cluster nodes using certification-based replication, guaranteeing zero replication lag, automatic node failover, and strict ACID consistency across the entire database topology.
In this practical engineering guide, we walk through configuring a 3-node MariaDB Galera 10.11+ cluster on Ubuntu/AlmaLinux, handling quorum arbitration to prevent split-brain scenarios, and setting up an automated health check proxy on Cloud VPS instances and bare-metal Dedicated Servers.
1. Galera Replication Architecture: Certification-Based Synchrony
Unlike traditional MySQL binary log replication, Galera intercepts transactions at commit time using the wsrep (Write Set Replication) API.
+--------------------------------------------------------------------------+
| MARIADB GALERA CLUSTER TOPOLOGY |
+--------------------------------------------------------------------------+
| Client Application (Laravel / WooCommerce / Fintech API) |
| │ |
| ▼ (Read/Write Queries) |
| [ HAProxy / Keepalived Layer-7 Load Balancer ] |
| │ |
| ├───────────────────────┬─────────────────────────┐ |
| ▼ (Port 3306) ▼ (Port 3306) ▼ (Port 3306) |
| [ Node 1: Karachi ] [ Node 2: Lahore ] [ Node 3: Islamabad ] |
| 192.168.1.10 192.168.1.20 192.168.1.30 |
| ▲ ▲ ▲ |
| └──────── wsrep Galera Mesh (Port 4567/4568) ─────┘ |
| Synchronous Certification & Group Communication |
+--------------------------------------------------------------------------+
Direct Technical Comparison: Asynchronous vs. Galera Synchronous
| Feature / Metric | Standard MySQL Asynchronous | Semi-Synchronous | MariaDB Galera Cluster |
|---|---|---|---|
| Replication Lag | Common (Seconds to Minutes) | Minimal (Waits for 1 ACK) | Strictly 0 (Certified prior to commit) |
| Write Topologies | Single Primary Only | Single Primary Only | True Active-Active Multi-Master |
| Failover Mechanism | Manual / Orchestrator | Complex Scripting | Instantaneous & Automatic |
| Split-Brain Risk | High without MHA | High | Zero (Built-in Quorum Voting) |
| Data Loss on Crash | Yes (Unreplicated binlogs) | Low | Absolute Zero (Zero-RPO) |
2. Firewall & Network Port Configuration
Galera requires specific network ports for group communication, state transfer (SST/IST), and incremental updates. Open these ports on all 3 nodes:
# Ubuntu (UFW)
sudo ufw allow 3306/tcp # MySQL/MariaDB client connections
sudo ufw allow 4567/tcp # Galera Cluster group communication
sudo ufw allow 4567/udp # Galera Cluster multicast
sudo ufw allow 4568/tcp # Incremental State Transfer (IST)
sudo ufw allow 4444/tcp # State Snapshot Transfer (SST via rsync/mariabackup)
sudo ufw reload
3. Configuring galera.cnf Across All Cluster Nodes
On each node, install MariaDB Server and the Galera plugin:
# Ubuntu / Debian
sudo apt-get install -y mariadb-server mariadb-backup
# AlmaLinux / Rocky Linux
sudo dnf install -y mariadb-server mariadb-backup
Create /etc/mysql/mariadb.conf.d/60-galera.cnf (or /etc/my.cnf.d/galera.cnf on AlmaLinux):
# /etc/mysql/mariadb.conf.d/60-galera.cnf
[mysqld]
binlog_format = ROW
default_storage_engine = InnoDB
innodb_autoinc_lock_mode = 2
innodb_doublewrite = 1
bind-address = 0.0.0.0
# WSREP (Write Set Replication) Settings
wsrep_on = ON
wsrep_provider = /usr/lib/galera/libgalera_smm.so
# On AlmaLinux: /usr/lib64/galera-4/libgalera_smm.so
# Cluster Identification
wsrep_cluster_name = "nextgen_galera_cluster"
wsrep_cluster_address = "gcomm://192.168.1.10,192.168.1.20,192.168.1.30"
# Node-Specific Configuration (Change IP & Name per node!)
wsrep_node_address = "192.168.1.10"
wsrep_node_name = "galera-node-khi"
# High-Speed SST Provider using Mariabackup (Non-blocking)
wsrep_sst_method = mariabackup
wsrep_sst_auth = "sstuser:StrongSecretPassword123"
Note: Update wsrep_node_address and wsrep_node_name to correspond to each respective node.
4. Bootstrapping the Primary Node and Joining the Cluster
A common rookie mistake is attempting to start mariadb with systemctl start mariadb on all nodes simultaneously. A brand new cluster must be bootstrapped from Node 1:
Step 1: Bootstrap Node 1 (Karachi)
# Initialize the initial cluster state
sudo galera_new_cluster
Verify that Node 1 is operating as a 1-node cluster:
SHOW STATUS LIKE 'wsrep_cluster_size';
-- Returns: Value = 1
SHOW STATUS LIKE 'wsrep_cluster_status';
-- Returns: Value = Primary
Step 2: Start Node 2 (Lahore) and Node 3 (Islamabad)
On Node 2 and Node 3, simply start the standard service:
sudo systemctl start mariadb
Query cluster status from any node:
SHOW STATUS LIKE 'wsrep_cluster_size';
The output will now display Value = 3. The three nodes are fully synchronized!
5. Live Replication Test: Proving Multi-Master Consistency
To demonstrate true active-active capabilities:
-
On Node 1 (Karachi):
CREATE DATABASE test_galera; USE test_galera; CREATE TABLE customers (id INT AUTO_INCREMENT PRIMARY KEY, name VARCHAR(100), city VARCHAR(50)) ENGINE=InnoDB; INSERT INTO customers (name, city) VALUES ('Ahmed Raza', 'Karachi'); -
On Node 2 (Lahore):
USE test_galera; SELECT * FROM customers; --Ahmed Raza appears instantaneously! INSERT INTO customers (name, city) VALUES ('Bilal Tariq', 'Lahore'); -
On Node 3 (Islamabad):
USE test_galera; SELECT * FROM customers; -- Both Ahmed Raza and Bilal Tariq appear immediately!
Any read or write executed on any node is propagated synchronously across all datacenters with mathematical consistency.
6. Enterprise Resiliency: Quorum & Split-Brain Prevention
A 3-node cluster guarantees that if any single node experiences a power failure or network cutoff, the remaining 2 nodes maintain quorum (2/3 > 50%) and continue serving traffic without hesitation.
Explore our related infrastructure tutorials:
- Ansible Automation for Fleet Management
- Linux eBPF & XDP DDoS Mitigation
- PostgreSQL Production Tuning: shared_buffers & work_mem
For large-scale transactional platforms, financial ledgers, and zero-downtime e-commerce giants requiring dedicated physical compute and multi-datacenter connectivity, explore our high-availability Dedicated Servers in Pakistan.
Deploy Zero-Downtime Database Clusters with Nextgen
Eliminate replication lag, single points of failure, and database downtime. Deploy high-availability MariaDB and PostgreSQL clusters across Pakistan with local low latency.
