Multi-Region RDBMS operational procedure
This runbook covers the day-2 operations of a Multi-Region RDBMS setup. It covers losing a region, bringing it back, and activating a zone you declared but never deployed.
Develop, test, and rehearse these procedures in a non-production environment before you need them. The commands below are examples from the reference implementation. Adapt them to your environment.
What is different from dual-region
In a dual-region setup, losing a region costs the Zeebe quorum. Processing stops, and the failover procedure exists to restore it. That procedure removes the lost brokers, disables the exporter to the lost region, and later restores secondary storage from a snapshot.
With three or more zones and no zone holding half the replicas or more, none of that applies. Every partition keeps a majority of its replicas. Zeebe keeps processing, and you need no Zeebe action to restore service. The dry run confirms this before you act. The failover procedure mostly reports. It only acts on the database writer, and only when the writer was in the lost region.
| Step | Dual-region | Multi-Region RDBMS |
|---|---|---|
| Restore processing | Force-remove the lost brokers | Nothing, processing never stopped |
| Secondary storage after failover | Disable the exporter to the lost region | Nothing, there is one exporter and one database |
| Promote the database | n/a | Only if the writer was in the lost region |
| Remove the lost zone | Same step as restoring processing | Recommended, not needed for quorum |
| Failback | Snapshot and restore secondary storage | Redeploy the region |
The dual-region procedure takes 10 operator steps: two to fail over and eight to fail back. The diagram above counts three operator actions here: promote the writer if needed, remove the lost zone, and redeploy the region at failback. The runbook below adds confirmations around them, for five steps in total.
This runbook applies only to a zone-aware cluster with three or more zones and RDBMS secondary storage. Don't run the dual-region procedure on it: force-removing brokers or restoring secondary storage from a snapshot is unnecessary here and can lose data. For a two-region cluster with Elasticsearch, use the dual-region procedure instead.
Terminology
| Term | Meaning |
|---|---|
| Slot | A position in the region list, numbered from 0. Fixed when the cluster is bootstrapped. |
| Zone | The Camunda-level name of a region, for example london. One zone per region. |
| Declared zone | A zone present in the zone list, whether or not it is deployed. |
| Active region | A slot that is actually deployed. |
| Writer | The single database instance accepting writes from every region. |
Prerequisites
You need a local copy of the aws/kubernetes/eks-multi-region-rdbms reference architecture, from the camunda-deployment-references repository. It holds the Terraform modules, the Helm values, and every procedure script this documentation refers to.
The following clones the repository and changes into the architecture directory. Every command in this documentation runs from there.
loading...
The reference architecture is a starting point you own and extend, not a module you consume, so the workflow is to copy it into your own repository rather than reference it remotely.
Source the environment before running any procedure. The scripts derive everything from the Terraform state, and refuse to run against an inconsistent topology:
cd procedure
. ./export-terraform-outputs.sh
. ./export_environment_prerequisites.sh
Source them with the leading dot (. ./script.sh). These scripts export variables into your current shell, not into a subshell. For what each variable means, see prepare the environment in the deployment guide.
You also need the credentials and the CLI tools the deployment used: kubectl contexts for every active region, helm, jq, and your cloud provider's CLI. The deployment guide lists them.
Confirm the cluster is healthy before you start, so you can tell what the procedure changed:
./check-cluster-topology.sh
Handle a region loss
1. Confirm the quorum is intact
The surviving zones keep processing if they hold a majority of each partition's replicas. The concept page explains when this holds.
Confirm this rather than assuming it. The script takes one lost slot and computes the surviving replicas without it. Its verdict only covers a single lost zone. If more than one zone is affected, the reference procedures don't cover the situation. Read the partition health of the surviving brokers from GET /actuator/cluster on a surviving region instead. Don't use ./check-cluster-topology.sh here: it expects every active region to be up.
./failover.sh <lost-region-slot> --dry-run
With --dry-run, the script reports the quorum state, prints the current cluster view, and warns if the surviving zones no longer hold a majority. It changes nothing, so you can read the verdict before deciding to act.
2. Promote the database writer if needed
If the writer was in the lost region, promote a surviving member. The mode depends on whether the lost region is still reachable:
- Planned
- Unplanned
The region is still reachable, for example during a scheduled evacuation. A switchover completes replication before promoting, so no data is lost in the RDBMS. It also takes considerably less time than an unplanned failover.
Run the same script as in step 1, without --dry-run. It repeats the quorum report, then promotes a surviving member if the writer was in the lost region:
./failover.sh <lost-region-slot>
The region is gone. Follow the Aurora Global Database unplanned recovery procedure. The reference script doesn't automate this operation because detaching and promoting a member changes the global topology outside Terraform.
Whatever had not replicated at the time of the outage can be missing from the promoted database. With LOG_SEQ, recovery depends on promoting a standby whose replication was confirmed within the configured minimum. With DELAY, recovery depends on the actual lag staying below the configured delay. Restore the global database membership before running the Camunda failback procedure.
Camunda needs no reconfiguration and no restart, as long as the JDBC URL keeps resolving to the current writer. The reference implementation gets that from the AWS Advanced JDBC Wrapper. Its failover plugin follows the writer on established connections, and on brokers that start after the promotion. This is not general JDBC behavior. With your own database, whether connections re-resolve the writer depends on your driver and endpoint. Confirm it or plan a restart.
If the writer was not in the lost region, you need no database action.
Move the Raft leaders to the new writer region
Once the writer moves, the zone priorities still favor the region that hosted the old one. Partition leaders keep exporting across regions and pay the inter-region round trip on every flush. Move the leaders next to the new writer:
- Raise the priority of the zone that now hosts the writer. See zone-aware clusters for the priority property, and the Partitioning API for applying it to a running cluster.
- Wait until the change reports
COMPLETED. The cluster rejects a new change while one is still in progress. - Run a rebalance with
POST /cluster/v2/rebalance. Priorities apply at the next election and don't move existing leaders on their own.
3. Route client traffic away from the lost region
Zeebe keeps processing, but the gateway in the lost region is unreachable. Update your DNS or load balancer to stop sending client traffic there. Traffic routing sits outside Camunda's control and depends on your own setup.
4. Remove the lost zone
Remove the brokers of the lost zone. One atomic change evicts them. It also drops the zone from the persisted partition distribution, so quorum stops counting replicas that cannot answer:
./failover.sh <lost-region-slot> --drain-brokers
This issues DELETE /actuator/cluster/zones/<zone>?force=true against a surviving region. Without force=true, the API tries a graceful drain, which fails when the zone is down. Only do this for a zone that is down and unreachable, and for one zone at a time. See the cluster management API.
In a planned evacuation, the zone is still reachable, so don't force-remove it. Drain it gracefully instead: send DELETE /actuator/cluster/zones/<zone> without force=true through the Remove a zone API. The engine moves the zone's partitions to the remaining zones before it removes the brokers. The request is asynchronous. Wait until GET /actuator/cluster reports the change as COMPLETED before you shut down the zone's brokers.
You must remove the zone when it held half the replicas or more. The replica count decides this, not the number of zones. See step 1. An evenly split two-zone cluster always needs it, which is why Dual-Region has a failover runbook and this architecture does not.
The trade-off is failback cost. You must add a removed zone back when you bring the region back, and its brokers start empty.
5. Verify the degraded cluster
./verify-degraded-cluster.sh <lost-region-slot>
The cluster should report the surviving brokers, all partitions healthy, and processing continuing.
Bring a region back
Failback is short by design. It has no secondary storage snapshot and restore step. The database holds a single copy of the exported data and replicates it itself. A returning region has nothing to catch up on at the Camunda level.
./failback.sh <recovered-region-slot>
The procedure does four things:
- Redeploys Camunda in the recovered region: namespace, database secret, Helm values, and chart.
- Re-exports the region's services to the ClusterSet, so brokers in other regions can resolve them again.
- Re-adds the zone if you force-removed it during failover. If you left the zone in place, its brokers rejoin and catch up from the Raft log with no membership change at all.
- Reports the database state, and stops if an unplanned recovery left the global topology incomplete.
Move the writer back to the recovered region if the other regions are further from the current writer:
./failback.sh <recovered-region-slot> --switch-writer
Leaving the writer where it is costs nothing but cross-region latency for the regions furthest from it.
An unplanned recovery can leave the promoted member detached from the global database. Restore a complete Aurora Global Database topology with the AWS recovery procedure before running failback.sh. The script refuses to continue while the global cluster has only one member.
Confirm the topology when done:
./check-cluster-topology.sh
Activate a declared zone
Activating a zone that you declared in the zone list but never deployed is an online operation.
This section applies only to a zone that you declared at bootstrap and never ran. A zone that you removed during failover comes back through Bring a region back instead.
The operation is online because the zone already exists as far as the cluster is concerned. Every region started with that zone in its zone list. The partition distribution already assigned it replicas, and every partition runs short of those replicas. Deploying the zone starts brokers that claim replicas already reserved for them. Nothing else changes:
- The cluster does not renumber any broker.
- The cluster does not redistribute any partition.
- No running region restarts.
- You do not need a cluster management API call.
1. Provision the infrastructure
Raise active_region_count so the region's cluster, Transit Gateway attachments, and security group rules exist. Use the same variable file as the initial deployment:
cd ../terraform/clusters
terraform apply -var-file=terraform-cluster.tfvars -var active_region_count=3
2. Update the environment
Re-source the environment so CAMUNDA_ACTIVE_REGIONS reflects the new count, and register a kubectl context for the new cluster:
cd ../../procedure
unset CAMUNDA_ACTIVE_REGIONS
. ./export-terraform-outputs.sh
. ./export_environment_prerequisites.sh
./register-kubecontexts.sh
3. Activate the slot
./activate-region.sh <slot>
The procedure does the following:
- Joins the new cluster to the ClusterSet.
- Prepares its storage class, namespace, and database secret.
- Renders the Helm values with the longer contact point list.
- Installs only the new region.
- Exports its services.
- Waits for the new brokers to join.
The regions already running keep their shorter contact point list, and they don't restart. The contact point list matters at bootstrap. Once a cluster forms, a newcomer only has to reach one member, and the rest learn about it by gossip. The running regions pick up the longer list on their next upgrade.
activate-region.sh fills a slot that already exists in the zone list. It does not add a new zone. Adding a zone that was never declared changes the zone list in every region and redistributes partitions. That is a migration rather than an online operation.
The script rejects any slot outside the provisioned range, 0 to CAMUNDA_REGION_SLOTS - 1. Before you run it, apply the Terraform step above. Then re-source the environment and register the kubectl context, so CAMUNDA_ACTIVE_REGIONS and CLUSTER_CONTEXTS include the new slot.
Upgrade the cluster
Upgrade one region at a time, and wait for the cluster to report healthy before starting the next:
./check-cluster-topology.sh
Upgrading several regions at the same time risks losing quorum.
Follow the general upgrade guidance and create a backup first.
Diagnose problems
| Symptom | Start here |
|---|---|
| Brokers do not reach the expected count | ./submariner/verify-submariner.sh, then ./submariner/diagnose-submariner.sh |
| Cross-region traffic is dropped | ./verify-cross-region-connectivity.sh |
| Export latency is higher than expected | ./measure-rdbms-latency.sh |
| Partition distribution looks wrong | ./check-cluster-topology.sh |
For the underlying causes and the AWS commands that confirm them, see troubleshooting in the EKS guide.
Related resources
- Multi-Region RDBMS: the architecture and its trade-offs.
- Multi-region setup with RDBMS on Amazon EKS: the reference implementation.
- Cluster management API: the endpoints these procedures call.
- Dual-region operational procedure: the equivalent runbook for two regions.