For the complete documentation index, see llms.txt.
Skip to main content
Version: 8.10 (unreleased)

Multi-Region RDBMS operational procedure

This runbook covers the day-2 operations of a Multi-Region RDBMS setup. It covers losing a region, bringing it back, and activating a zone you declared but never deployed.

caution

Develop, test, and rehearse these procedures in a non-production environment before you need them. The commands below are examples from the reference implementation. Adapt them to your environment.

What is different from dual-region​

In a dual-region setup, losing a region costs the Zeebe quorum. Processing stops, and the failover procedure exists to restore it. That procedure removes the lost brokers, disables the exporter to the lost region, and later restores secondary storage from a snapshot.

With three or more zones and no zone holding half the replicas or more, none of that applies. Every partition keeps a majority of its replicas. Zeebe keeps processing, and you need no Zeebe action to restore service. The dry run confirms this before you act. The failover procedure mostly reports. It only acts on the database writer, and only when the writer was in the lost region.

Side-by-side timelines of the same zone loss. In a two-zone cluster, Zeebe loses quorum and processing stops until an operator force-removes the lost brokers and disables the exporter. Failback also requires a secondary storage snapshot and restore, for four operator steps in total. In a three-zone cluster, quorum holds and processing continues. Three operator steps remain: promoting the database writer if it was in the lost zone, removing the lost zone, which is recommended but not needed for quorum, and redeploying the zone at failback.What a region loss costs youThe same event, in a two-zone and in a three-zone cluster. Time runs downward.Red needs an operator, amber is impact you absorb, grey is a step that does not exist here.Dual-region, two zonesMulti-Region RDBMS, three zonesZone B is lostZeebe loses quorum, processing STOPSevery partition is down to 1 replica of 2Force-remove the lost brokersprocessing only resumes after thisDisable the exporter to zone BProcessing resumesFailback: redeploy zone BSnapshot and restore the secondary storagezone B has no copy of the exported dataCluster is whole againZone C is lostQuorum holds, processing CONTINUESevery partition still has 3 replicas of 5Promote the database writeronly if the writer was in zone CRemove the lost zonerecommended, not needed for quorumNothing to disablethere is one exporter, and one databaseFailback: redeploy zone CNothing to restorethe database already holds the exported dataCluster is whole again4 operator steps.Processing is down until the first two are done.3 operator steps.Processing never stopped.
StepDual-regionMulti-Region RDBMS
Restore processingForce-remove the lost brokersNothing, processing never stopped
Secondary storage after failoverDisable the exporter to the lost regionNothing, there is one exporter and one database
Promote the databasen/aOnly if the writer was in the lost region
Remove the lost zoneSame step as restoring processingRecommended, not needed for quorum
FailbackSnapshot and restore secondary storageRedeploy the region

The dual-region procedure takes 10 operator steps: two to fail over and eight to fail back. The diagram above counts three operator actions here: promote the writer if needed, remove the lost zone, and redeploy the region at failback. The runbook below adds confirmations around them, for five steps in total.

Use this runbook only for Multi-Region RDBMS

This runbook applies only to a zone-aware cluster with three or more zones and RDBMS secondary storage. Don't run the dual-region procedure on it: force-removing brokers or restoring secondary storage from a snapshot is unnecessary here and can lose data. For a two-region cluster with Elasticsearch, use the dual-region procedure instead.

Terminology​

TermMeaning
SlotA position in the region list, numbered from 0. Fixed when the cluster is bootstrapped.
ZoneThe Camunda-level name of a region, for example london. One zone per region.
Declared zoneA zone present in the zone list, whether or not it is deployed.
Active regionA slot that is actually deployed.
WriterThe single database instance accepting writes from every region.

Prerequisites​

You need a local copy of the aws/kubernetes/eks-multi-region-rdbms reference architecture, from the camunda-deployment-references repository. It holds the Terraform modules, the Helm values, and every procedure script this documentation refers to.

The following clones the repository and changes into the architecture directory. Every command in this documentation runs from there.

aws/kubernetes/eks-multi-region-rdbms/procedure/get-your-copy.sh
loading...

The reference architecture is a starting point you own and extend, not a module you consume, so the workflow is to copy it into your own repository rather than reference it remotely.

Source the environment before running any procedure. The scripts derive everything from the Terraform state, and refuse to run against an inconsistent topology:

cd procedure
. ./export-terraform-outputs.sh
. ./export_environment_prerequisites.sh

Source them with the leading dot (. ./script.sh). These scripts export variables into your current shell, not into a subshell. For what each variable means, see prepare the environment in the deployment guide.

You also need the credentials and the CLI tools the deployment used: kubectl contexts for every active region, helm, jq, and your cloud provider's CLI. The deployment guide lists them.

Confirm the cluster is healthy before you start, so you can tell what the procedure changed:

./check-cluster-topology.sh

Handle a region loss​

1. Confirm the quorum is intact​

The surviving zones keep processing if they hold a majority of each partition's replicas. The concept page explains when this holds.

Confirm this rather than assuming it. The script takes one lost slot and computes the surviving replicas without it. Its verdict only covers a single lost zone. If more than one zone is affected, the reference procedures don't cover the situation. Read the partition health of the surviving brokers from GET /actuator/cluster on a surviving region instead. Don't use ./check-cluster-topology.sh here: it expects every active region to be up.

./failover.sh <lost-region-slot> --dry-run

With --dry-run, the script reports the quorum state, prints the current cluster view, and warns if the surviving zones no longer hold a majority. It changes nothing, so you can read the verdict before deciding to act.

2. Promote the database writer if needed​

If the writer was in the lost region, promote a surviving member. The mode depends on whether the lost region is still reachable:

The region is still reachable, for example during a scheduled evacuation. A switchover completes replication before promoting, so no data is lost in the RDBMS. It also takes considerably less time than an unplanned failover.

Run the same script as in step 1, without --dry-run. It repeats the quorum report, then promotes a surviving member if the writer was in the lost region:

./failover.sh <lost-region-slot>

Camunda needs no reconfiguration and no restart, as long as the JDBC URL keeps resolving to the current writer. The reference implementation gets that from the AWS Advanced JDBC Wrapper. Its failover plugin follows the writer on established connections, and on brokers that start after the promotion. This is not general JDBC behavior. With your own database, whether connections re-resolve the writer depends on your driver and endpoint. Confirm it or plan a restart.

If the writer was not in the lost region, you need no database action.

Move the Raft leaders to the new writer region​

Once the writer moves, the zone priorities still favor the region that hosted the old one. Partition leaders keep exporting across regions and pay the inter-region round trip on every flush. Move the leaders next to the new writer:

  1. Raise the priority of the zone that now hosts the writer. See zone-aware clusters for the priority property, and the Partitioning API for applying it to a running cluster.
  2. Wait until the change reports COMPLETED. The cluster rejects a new change while one is still in progress.
  3. Run a rebalance with POST /cluster/v2/rebalance. Priorities apply at the next election and don't move existing leaders on their own.

3. Route client traffic away from the lost region​

Zeebe keeps processing, but the gateway in the lost region is unreachable. Update your DNS or load balancer to stop sending client traffic there. Traffic routing sits outside Camunda's control and depends on your own setup.

4. Remove the lost zone​

Remove the brokers of the lost zone. One atomic change evicts them. It also drops the zone from the persisted partition distribution, so quorum stops counting replicas that cannot answer:

./failover.sh <lost-region-slot> --drain-brokers

This issues DELETE /actuator/cluster/zones/<zone>?force=true against a surviving region. Without force=true, the API tries a graceful drain, which fails when the zone is down. Only do this for a zone that is down and unreachable, and for one zone at a time. See the cluster management API.

In a planned evacuation, the zone is still reachable, so don't force-remove it. Drain it gracefully instead: send DELETE /actuator/cluster/zones/<zone> without force=true through the Remove a zone API. The engine moves the zone's partitions to the remaining zones before it removes the brokers. The request is asynchronous. Wait until GET /actuator/cluster reports the change as COMPLETED before you shut down the zone's brokers.

You must remove the zone when it held half the replicas or more. The replica count decides this, not the number of zones. See step 1. An evenly split two-zone cluster always needs it, which is why Dual-Region has a failover runbook and this architecture does not.

The trade-off is failback cost. You must add a removed zone back when you bring the region back, and its brokers start empty.

5. Verify the degraded cluster​

./verify-degraded-cluster.sh <lost-region-slot>

The cluster should report the surviving brokers, all partitions healthy, and processing continuing.

Bring a region back​

Failback is short by design. It has no secondary storage snapshot and restore step. The database holds a single copy of the exported data and replicates it itself. A returning region has nothing to catch up on at the Camunda level.

./failback.sh <recovered-region-slot>

The procedure does four things:

  1. Redeploys Camunda in the recovered region: namespace, database secret, Helm values, and chart.
  2. Re-exports the region's services to the ClusterSet, so brokers in other regions can resolve them again.
  3. Re-adds the zone if you force-removed it during failover. If you left the zone in place, its brokers rejoin and catch up from the Raft log with no membership change at all.
  4. Reports the database state, and stops if an unplanned recovery left the global topology incomplete.

Move the writer back to the recovered region if the other regions are further from the current writer:

./failback.sh <recovered-region-slot> --switch-writer

Leaving the writer where it is costs nothing but cross-region latency for the regions furthest from it.

After an unplanned failover

An unplanned recovery can leave the promoted member detached from the global database. Restore a complete Aurora Global Database topology with the AWS recovery procedure before running failback.sh. The script refuses to continue while the global cluster has only one member.

Confirm the topology when done:

./check-cluster-topology.sh

Activate a declared zone​

Activating a zone that you declared in the zone list but never deployed is an online operation.

This section applies only to a zone that you declared at bootstrap and never ran. A zone that you removed during failover comes back through Bring a region back instead.

The operation is online because the zone already exists as far as the cluster is concerned. Every region started with that zone in its zone list. The partition distribution already assigned it replicas, and every partition runs short of those replicas. Deploying the zone starts brokers that claim replicas already reserved for them. Nothing else changes:

  • The cluster does not renumber any broker.
  • The cluster does not redistribute any partition.
  • No running region restarts.
  • You do not need a cluster management API call.

1. Provision the infrastructure​

Raise active_region_count so the region's cluster, Transit Gateway attachments, and security group rules exist. Use the same variable file as the initial deployment:

cd ../terraform/clusters
terraform apply -var-file=terraform-cluster.tfvars -var active_region_count=3

2. Update the environment​

Re-source the environment so CAMUNDA_ACTIVE_REGIONS reflects the new count, and register a kubectl context for the new cluster:

cd ../../procedure
unset CAMUNDA_ACTIVE_REGIONS
. ./export-terraform-outputs.sh
. ./export_environment_prerequisites.sh
./register-kubecontexts.sh

3. Activate the slot​

./activate-region.sh <slot>

The procedure does the following:

  1. Joins the new cluster to the ClusterSet.
  2. Prepares its storage class, namespace, and database secret.
  3. Renders the Helm values with the longer contact point list.
  4. Installs only the new region.
  5. Exports its services.
  6. Waits for the new brokers to join.

The regions already running keep their shorter contact point list, and they don't restart. The contact point list matters at bootstrap. Once a cluster forms, a newcomer only has to reach one member, and the rest learn about it by gossip. The running regions pick up the longer list on their next upgrade.

warning

activate-region.sh fills a slot that already exists in the zone list. It does not add a new zone. Adding a zone that was never declared changes the zone list in every region and redistributes partitions. That is a migration rather than an online operation.

The script rejects any slot outside the provisioned range, 0 to CAMUNDA_REGION_SLOTS - 1. Before you run it, apply the Terraform step above. Then re-source the environment and register the kubectl context, so CAMUNDA_ACTIVE_REGIONS and CLUSTER_CONTEXTS include the new slot.

Upgrade the cluster​

Upgrade one region at a time, and wait for the cluster to report healthy before starting the next:

./check-cluster-topology.sh

Upgrading several regions at the same time risks losing quorum.

Follow the general upgrade guidance and create a backup first.

Diagnose problems​

SymptomStart here
Brokers do not reach the expected count./submariner/verify-submariner.sh, then ./submariner/diagnose-submariner.sh
Cross-region traffic is dropped./verify-cross-region-connectivity.sh
Export latency is higher than expected./measure-rdbms-latency.sh
Partition distribution looks wrong./check-cluster-topology.sh

For the underlying causes and the AWS commands that confirm them, see troubleshooting in the EKS guide.