For the complete documentation index, see llms.txt.
Skip to main content
Version: 8.10

Multi-Region RDBMS

Multi-Region RDBMS spreads an Orchestration Cluster across two or more regions and delegates secondary storage replication to the database. With three or more regions, a region loss leaves the Raft quorum intact.

Multi-Region RDBMS spreads a single Orchestration Cluster across two or more regions. It uses a relational database as its secondary storage, and leaves replication to that database. You need three or more regions to keep processing through a region loss. Every partition then keeps a majority when one region disappears, as long as no region holds half the replicas of a partition or more. The engine keeps processing instead of stopping for an operator.

Before you begin

Running a multi-region setup requires you to develop, test, and execute operational procedures specific to your environment. Review the limitations and requirements before you commit to this configuration.

To have your multi-region setup covered by Camunda enterprise support, get your configuration and runbooks reviewed by Camunda before going to production. Contact your Customer Success Manager as soon as you plan your multi-region setup.

That review covers the architecture you build. It does not make the reference implementation a supported product. That repository is explicitly experimental. It is meant for learning, evaluation, and design review, and is not production-ready as published.

How Multi-Region RDBMS differs from Dual-Region​

The two multi-region architectures differ in region count and in secondary storage.

Dual-Region is the two-region architecture with Elasticsearch secondary storage and parity-numbered brokers. It has two properties that come from the region count rather than from any implementation choice.

With two regions, no replica placement survives losing half of them. A region loss therefore costs the Raft quorum, and Zeebe stops processing until an operator force-removes the lost brokers. Each region also owns its own copy of the secondary storage, so a returning region has to be re-seeded. Failback therefore includes a secondary storage snapshot and a cross-region restore.

Multi-Region RDBMS removes both. It changes the number of regions, and it changes who owns replication of the secondary storage.

ConsiderationDual-RegionMulti-Region RDBMS
RegionsExactly twoTwo or more. Three or more to keep processing through a region loss
Region lossQuorum lost, processing stops until brokers are force-removedWith three or more regions, quorum preserved and processing continues
FailbackMulti-step runbook including a secondary storage snapshot and restoreRedeploy the region, nothing to restore
Secondary storageElasticsearch, one cluster per region, one Camunda exporter per regionRDBMS, one database, one exporter, replication inside the database
OptimizeSupportedNot available, Optimize requires Elasticsearch or OpenSearch
Relative cost$$$: two regions of Orchestration Cluster capacity, plus cross-region traffic$$$$: two or more regions of Orchestration Cluster capacity, three or more to survive a region loss, plus cross-region traffic

Choose Multi-Region RDBMS when processing must continue through a region loss without operator intervention, and when you can run without Optimize. Choose Dual-Region when two regions are sufficient, or when you need Optimize on the same cluster.

Architecture​

Three infrastructure layers let one cluster span several regions.

Three regions each running an Orchestration Cluster, connected by a private inter-region network and a cross-cluster service discovery layer, all writing to a single replicated relational databaseOne Orchestration Cluster, three regions, one databaseEvery zone processes. Every zone writes to the same database.ZONES ( one per region )zone londoneu-west-2 priority 1000broker london_0broker london_1gateway + connectors2 replicas of every partitionzone pariseu-west-3 priority 900broker paris_0broker paris_1gateway + connectors2 replicas of every partitionzone zuricheu-central-2 priority 800broker zurich_0broker zurich_1gateway + connectors1 replica of every partitionPrivate inter-region network · Cross-cluster service discoveryRaft, full meshone JDBC URL,shared by every zoneSECONDARY STORAGE ( one, shared )Replicated relational databaseone writer -> standbysreplication is the database's job,not Camunda's

One Orchestration Cluster spans every region. Each region runs its own brokers and connectors, and all of them are members of the same Zeebe cluster. Three infrastructure layers make that possible:

LayerResponsibility
Inter-region networkCarries broker-to-broker traffic, including Raft, between regions on private addresses.
Cross-cluster service discoveryPublishes each region's Zeebe service under a name every other region can resolve.
Relational secondary storageAccepts writes from every region through a single endpoint, and replicates them itself.

The Camunda layer sees one cluster and one database. Everything region-specific lives in the infrastructure layers, which keeps the architecture portable across deployment platforms.

Partition placement across zones​

Multi-Region RDBMS relies on zone-aware clusters. Each region is one zone, and every zone declares how many brokers it holds and how many replicas of each partition live in it.

A partition survives while a majority of its replicas answer. Zeebe counts replicas, not zones. A zone holds as many replicas as you give it. The replication factor is the sum across the zones rather than the number of zones.

The simplest case is one replica per zone, which is what the following illustration uses. The replication factor then equals the zone count, and a partition keeps its majority while N - 1 > N / 2.

With two zones, losing one leaves one replica of two and no majority, so processing stops. With three zones, losing one leaves two replicas of three, a majority, so processing continues.Why three zones, and not twoOne partition, one replica per zone.A partition works while a majority of its replicas answer.2 zones, one lostzone Areplicaof P1zone BreplicaLOST1 of 2 replicas answerno majorityprocessing stops3 zones, one lostzone Areplicaof P1zone Breplicaof P1zone CreplicaLOST2 of 3 replicas answermajority holdsthe engine keeps processingThe same holds for every partition, because every zone carries one replica of all of them.A fourth zone does not raise the tolerance either: four replicas still need three to answer.
ZonesReplicas per zoneReplication factorZone losses tolerated
2120
3131
4141
5152

Three zones is the smallest topology in which losing one does not stop the engine. A fourth zone at one replica each does not change that. The replication factor becomes four, and a majority is still three. A second loss leaves two. Tolerating two losses takes five replicas.

Choosing an asymmetric layout​

Zones do not have to be equal, and making them equal is rarely what you want. A common shape places more replicas in the regions that also host a database member. It places one replica in a region that exists to break ties:

ZoneDatabase memberReplicasLosing this zone leaves
Ayes23 of 5, majority holds
Byes23 of 5, majority holds
Cno14 of 5, majority holds

That is replicationFactor: 5 across three zones. The third region carries a vote without carrying a database, and it is the zone that decides a quorum when the other two disagree.

It is still a full region: the brokers there hold data and process work like any others, and the region serves clients. Only the database member is absent. This is not a lightweight arbiter or a "2.5 region" topology.

The only rule is that no single zone may hold half the replicas or more, or losing that zone stops the engine. A 4-1-1 layout across three zones fails it: losing the first leaves two replicas of six.

Zone awareness also assigns a Raft election priority per zone. Give the zone that hosts the database writer the highest priority. Elections then favor leaders next to the writer, which reduces inter-region round trips on export flushes. The priority biases elections but does not pin leaders. Move existing leaders with the coordinated rebalancing API (POST /cluster/v2/rebalance), described in rebalancing.

Replication-agnostic secondary storage​

Camunda uses one JDBC URL per Orchestration Cluster, through a connection pool on each broker, and the RDBMS exporter has no multi-region mode. As the RDBMS multi-region support documentation states, multi-region replication must be handled within the database itself.

Multi-Region RDBMS adopts that constraint rather than working around it:

  • Every broker in every region writes to the same JDBC URL.
  • There are no per-region exporters to enable, disable, or reinitialize.
  • Region loss desynchronizes nothing at the Camunda layer, so failback has no restore step.
  • Swapping the database changes one value.
Brokers in london, paris, and zurich all write to one JDBC URL, which is the same in every region. The URL points at the database writer in london today, and at the paris standby after a promotion. The writer replicates asynchronously to the standby. Any mechanism that keeps one endpoint on the current writer fits: a globally replicated managed database, a PostgreSQL cluster behind a floating endpoint, a connection proxy, or a DNS record you repoint at failover.Every broker in every region writes to one JDBC URLThe database follows its own writer. Camunda sees one endpoint and never changes it.london2 brokersparis2 brokerszurich2 brokersone JDBC URLsame in every regionno per-region exporternothing to restoretodayafter promotionwriterregion londonstandbyregion parisasyncreplicationAny mechanism that keeps one endpoint on the current writer fits:globally replicatedmanaged databasePostgreSQL cluster behinda floating endpointconnectionproxyDNS record yourepoint at failover

Any database that presents a single endpoint following its own writer fits. See multi-region support for the supported databases. For example:

  • A globally replicated managed database.
  • A PostgreSQL cluster behind a floating endpoint.
  • A connection proxy.
  • A DNS record you repoint during failover.

Whichever mechanism you choose, test it with the failover procedure before you go to production.

The Camunda configuration does not change between them.

Monitor asynchronous replication​

Asynchronous replication monitoring is required, not a tuning option. Without it the RDBMS exporter acknowledges records the standby has not received yet, and a writer failover loses exported data. This architecture treats a writer failover as a routine operation rather than an incident, so set camunda.data.secondary-storage.rdbms.async-replication.enabled to true.

Camunda turns this monitoring off by default. The default suits a single database without asynchronous replicas, and the monitoring needs extra database privileges, for example the PG_MONITOR role on PostgreSQL. This architecture needs it, so the reference implementation sets it to true. Once you turn it on, the strategy defaults to LOG_SEQ.

The strategy you can use depends on the database engine, not on the cloud provider. Use LOG_SEQ when your database is in its vendor support list. Otherwise, choose TIME_LAG or DELAY. The reference implementation covers only Aurora Global Database. Managed databases on other providers, such as Azure or Google Cloud, follow the same rules but have no reference implementation.

StrategyWhen to use itWhat you configure
LOG_SEQ (LSN monitoring)Default and preferred. Reads the database's own replication position. Supported on Aurora Global Database with PostgreSQL, Aurora Global Database with MySQL, MSSQL, and PostgreSQL.async-replication.type: LOG_SEQ
TIME_LAGLess precise than LOG_SEQ. Reads the replication lag the primary reports, and acknowledges less often. Supported on the same databases as LOG_SEQ.async-replication.type: TIME_LAG
DELAYFallback for a database that supports neither LOG_SEQ nor TIME_LAG, for example Azure SQL Database. Carries no replication signal.async-replication.type: DELAY, a delay value, and your own monitoring of the actual lag

Camunda doesn't switch strategies for you. See multi-region support for the supported backends and the settings.

The database tier is active-standby​

The Zeebe data plane is active-active: every region processes. The database tier is not. A single writer serves every region, and brokers that are not co-located with it pay the inter-region round trip on every export flush.

Two consequences follow, and both are sizing decisions rather than configuration:

  • Keep regions inside the round-trip time budget described in network requirements.
  • Size the exporter queue for the latency of the furthest region, not the nearest.

Skewing partition leadership to the writer's zone through zone priority reduces how often that round trip is paid, but it does not remove it.

Three regions, london, paris, and zurich, each run Zeebe brokers. Only london hosts the database writer. The london brokers export to it locally, while the paris and zurich brokers export across a region. The writer replicates asynchronously to the database standby in paris. Zurich has no database member.londonZeebe brokersparisZeebe brokerszurichZeebe brokersDatabase writerDatabase standbyno databaselocalcross-regionasync replication

Requirements​

The architecture only works under the cluster, network, platform, and upgrade requirements below.

Zeebe cluster configuration​

SettingRequirement
Partitioning schemeZONE_AWARE. The parity-based broker numbering only supports exactly two regions.
ZonesOne zone per region, two or more. Three or more to keep processing through a region loss.
number-of-replicas per zoneDeclared per zone, and free to differ between them. To survive a zone loss, no zone may hold half the replication factor or more.
number-of-brokers per zoneDeclared per zone. Keep zones balanced so a zone loss removes an equal share of capacity.
Surviving capacitySize the cluster so the regions left after a loss carry the full workload. Quorum surviving is not the same as the cluster keeping up.
priority per zoneHighest for the zone hosting the database writer, to keep partition leaders next to it.
partitionCountUnrestricted. Size it from your workload. See sizing your environment.

Each broker sets its own zone, while the zone list is identical in every region. For the full property reference, see zone-aware clusters. For the matching Helm keys, see multi-region zone awareness.

Network requirements​

  • Kubernetes clusters, services, and pods must use distinct, non-overlapping CIDRs across every region.
  • Every region must reach every other region. Zeebe uses a full mesh, not a hub and spoke, so partial connectivity leaves partitions unable to form a quorum.
  • Every other cluster must resolve and reach the Kubernetes services in one cluster. The name must be the same from each region's point of view.
  • Round-trip time between regions directly affects Raft commit latency and throughput. As a guideline, keep it at or below 100 ms. Higher latencies degrade performance, but are not a hard limit enforced by the engine.
  • Required open ports between regions:
    • 26500: Zeebe gateway, client and worker communication
    • 26501 and 26502: Zeebe broker and gateway communication, including Raft
    • 8080: Orchestration Cluster REST API
    • 53: DNS, for cross-cluster service resolution

These are the default ports. Change them if you customize them in your configuration.

A private inter-region network is preferred for the database, but it is not required. A public path also works, and it adds egress cost and exposure. You must measure the latency between the regions, and the inter-region connectivity must stay stable.

Infrastructure and deployment platform considerations​

Multi-region setups require careful planning. You must manage the following areas independently, and Camunda does not control or document them:

  • Kubernetes cluster management: managing three or more Kubernetes clusters and their deployments
  • Monitoring and alerting: multi-region monitoring with cross-region correlation
  • Cost implications: three or more clusters and cross-region traffic increase costs, and inter-region data transfer is billed per gigabyte
  • Network reliability: increased latency affects Raft commit latency and export throughput. Even short latency bursts have an impact.
  • Traffic management: DNS and incoming traffic routing across more than two regions
  • Database operations: replication, failover, and backup of the secondary storage are the database's responsibility, and therefore yours
  • Security: consistent security policies and network controls across every region

Upgrade considerations​

Upgrade one region at a time, so the other regions keep the quorum. The operational procedure lists the steps.

Growth and region loss​

Two companion pages describe how the cluster changes over its life:

Limitations​

Recovery behaves as described only inside the boundaries this architecture sets. The following table lists them.

AspectDetails
Installation methodsKubernetes with the Camunda Helm chart. Alternative installation methods are not covered by these guides.
Secondary storageRDBMS only. Elasticsearch and OpenSearch replicate per region and do not fit the single-endpoint model this architecture depends on.
Database availabilityThe database tier is active-standby. A single writer serves every region, and regions further from it pay more export latency.
Management Identity supportNot deployed by the reference configuration, which uses Basic authentication. The Orchestration Cluster-level Admin supports multi-tenancy and role-based access control instead. To use Camunda Hub, deploy Management Identity in a single region as part of a management platform.
Optimize supportNot available. Optimize requires Elasticsearch or OpenSearch, regardless of the region count.
Camunda HubNot deployed by the reference configuration. Camunda Hub runs in a single region, with no multi-region guarantees, and requires OIDC authentication, Management Identity, and PostgreSQL. See Management platform and Orchestration Cluster.
Connectors deploymentConnectors run in every region and are not deduplicated. Account for idempotency to avoid event duplication.
Zone list changesAdding a zone is online, through the cluster management API: the engine places the new zone's replicas without renumbering brokers. Adding a zone leaves the partition count unchanged.
Backup and restoreRDBMS backup relies on continuous primary storage backups plus a database-native backup. See backup and restore.

Reference implementation​

Camunda publishes one implementation of this architecture, on Amazon Web Services:

The architecture is not AWS-specific. Each of its three layers has an equivalent on other platforms. For example, Red Hat OpenShift provides Submariner through Advanced Cluster Management, as the OpenShift dual-region setup already uses.

These pages cover the concepts and settings this page refers to.