For the complete documentation index, see llms.txt.
Skip to main content
Version: 8.9

Dual-Region

Dual-Region is Camunda's production-proven dual-region configuration with continuous replication.

Dual-Region is Camunda's certified, continuous-replication multi-region configuration. A Camunda Orchestration Cluster runs in both a primary and a secondary region at all times, with Zeebe replicating its log stream across both regions using the Raft protocol. Secondary storage (Elasticsearch) is populated independently in each region via the Camunda Exporter. It includes published RTO/RPO targets, a reference architecture, and a documented failover runbook.

Before you begin

Before implementing a dual-region setup, review the limitations and infrastructure considerations for this configuration.

ConsiderationValue
Recovery time (RTO)~15 minutes (see Recovery objectives)
Data loss (RPO)0
Failover modeManual, operator-initiated
Standing second region requiredYes. Full Orchestration Cluster running in both regions

Architecture​

The dual-region setup uses two Kubernetes clusters, each running a complete set of Camunda 8 components.

Camunda dual-region architecture
  • With v2 APIs (default in 8.9+), both regions serve user traffic simultaneously.
  • With v1 APIs, Region 0 is the primary region serving user traffic; Region 1 is operational but doesn't serve user traffic under normal conditions.
note

The diagram shows both regions as operational. Any grayed-out appearance represents user traffic routing, not system operational status. All components in both regions must be running.

ComponentModeBoth regions runningUser trafficRPO

Orchestration Cluster

✅ Required
ZeebeActive-active✅ RequiredBoth regions process data0
AdminActive-active✅ RequiredCluster-level identity0
OperateActive-active with v2 API (active-passive with v1)✅ RequiredBoth regions serve users with v2 API0
TasklistActive-active with Tasklist V2 (active-passive with v1)✅ RequiredBoth regions serve users with Tasklist V20

Elasticsearch

Active-active✅ RequiredData replicated to both0

In a dual-region setup, each Orchestration Cluster component operates as follows:

  • Orchestration cluster runs across both regions using the Raft protocol, distributing partition leaders and followers for continuous replication.
  • Camunda exporters push identical data to the Elasticsearch instance in each region. Camunda orchestration cluster dual export mechanism (not Elasticsearch replication) maintains data consistency. The two Elasticsearch clusters don't communicate directly with each other.
  • Operate and Tasklist maintain synchronized data state across both regions via the Camunda Exporter. See Active-active and active-passive modes.
  • Admin is embedded in the Orchestration Cluster and provides cluster-level identity management.

Component requirements​

ComponentModeRequirementFunctionData loss risk

Orchestration Cluster

ZeebeActive-activeAll brokers in both regions must run
  • Leaders and followers distributed across regions
  • Continuous replication via Raft protocol
  • Both regions required for quorum maintenance
Can handle region failure without data loss when properly configured
AdminActive-activeEmbedded in the Orchestration Cluster
  • Provides unified, cluster-level identity management and authorization
Can handle region failure without data loss
OperateActive-active with v2 API (active-passive with v1)Embedded in the Orchestration Cluster
  • Both regions maintain synchronized data state
  • Both regions serve users when using v2 API
  • Region-specific data: Uncompleted batch operations only when using v1 API
Data loss possible only when using v1 API, as changes are isolated to the initiated region
TasklistActive-active with Tasklist V2 (active-passive with v1)Embedded in the Orchestration Cluster
  • Both regions maintain synchronized data state
  • Both regions serve users when using Tasklist V2
  • Region-specific data: Task assignments only when using v1 API
Data loss possible only when using v1 API, as changes are isolated to the initiated region

Elasticsearch

Active-activeBoth clusters must run
  • Independent clusters in each region
  • Zeebe exports identical data to both continuously and directly
  • Zeebe's dual export mechanism (not Elasticsearch replication) maintains data consistency
  • The clusters don't communicate with each other
Zeebe exporters may fail globally if the secondary Elasticsearch cluster is down

Management platform and Orchestration Cluster​

A dual-region deployment stretches the Orchestration Cluster across both regions and runs the management platform in a single region. The management platform groups the design and management components that sit outside the Orchestration Cluster. Whether a component belongs to the Orchestration Cluster or to the management platform determines what happens to it when a region is lost, and how you protect it.

LayerComponentsStretched across regionsOn region lossProtection
RuntimeOrchestration Cluster (Zeebe, Operate, Tasklist, Admin) and its secondary storageYesProcessing stops until failover completes, then resumesDual-region failover procedure
Management platformManagement Identity, Web Modeler, Console, and OptimizeNoUnavailable until you restore itBackup and restore

Deploy the management platform in a single region​

Deploy the management platform as a separate release in one region, next to the two Orchestration Clusters rather than inside them. That region can be one of the two dual-region regions or a third region. Camunda doesn't stretch these components across regions, and they take no part in the dual-region failover procedure.

The dual-region reference architecture doesn't deploy the management platform. It uses Basic authentication, disables Management Identity, and sets optimize.enabled: false. That's the scope of the reference configuration, not a product restriction: you can run Optimize and Web Modeler alongside a dual-region Orchestration Cluster.

Both components authenticate through Management Identity, so a management platform requires OpenID Connect (OIDC) authentication rather than the Basic authentication the reference configuration uses. Point each component at your Management Identity instance with global.identity.service.url, and give Web Modeler its own PostgreSQL database.

Optimize imports the zeebe-record indices that the legacy Elasticsearch exporter writes. The dual-region reference configuration doesn't write them: it sets orchestration.exporters.zeebe.enabled: false, and the Helm chart doesn't enable this exporter automatically when the Orchestration Cluster spans two regions. Each broker exports only the partitions it leads, so configure the Elasticsearch exporter on every broker in both regions, with the URL of the Elasticsearch cluster that Optimize reads. Add it as an additional exporter through environment variables, the same way the reference configuration adds its regional Camunda Exporters.

note

The Helm chart doesn't reject optimize.enabled: true in a release that has no Management Identity. That combination installs successfully and then fails to authenticate at runtime. Confirm Management Identity is reachable before you enable Optimize.

Region loss behavior for management platform components​

Management platform components don't replicate across regions, so losing the region they run in makes them unavailable until you restore them. If the management platform runs in one of the two dual-region regions, losing that region also stops the Orchestration Cluster until the dual-region failover procedure completes. Failover restores process execution, and deployed processes keep running. It doesn't restore the management platform, which you recover from backups.

Broker processing and authenticated client access recover separately. If the Orchestration Cluster uses OIDC authentication with a provider, such as Keycloak, that runs in the lost region, failover restores broker processing only. Workers, Connectors, and other clients can't get new tokens after their current tokens expire, and users can't sign in to the Orchestration Cluster. Process instances that wait for job workers then stop progressing. Run the OIDC provider where it survives the loss of either dual-region region, or recover it as part of the failover.

ComponentState it holdsIf its region is lost
Management IdentityUsers, groups, roles, tenants, and OIDC clientsAuthentication to Optimize, Web Modeler, and Console fails until you restore it
OIDC provider, such as KeycloakUsers, credentials, and client registrationsClients of the Orchestration Cluster can't get new tokens, and users can't sign in, until you restore it
Web ModelerDiagrams, projects, and collaboration history in PostgreSQLModeling and deployment from Web Modeler stop until you restore it
ConsoleNo persistent state of its ownThe Console UI is unavailable until you redeploy it
OptimizeReports, dashboards, collections, alerts, and its own import positionReporting stops until you restore it, and content created since the last backup is lost

Protect the management platform with backup and restore​

Back up Management Identity and Web Modeler on their own schedule. If Optimize shares the Elasticsearch instance of the Orchestration Cluster, back it up together with the Orchestration Cluster, using the same backup ID. Replicate all backups to a second region, so they survive the loss of the management platform region.

ComponentBackup method
Management IdentityBack up its PostgreSQL database using your database's native tooling. Also back up your OIDC provider, such as Keycloak, and keep user IDs unchanged when you restore it
Web ModelerBack up its PostgreSQL database. See Web Modeler backup and restore
ConsoleNo backup required. Console is stateless, so redeploy it alongside the rest of the management platform
OptimizeUse the Optimize backup API with the same backup ID as the Orchestration Cluster backup. See Optimize backup and restore

The management platform's recovery point and recovery time follow from your backup interval and restore procedure. The dual-region recovery objectives don't cover them, because those objectives apply to the Orchestration Cluster only.

Active-active and active-passive modes​

Starting in Camunda 8.9, active-active is the default user traffic routing for dual-region deployments:

  • Active-active: Both regions serve user traffic simultaneously. All writes flow through the Camunda Exporter, which maintains data consistency regardless of which region handles the request. This is the default with v2 REST API and Tasklist V2 (8.9+).
  • Active-passive: One region handles all user traffic; the other is operational but doesn't serve user requests. Required for deployments using v1 APIs, because v1 stores some state locally per region (batch operations, task assignments).

In both modes, both regions participate in data processing and replication at all times. "Passive" refers only to the user traffic layer.

Active-active

Starting in Camunda 8.8, the v2 REST API removed previous region-specific limitations. In current releases, Tasklist also uses only the Orchestration Cluster REST API, so user task operations are no longer tied to the legacy Tasklist V1 behavior. These improvements make a user-facing active-active setup possible. Starting with version 8.9, active-active routing is the default for dual-region deployments.

Traffic routing​

Primary and secondary regions​

In v2 API deployments (default in 8.9+), both regions serve user traffic simultaneously and there's no primary/secondary distinction at the UI layer. See Active-active and active-passive modes.

In v1 API deployments, one region is designated as primary and the other as secondary:

  • Primary region: Serves user traffic (UI access, API calls).
  • Secondary region: Operational but doesn't serve user traffic under normal conditions.

In both cases, both regions are operationally active with all components running and replicating data.

Managing user traffic​

With v2 APIs (default in 8.9+), distribute traffic across both regions using DNS, a load balancer, or network routing policies.

With v1 APIs, route all user traffic exclusively to the primary region. You're responsible for configuring and maintaining:

  • DNS routing to the primary region
  • Load balancer rules and health checks
  • Traffic redirection to the secondary region during primary region failure, as part of the complete failover procedure
warning

Redirecting traffic without following the full operational procedure can cause system inconsistencies and data issues.

Requirements​

caution

Running dual-region setups requires you to develop, test, and execute custom operational procedures specific to your environment.

Installation requirements​

To install with the Helm chart, you need two Kubernetes clusters.

Database support

Dual-region configurations don't support OpenSearch or RDBMS (relational database) secondary storage.

Network requirements​

  • Kubernetes clusters, services, and pods must use distinct, non-overlapping CIDRs to avoid routing issues.
  • Both regions must be able to communicate with each other (for example, via VPC peering). See the example AWS EKS implementation.
  • Network round-trip time (RTT) between regions directly affects Raft commit latency and throughput. As a guideline, keep RTT at or below 100 ms. Higher latencies degrade performance, but are not a hard limit enforced by the engine.
  • Required open ports between regions:
    • 9200: Elasticsearch (cross-region data pushed by Zeebe)
    • 26500: Zeebe Gateway (client/worker communication)
    • 26501 and 26502: Zeebe broker and Zeebe Gateway communication

Zeebe cluster configuration​

Zeebe supports the following broker and replication configurations for dual-region setups:

  • clusterSize must be a multiple of 2 and at least 4 to distribute brokers evenly across both regions.
  • replicationFactor must be 4 to ensure even partition distribution across regions.
  • partitionCount is unrestricted but should be based on workload requirements. See understanding sizing and scalability behavior and partitions.

Zeebe creates partitions in a round-robin fashion. The Helm chart places all brokers with even numbers (0, 2, 4, 6, ...) in one region and all brokers with odd numbers (1, 3, 5, 7, ...) in the other. This distribution ensures even partition replication across both regions.

Scaling the Zeebe cluster​

When scaling, follow the cluster scaling steps and ensure you meet the Zeebe cluster configuration requirements.

Keep both regions balanced. They should always have the same number of brokers.

Infrastructure and deployment platform considerations​

Multi-region setups require careful planning. You must manage the following areas independently. Camunda doesn't control or document them:

  • Kubernetes cluster management: Managing multiple Kubernetes clusters and deployments across regions
  • Monitoring and alerting: Dual-region monitoring with cross-region correlation
  • Cost implications: Multiple clusters and cross-region traffic increase costs
  • Network reliability: Increased latency can affect data consistency and synchronization. Even short latency bursts can have an impact.
  • Traffic management: DNS and incoming traffic routing
  • Security: Consistent security policies and network controls across regions
Operational readiness

Before implementing dual-region, ensure your organization has:

  • Experience managing multi-cluster Kubernetes environments
  • Established procedures for cross-region networking and security
  • Monitoring and alerting systems with cross-region correlation capability
  • Defined RTO/RPO requirements and tested recovery procedures

Upgrade considerations​

Follow the upgrade recommendations in the Camunda Helm chart and the component-specific upgrade guides.

Review the upgrade overview before starting, and always create a Camunda-supported backup first.

For dual-region setups, use a staged upgrade approach: upgrade one region at a time. Upgrading both regions simultaneously risks quorum loss in Zeebe partitions. Complete the upgrade in one region before starting the other, updating only one Zeebe broker at a time.

Certain minor version upgrades might require you to upgrade both regions simultaneously to complete migration steps. Always check the release notes and migration instructions for your version before proceeding.

Limitations​

AspectDetails
Installation methodsOnly Kubernetes with the Camunda Helm chart is supported. Alternative installation methods such as docker-compose aren't covered by our guides.
v1 API user trafficDeployments using v1 APIs don't support active-active user traffic routing. All user traffic must be routed to a single primary region. See Active-active and active-passive modes.
Management IdentityNot deployed by the dual-region reference configuration, which uses Basic authentication. The Orchestration Cluster-level Admin provides multi-tenancy and role-based access control (RBAC) instead. Deploy Management Identity in a single region if you need Optimize or Web Modeler. See Management platform and Orchestration Cluster.
OptimizeRuns in a single region alongside a dual-region cluster, with no multi-region guarantees. Requires OIDC authentication and Management Identity. If its region is lost, Optimize is unavailable until you restore it. See Management platform and Orchestration Cluster.
Connectors deploymentConnectors can be deployed in a dual-region setup, but you must account for idempotency to avoid event duplication. With two connector deployments running, message idempotency is critical.
ConnectorsIf you run Connectors with an inbound connector deployed in a dual-region setup:
  • To delete a process deployment, do so via Operate, otherwise the inbound connector won't deregister.
  • If you have multiple Operate instances running, delete the process in both instances. This is a known limitation.
Zeebe cluster scalingSupported. See Zeebe cluster configuration.
Web ModelerRuns in a single region alongside a dual-region cluster, with no multi-region guarantees. Requires OIDC authentication, Management Identity, and PostgreSQL. If its region is lost, Web Modeler is unavailable until you restore it. See Management platform and Orchestration Cluster.

Region failure and recovery​

In a dual-region setup, losing either region affects Camunda 8 processing because of Zeebe's quorum requirements.

When a region becomes unavailable, the Zeebe cluster loses quorum (half of its brokers become unreachable) and immediately stops processing new data. All components stop processing until the failover procedure completes.

This section covers the Orchestration Cluster. Failover doesn't recover Management Identity, Web Modeler, Console, or Optimize. For their behavior on region loss, see Management platform and Orchestration Cluster.

Immediate impact

Region failure causes immediate service interruption:

  • No new process instances can start.
  • Running process instances are suspended.
  • User interfaces become unavailable if the primary region is lost.

See the operational procedure for recovery and re-establishment steps.

caution

Monitor for region failures and execute the operational procedures promptly to ensure smooth recovery.

Primary region failure​

If the primary region fails:

  • Service disruption: User traffic is unavailable.
  • Zeebe halt: Processing stops due to quorum loss.
  • Data loss: With v1 APIs, region-specific data (batch operations, task assignments) is lost. With v2 REST API and Tasklist V2, the Camunda Exporter replicates all data to both regions, so data remains available after region failure.

Recovery steps for primary region failure​

  1. Temporary recovery: Follow the operational procedure to restore functionality and unblock the process automation engine (Zeebe).
  2. Traffic rerouting: With v2 APIs (default in 8.9+), remove the failed region from serving traffic (for example, via DNS or load balancer health checks). With v1 APIs, redirect user traffic to the secondary region (now primary).
  3. Data and task management (v1 API setups only):
    • Reassign uncompleted tasks lost from the previous primary region.
    • Recreate batch operations in Operate.
  4. Permanent region setup: Follow the operational procedure to create a new secondary region.

Secondary region failure​

If the secondary region fails:

  • Zeebe halt: Processing stops due to quorum loss.
  • UI availability: Operate and Tasklist remain accessible from the primary region.
  • Process execution: No new process instances can start. Zeebe suspends running instances until quorum is restored.

Recovery steps for secondary region failure​

  1. Temporary recovery: Follow the operational procedure to restore processing.
  2. Permanent region setup: Follow the operational procedure to create a new secondary region.
note

Unlike primary region failure, no user-facing data is lost and no traffic rerouting is necessary.

Recovery objectives (RPO and RTO)​

Based on the requirements and limitations outlined in this page, you can use the Recovery Point Objective (RPO) and Recovery Time Objective (RTO) values below to inform your risk assessment.

The RPO is the maximum tolerable data loss measured in time.

The Recovery Time Objective (RTO) is the time required to restore services to a functional state.

For Operate, Tasklist, and Zeebe, the RPO is 0.

The RTO applies to both the failover and failback procedures:

  • Failover RTO: < 1 minute to restore a functional state, excluding DNS reconfiguration and network considerations.
  • Failback RTO: Five minutes plus the time required to back up and restore Elasticsearch. This depends on your setup and chosen Elasticsearch backup type.

In internal tests, reinstalling and reconfiguring Camunda 8 takes approximately five minutes. Treat this as a general guideline. Actual times vary depending on available resources and your familiarity with the procedure.