For the complete documentation index, see llms.txt.
Skip to main content
Version: 8.10 (unreleased)

Multi-region setup with RDBMS (EKS)

This guide deploys one Camunda 8 Orchestration Cluster across three AWS regions. It uses Amazon EKS for compute and AWS Transit Gateway for inter-region routing. It uses Submariner for cross-cluster service discovery and Aurora Global Database as relational secondary storage.

caution

Review the Multi-Region RDBMS concept documentation before continuing, to understand the limitations and requirements of this configuration.

The result is a cluster where losing a region does not stop processing. Bringing the region back is a redeployment rather than a data restore. For the reasoning behind the topology, see partition placement across zones.

High-level design​

Three AWS regions, each with an EKS cluster in its own VPC and a Camunda zone. A Transit Gateway per region is peered in a full mesh, Submariner publishes each region's Zeebe service under a clusterset name, and an Aurora Global Database with a writer in eu-west-2 and a reader in eu-west-3 backs all three regions through a single JDBC URL.Multi-region with RDBMS on Amazon EKSThree EKS clusters, one Orchestration Cluster, one Aurora Global Database.Compute spans three regions; the database spans two.COMPUTE one EKS cluster and one Camunda zone per regioneu-west-2 · LondonVPC 10.192.0.0/16svc 10.190.0.0/16EKS camunda-londonnamespace camundazeebe london_0, london_1gateway, connectorszone londoneu-west-3 · ParisVPC 10.202.0.0/16svc 10.200.0.0/16EKS camunda-parisnamespace camundazeebe paris_0, paris_1gateway, connectorszone pariseu-central-2 · ZurichVPC 10.212.0.0/16svc 10.210.0.0/16EKS camunda-zurichnamespace camundazeebe zurich_0, zurich_1gateway, connectorszone zurichWith the AWS VPC CNI a pod address is an ordinary VPC address, so no overlay is needed.L3 AWS Transit Gateway, one hub per regionTGW eu-west-2TGW eu-west-3TGW eu-central-2peering is not transitive: one edge per region pairL7 Submariner, service discovery onlypublishes each region's Zeebe service under one clusterset namelondon.camunda-zeebe.camunda.svc.clusterset.local.One namespace name in every cluster: the cluster ID prefix disambiguates.SECONDARY STORAGE Aurora Global Database, reached over the Transit Gatewaywriter eu-west-2 reader eu-west-3replication and writer failover are the database's jobCamunda is never reconfiguredno database memberin this regionjdbc:aws-wrapper:postgresql://<writer>:5432/camunda?wrapperPlugins=failoverThe same URL in every region. The driver follows the writer across a global failover.

Each layer of the design has one job, and the layers are independent:

LayerComponentWhat it provides
ComputeOne EKS cluster per regionZeebe brokers, gateway, and connectors for one Camunda zone.
L3, inter-region routingAWS Transit GatewayCarries all cross-region traffic, including Raft and the database writes, on private addresses.
L7, service discoverySubmarinerPublishes each region's Zeebe service under a name every other region can resolve.
Secondary storageAurora Global DatabaseOne writer and its readers, replicated by the database, reached through a single JDBC URL.

The database regions are decoupled from the compute regions. The Aurora members live in London and Paris, while compute spans all three regions. That keeps the database cheaper, and the two topologies stay independent.

tip

New to Terraform or to running Camunda on EKS? Start with the single-region EKS Terraform setup, which covers AWS authentication, Terraform state management, and the essentials of an EKS cluster. This guide assumes you have completed a single-region deployment at least once.

Requirements​

  • AWS account: required to create AWS resources in every target region. See What is an AWS account?.
  • AWS CLI: command-line tool to manage AWS resources. Install AWS CLI.
  • Terraform: IaC tool to provision resources. Install Terraform.
  • kubectl: CLI to interact with Kubernetes clusters. Install kubectl.
  • Helm: package manager for Kubernetes. Install Helm.
  • jq: lightweight JSON processor. Download jq.
  • subctl: Submariner CLI. The reference architecture installs it for you.

For the tool versions used in testing, see the repository's .tool-versions file.

AWS service quotas​

Verify your quotas in every region before deploying, and request increases where needed:

  • Elastic IPs: at least three per region, one per availability zone.
  • VPCs, EC2 instances, and EBS storage: enough for one EKS cluster per region.
  • Transit Gateways: one per region, plus one peering attachment per region pair.
  • Aurora Global Database: available in the regions you choose for the database. Aurora Global Database is not offered in every region.

Some AWS regions are opt-in and must be enabled on the account before anything can be created in them. eu-central-2 (Zurich), used as the third region in this guide, is one of them:

aws account enable-region --region-name eu-central-2

Considerations​

  • This is a multi-region deployment, and costs scale with the region count. You pay for three EKS control planes and node groups. You also pay for three Transit Gateways with a full peering mesh, billed per attachment-hour. An Aurora Global Database adds a member per database region. AWS bills inter-region data transfer per gigabyte. Destroy the environment when you are done evaluating.
  • Non-overlapping CIDRs are mandatory. Transit Gateway cannot route duplicate prefixes, and Submariner runs without Globalnet, so every CIDR must identify exactly one cluster.
  • Round-trip time between regions matters. Keep it at or below 100 ms. The regions used in this guide are London, Paris, and Zurich, whose pairwise round-trip times are well inside that budget.
  • Management Identity, Web Modeler, Console, and Optimize are not part of this deployment. See limitations.
  • This guide is a blueprint, not a production deployment. It shows the moving parts and how they fit together. Adapt sizing, security, and traffic routing to your environment.

Outcome​

Following this guide gives you:

  • Three EKS clusters, one per region, each with its own VPC and a dedicated non-overlapping CIDR.
  • A Transit Gateway per region, peered in a full mesh, routing every VPC and Kubernetes service range between regions.
  • Submariner service discovery, publishing each region's Zeebe service as <clusterID>.<service>.<namespace>.svc.clusterset.local.
  • An Aurora Global Database with a writer in one region and readers in the others, reached through a single JDBC URL.
  • One Orchestration Cluster with six brokers, six partitions, and a replication factor of five. Each database region holds two replicas of every partition, and the third region holds one.

Topology​

The default topology uses three regions and three zones:

SettingDefaultMeaning
Regionseu-west-2, eu-west-3, eu-central-2London, Paris, Zurich
Zone nameslondon, paris, zurichOne zone per region
orchestration.partitioning.schemezone-awareZone-aware partitioning
numberOfBrokers per zone2Brokers deployed in that zone
numberOfReplicas per zone2, 2, 1Two in each database region, one in the tie-breaker
orchestration.clusterSize6Sum of numberOfBrokers across zones. The zone list derives it
Replication factor5Sum of numberOfReplicas across zones
orchestration.partitionCount6One partition per broker
Database regionsSlots 0 and 1Aurora members, writer first

Each broker has the name <zone>_<index>, so paris_1 is the second broker in the Paris zone. The zone list is identical in every region, as the values section shows.

CIDR allocation​

Every region owns a distinct VPC range and a distinct Kubernetes service range. Both are routed over the Transit Gateway:

SlotRegionVPC and pod CIDRKubernetes service CIDR
0eu-west-210.192.0.0/1610.190.0.0/16
1eu-west-310.202.0.0/1610.200.0.0/16
2eu-central-210.212.0.0/1610.210.0.0/16

A fourth slot (eu-south-1, VPC 10.222.0.0/16, service CIDR 10.220.0.0/16) is prepared but disabled. Enable it in variables.tf before you bootstrap the cluster.

There is no separate pod range. With the AWS VPC CNI, pod IPs are VPC addresses. The Transit Gateway routes VPC CIDRs to enable cross-region pod-to-pod traffic. The connectivity check verifies reachability using the per-pod DNS records that Zeebe brokers dial; it does not test remote ClusterIP access. Service CIDRs are also routed in this reference topology.

1. Configure AWS and apply Terraform​

Obtain a copy of the reference architecture​

You need a local copy of the aws/kubernetes/eks-multi-region-rdbms reference architecture, from the camunda-deployment-references repository. It holds the Terraform modules, the Helm values, and every procedure script this documentation refers to.

The following clones the repository and changes into the architecture directory. Every command in this documentation runs from there.

aws/kubernetes/eks-multi-region-rdbms/procedure/get-your-copy.sh
loading...

The reference architecture is a starting point you own and extend, not a module you consume, so the workflow is to copy it into your own repository rather than reference it remotely.

Review the region topology​

The region slots are declared in variables.tf. Adjust the regions, short names, and CIDR blocks to your environment before applying.

Two variables control the topology, and they are not interchangeable:

VariableMeaning
regionsThe full list of region slots the cluster will ever have. Every slot contributes a zone to the Camunda zone list.
active_region_countHow many of those slots are actually deployed. At most one slot may be left empty.

Declaring a slot without deploying it is the supported growth path. The deployed slots are always the first ones in the list, so the empty slot is the last. Activating it later fills in its replicas without redistributing anything. What that costs depends on how many replicas the empty slot holds, not on how many slots there are. The default layout gives the trailing slots one replica each. Every partition therefore runs at four of five with three slots, and five of six with four. The cluster tolerates no further zone loss while a slot is empty.

Leaving two or more slots empty fails at plan time. The guard counts region slots rather than replicas, and requires the deployed slots to be a majority of the declared ones. It can therefore also refuse a layout whose deployed zones would still hold a majority of the replicas. The replica-level test runs separately in export_environment_prerequisites.sh, which catches a layout whose undeployed zones hold the majority.

Apply the infrastructure​

The root module creates every EKS cluster, the Transit Gateway mesh, the security group rules, and the Aurora Global Database in a single state.

Keep your settings in a variable file. The reference architecture ships no variable file, so create one:

terraform-cluster.tfvars
cluster_name            = "camunda"
active_region_count = 3
np_desired_node_count = 4
single_nat_gateway = false
database_instance_class = "db.r6g.large"
default_tags = {
environment = "evaluation"
}

Then apply it:

cd terraform/clusters
terraform init
terraform apply -var-file=terraform-cluster.tfvars

Expect roughly 25 minutes for the EKS clusters and 15 minutes for the Aurora Global Database. They are created in parallel.

For a cheaper evaluation, deploy two of the three slots and reduce the node count. This is a valid state. With the default 2-2-1 layout, every partition holds four of its five replicas. The third slot's replica stays reserved until you deploy it.

terraform-cluster.tfvars
cluster_name          = "camunda"
active_region_count = 2
single_nat_gateway = true
np_desired_node_count = 2
note

Set up remote Terraform state before deploying anything you intend to keep. The single-region EKS guide covers creating an S3 backend.

Bring your own database​

To run this architecture on a database other than Aurora Global Database, set deploy_database = false and supply your own JDBC URL through CAMUNDA_RDBMS_URL. Anything that presents a single endpoint following its own writer works the same way. Examples are a PostgreSQL cluster behind a floating endpoint, a connection proxy, or a DNS record you repoint during failover.

The generated Helm values use the LOG_SEQ replication strategy, which fails at startup on a backend without LSN support. For such a backend, see replication-agnostic secondary storage to choose DELAY instead.

2. Prepare the environment​

Export the Terraform outputs​

The procedure scripts derive the entire environment from the Terraform state, down to the kubectl context aliases, so nothing has to be typed twice.

cd ../../procedure
. ./export-terraform-outputs.sh
. ./export_environment_prerequisites.sh
note

The dot is required. These scripts export variables into your current shell, not into a subshell.

export_environment_prerequisites.sh defines the environment contract of the architecture. You can override every value by exporting it beforehand. Region-indexed values are space-separated lists in slot order.

export-terraform-outputs.sh always sets CAMUNDA_RDBMS_URL, CAMUNDA_RDBMS_USERNAME, and CAMUNDA_RDBMS_PASSWORD from the Terraform outputs. They are empty with deploy_database = false. If you bring your own database, export these three values after that script and before export_environment_prerequisites.sh.

See the export_environment_prerequisites.sh script
aws/kubernetes/eks-multi-region-rdbms/procedure/export_environment_prerequisites.sh
loading...

The script refuses to continue if the topology is inconsistent, for example if more than one slot is left empty.

One namespace in every cluster

The dual-region setup needs a different namespace per region, because CoreDNS stub forwarding cannot distinguish local from remote traffic. This architecture instead uses the same namespace name in every cluster. Submariner disambiguates identically named services with the cluster ID prefix.

Register the kubectl contexts​

Create one kubectl context per active region, named after the region's short name, for example cluster-london. The rest of the guide selects regions by these names.

./register-kubecontexts.sh
See the register-kubecontexts.sh script
aws/kubernetes/eks-multi-region-rdbms/procedure/register-kubecontexts.sh
loading...

Configure the storage class​

Zeebe brokers need a storage class backed by fast disks. Apply it in every region:

./storageclass-configure.sh
See the storageclass-configure.sh script
aws/kubernetes/eks-multi-region-rdbms/procedure/storageclass-configure.sh
loading...

Verify it before continuing. A missing storage class leaves broker PVCs unbound and the pods pending, which is easy to misread later as a networking failure:

./storageclass-verify.sh
See the storageclass-verify.sh script
aws/kubernetes/eks-multi-region-rdbms/procedure/storageclass-verify.sh
loading...

3. Connect the clusters​

Two layers connect the regions, and they have different jobs:

  • Transit Gateway: carries the traffic.
  • Submariner: publishes service names across clusters.

Install subctl​

Install the Submariner CLI and put it on your PATH. Source the script rather than executing it, so the PATH change survives in your shell.

source ./submariner/install-subctl.sh
See the install-subctl.sh script
aws/kubernetes/eks-multi-region-rdbms/procedure/submariner/install-subctl.sh
loading...

Deploy the Submariner broker​

Deploy the ClusterSet broker into one region. It stores ClusterSet metadata only, so any cluster can host it and its loss does not interrupt anything already established.

./submariner/deploy-broker.sh
See the deploy-broker.sh script
aws/kubernetes/eks-multi-region-rdbms/procedure/submariner/deploy-broker.sh
loading...

Submariner is deployed with its service-discovery component only. It provides multi-cluster DNS and nothing else: no gateway nodes, no IPsec tunnel, no route agent. subctl show connections is empty by design.

Why there is no encrypted overlay

Running Submariner's connectivity component alongside the AWS VPC CNI puts two owners on the same prefixes. Submariner installs node routes for every remote cluster CIDR so it can pull that traffic into its tunnel. With the VPC CNI those CIDRs are the VPC ranges, which the Transit Gateway also routes, including the node addresses the tunnels are built on. The result is tunnels that report connected, cross-cluster DNS that resolves correctly, and Raft messages that are silently dropped.

Removing one of the two owners removes the whole class of problem. You cannot remove the Transit Gateway, so remove the other owner.

The traffic is still encrypted. AWS encrypts inter-region Transit Gateway peering itself. AES-256 protects the traffic at the virtual network layer as it travels between regions. AWS encrypts it again at the physical layer, on links outside its physical control. See transit gateway peering attachments and encryption in transit.

You give up control of the encryption. The keys are AWS-managed. If a control requires customer-managed keys, enable TLS in the workload, or replace the VPC CNI with Cilium in ENI mode plus WireGuard or IPsec. In ENI mode pod addresses stay ordinary VPC addresses, so the Transit Gateway remains the only owner of the routes.

Join the clusters to the ClusterSet​

Join every active region to the ClusterSet, so each one can publish and resolve the others' services.

./submariner/join-clusters.sh
See the join-clusters.sh script
aws/kubernetes/eks-multi-region-rdbms/procedure/submariner/join-clusters.sh
loading...

Then verify:

./submariner/verify-submariner.sh
See the verify-submariner.sh script
aws/kubernetes/eks-multi-region-rdbms/procedure/submariner/verify-submariner.sh
loading...

Verify the cross-region substrate​

Check that pods in one region can reach pods in another before you deploy Camunda. Submariner does not carry this traffic, so nothing else covers the Transit Gateway routes and the security group rules.

./setup-namespaces.sh
./verify-cross-region-connectivity.sh

If this fails, the problem is routing or firewalling, not Camunda. See troubleshooting.

The probe image is set in diagnose-submariner.sh. Set PROBE_IMAGE before the script to pull from an approved mirror instead:

export PROBE_IMAGE=my-registry.example.com/busybox

Ports open between regions​

The security group rules are declared explicitly in security.tf rather than allowing all traffic between VPCs:

PortProtocolPurpose
26500 - 26502TCPZeebe gateway gRPC, command API, and the internal API carrying Raft
8080TCPOrchestration Cluster REST API
53TCP/UDPCoreDNS and Submariner service discovery
n/aICMPCross-region connectivity diagnostics

Terraform creates each rule once per remote VPC range and once per remote service range. The rule count therefore grows linearly with the region count: 20 inbound rules at three regions and 30 at four. The AWS limit is 60 per security group. Terraform asserts that budget at plan time rather than letting the apply fail after the clusters exist.

4. Deploy Camunda 8​

Create the database secret​

Create the Kubernetes secret holding the database password, in every active region. The Helm values reference it by name rather than carrying the password.

./create-rdbms-secret.sh
See the create-rdbms-secret.sh script
aws/kubernetes/eks-multi-region-rdbms/procedure/create-rdbms-secret.sh
loading...

Generate the region-dependent values​

Three values cannot be hardcoded in the Helm values, because they depend on the deployed topology:

ValueContent
CAMUNDA_CLUSTER_INITIALCONTACTPOINTSOne entry per active region, pointing at that region's headless Zeebe service.
REGION_<slot>_ZEEBE_SERVICE_NAMEThe suffix each broker advertises, so peers in other regions can resolve it.
CAMUNDA_MULTIREGION_ZONESThe zone list, covering every slot including any not yet deployed.
. ./generate-zeebe-helm-values.sh
./assemble-envsubst-values.sh
Contact points are fully qualified

The generated contact points end with a trailing dot, which marks them as fully qualified names. Without it, the resolver walks the pod's search domains first. On a cold multi-region start, a broker whose peer is not yet published can then exhaust its DNS budget and never finish starting. Always keep the trailing dot.

Review the Helm values​

The values file is the same in every region. Only orchestration.partitioning.zone and the advertised host differ, so one file describes the whole topology.

See the full camunda-values.yml
aws/kubernetes/eks-multi-region-rdbms/helm-values/camunda-values.yml
loading...

The parts worth reading before you install:

  • orchestration.partitioning.scheme: zone-aware selects zone-aware partitioning. The chart rejects numberOfZones and zoneIndex with this scheme, because the zone list describes the topology instead. The chart derives the cluster size, replication factor, and broker node IDs from that list. See configure zone-aware multi-region deployments.
  • orchestration.partitioning.zones lists every zone with its broker count, replica count, and priority. Zone 0 has the highest priority because it hosts the database writer.
  • orchestration.data.secondaryStorage.type: rdbms with a single url shared by every broker in every region.
  • The AWS Advanced JDBC Wrapper uses initialConnection,failover. initialConnection discovers the current writer when a broker starts after a switchover. failover follows a writer change on an established connection.
  • CAMUNDA_DATA_SECONDARYSTORAGE_RDBMS_ASYNCREPLICATION_ENABLED: "true" is required. Without it the exporter acknowledges records the standby has not received, and a writer failover loses exported data.
  • The reference architecture also pins CAMUNDA_DATA_SECONDARYSTORAGE_RDBMS_ASYNCREPLICATION_TYPE: LOG_SEQ, CAMUNDA_DATA_SECONDARYSTORAGE_RDBMS_ASYNCREPLICATION_MAXLAG: PT1H, and CAMUNDA_DATA_SECONDARYSTORAGE_RDBMS_ASYNCREPLICATION_PAUSEONMAXLAGEXCEEDED: "false". The first two set a strategy and a lag budget. The third keeps the engine default, because pausing stops exporting. Enable pausing only on purpose. The lag budget has no effect until you enable pausing. The engine compares it only inside the pause condition.
  • Cross-region SWIM membership timeouts are relaxed. The defaults are tuned for intra-region latency, and on a cold start brokers otherwise see remote peers as unreachable, eject them, and never converge.
  • identity, console, and optimize are disabled. See limitations.

Install the chart​

Install the same release in every active region, from the values assembled in the previous step.

./install-chart.sh
See the install-chart.sh script
aws/kubernetes/eks-multi-region-rdbms/procedure/install-chart.sh
loading...

Then export the Camunda services to the ClusterSet so brokers in other regions can resolve them:

./submariner/export-services.sh
See the export-services.sh script
aws/kubernetes/eks-multi-region-rdbms/procedure/submariner/export-services.sh
loading...

5. Verify the deployment​

Verify that every broker joined and that the partition distribution matches the zone list:

./check-cluster-topology.sh
See the check-cluster-topology.sh script
aws/kubernetes/eks-multi-region-rdbms/procedure/check-cluster-topology.sh
loading...

Expect roughly 10 minutes for the Zeebe cluster to converge across regions. A healthy three-zone cluster reports six brokers, six partitions, and a replication factor of five.

Measure the cost of the write path from each region to the database writer. Regions that are not co-located with the writer pay the inter-region round trip on every export flush. That number tells you whether the exporter queue is sized correctly:

./measure-rdbms-latency.sh
See the measure-rdbms-latency.sh script
aws/kubernetes/eks-multi-region-rdbms/procedure/measure-rdbms-latency.sh
loading...

6. Operate the cluster​

Day-2 procedures, including region loss, failback, and activating a declared zone, are documented separately in the Multi-Region RDBMS operational procedure.

Troubleshooting​

Cross-region pod traffic is dropped​

Prove the substrate before investigating Camunda:

./verify-cross-region-connectivity.sh

Submariner does not carry this traffic, so do not start with subctl. The data plane is the Transit Gateway. Check that the remote ranges are routed, expecting one route per remote VPC and service CIDR:

aws ec2 describe-route-tables --region eu-west-2 \
--filters "Name=vpc-id,Values=<vpc-id>" \
--query 'RouteTables[].Routes[?TransitGatewayId!=null].[DestinationCidrBlock]' --output text

Then check that the remote security group allows the port:

aws ec2 describe-security-group-rules --region eu-west-3 \
--filters "Name=group-id,Values=<security-group-id>" \
--query 'SecurityGroupRules[?!IsEgress].[CidrIpv4,IpProtocol,FromPort,ToPort]' --output text

Names do not resolve across regions​

If routing is correct but names do not resolve, the problem is service discovery:

subctl show networks --contexts cluster-london
kubectl --context cluster-london -n submariner-operator get clusters.submariner.io
kubectl --context cluster-london -n camunda get serviceexports,serviceimports

A full diagnostic dump, including cross-cluster name resolution, is available:

./submariner/diagnose-submariner.sh
See the diagnose-submariner.sh script
aws/kubernetes/eks-multi-region-rdbms/procedure/submariner/diagnose-submariner.sh
loading...

Zeebe never reaches the expected broker count​

This is almost always cross-region DNS. Check the service discovery layer first, then the broker's startup gate:

./submariner/verify-submariner.sh
kubectl --context cluster-london -n camunda logs camunda-zeebe-0 -c wait-clusterset-dns

If brokers are Pending rather than Running, the storage class is missing in that region. See configure the storage class.

Next steps​