For the complete documentation index, see llms.txt.
Skip to main content
Version: 8.10 (unreleased)

Migrate to zone-aware brokers

This procedure migrates an existing Orchestration Cluster from numbered broker identities to zone-aware broker identities with the Camunda 8.10 Helm chart. You can migrate both single-region and dual-region clusters.

The migration creates replacement brokers instead of changing the persisted identity of existing brokers. The numbered and zone-aware broker generations run together until the cluster has moved its partitions and membership to the new brokers.

Migration overview​

The migration follows the same steps for single-region and dual-region clusters. The only difference is the number of zones and Helm releases involved.

TopologyHelm releasesZonesZone migration order
Single-regionOne release in one Kubernetes clusterOne zoneThe single zone
Dual-regionOne release per Kubernetes clusterOne zone per region, 2 totalzoneIndex: 1 first, then zoneIndex: 0

In a dual-region cluster, the primary zone is the region with zoneIndex: 0, whose numbered brokers have even node IDs. The secondary zone is the region with zoneIndex: 1, whose numbered brokers have odd node IDs. You must start a dual-region migration with the secondary zone (zoneIndex: 1). The numbered broker with node ID 0 belongs to the primary zone and coordinates cluster configuration changes. Migrating the secondary zone first keeps this coordinator in place while the other zone migrates. If you start with the primary zone, the management API rejects the request with an error similar to:

Zone migration must proceed from the highest remaining zone index to the lowest. Expected next zoneIndex 1 but got 0.

In a dual-region cluster, run every Helm step once per release, one release at a time, and wait for each release to be healthy before you continue with the next one. You can send each management API request to any node in the cluster. However, configure the port-forward on the region with zoneIndex: 0, since that region is the last to be migrated.

The procedure consists of these steps:

  1. Upgrade every release to the chart version that supports zone-aware migration.
  2. Upgrade every release with zone-aware values that keep the numbered brokers.
  3. Update the persisted partitioning configuration once.
  4. For each zone, add the zone's zone-aware brokers to the cluster with the management API, then remove the numbered brokers of that zone's release.

The migration is not reversible. After you update the persisted partitioning configuration, you can't return the cluster to numbered brokers. If a later step fails, complete the migration instead of reverting it. See Recover from an incomplete migration.

Before you begin​

  • Back up the Orchestration Cluster before you start this procedure.
  • Confirm that the existing Helm releases and their numbered brokers are healthy. See Check broker health.
  • Confirm that each Kubernetes cluster has enough capacity (nodes) for both broker generations and their persistent volume claims (PVCs).
  • Back up the values used by each Helm release.
  • Suspend planned node drains, autoscaler scale-down, and other maintenance that could evict broker pods until the migration is complete.

The examples in this procedure use these variables. Set them separately for each release.

export RELEASE="camunda-platform"
export NAMESPACE="camunda"
export CHART="camunda/camunda-platform"
export CHART_VERSION="<chart-version>"
export VALUES="values.yaml"
export LOCAL_ZONE="zone-a"
export MANAGEMENT_URL="http://127.0.0.1:9600"

Replace the example values with values from your installation. Set CHART_VERSION to the Camunda 8.10 chart version that supports zone-aware migration.

Access the management API​

Use the Orchestration management API to change the cluster topology. To reach it, and for its port, security, and TLS options, see About this API. Set MANAGEMENT_URL to the resulting address. To check broker health, you also need access to the brokers of every release.

For example, forward the management port of the release's gateway Service to your machine:

kubectl port-forward "svc/$RELEASE-zeebe-gateway" 9600:9600 --namespace "$NAMESPACE"

During the migration, this Service selects both the numbered and the zone-aware brokers. You can send the requests through any of them, because each broker forwards cluster configuration requests to the broker that coordinates the change.

Several requests in this procedure start an asynchronous configuration change and return a changeId. Track each change as described in Monitor a configuration change, and continue only after it reaches the COMPLETED status.

Check broker health​

Check broker health with the health check endpoint of each broker pod. A finished rollout and an ACTIVE state in GET /actuator/cluster don't prove that the brokers can process: a broker passes its readiness probe and stays ACTIVE in the topology even if its partitions fail to start.

Query each broker pod directly, not through a Service, for example after kubectl port-forward pod/<broker-pod> 9600:9600 --namespace "$NAMESPACE":

curl --fail "$MANAGEMENT_URL/actuator/health/status"

A healthy broker returns HTTP 200. If a broker returns 503, check its logs and resolve the problem before you continue.

Only brokers that belong to the logical cluster report healthy. During the migration, these brokers are expected to report unhealthy:

  • Zone-aware brokers of a zone that you haven't migrated yet.
  • Numbered brokers of a zone that you have migrated but not yet removed from the release.

Upgrade to the migration chart​

Upgrade each existing release to the chart version that supports zone-aware migration, with its existing values unchanged.

Keep orchestration.partitioning.keepUnzonedBrokers disabled during this chart upgrade. Wait for the rollout to finish before you enable the migration flag. If you enable the flag in the same command as the chart upgrade, numbered brokers can restart during migration.

helm upgrade "$RELEASE" "$CHART" \
--version "$CHART_VERSION" \
--namespace "$NAMESPACE" \
--values "$VALUES" \
--wait \
--timeout 15m

The --timeout value is an example. Increase it if broker rollouts in your cluster take longer.

After each rollout, check that every broker is healthy before you upgrade the next release or enable the migration flag.

Configure the zone-aware values​

Add the zone-aware configuration to the values of each release. The zones list describes the complete target topology and must be identical in every release. The zone value selects the zone owned by this release.

Keep these values unchanged while keepUnzonedBrokers is true, because they still describe the retained numbered brokers:

ValueRequired value during migration
orchestration.clusterSizeThe existing numbered cluster size. The chart divides it by numberOfZones to size the numbered StatefulSet.
orchestration.replicationFactorThe current replication factor of the cluster.
orchestration.partitioning.numberOfZonesThe number of regions of the numbered deployment: 1 for single-region, 2 for dual-region. clusterSize must be divisible by it.
orchestration.partitioning.zoneIndexThe region index of this release in the numbered deployment: 0 for single-region, 0 or 1 for dual-region.

For a single-region cluster, you can omit numberOfZones and zoneIndex. The chart defaults to numberOfZones: 1 and zoneIndex: 0.

If the existing values use the deprecated global.multiregion block, remove it when you add orchestration.partitioning. Set numberOfZones to the former regions value and zoneIndex to the former regionId value. The chart rejects values that configure both blocks.

The sum of the zones' numberOfReplicas values must equal the cluster's current replication factor. If the target topology needs a higher factor, see Update the partitioning configuration.

Values examples​

This example migrates a single-region cluster with clusterSize: 3 and replicationFactor: 3:

orchestration:
clusterSize: "3"
replicationFactor: "3"
partitioning:
scheme: zone-aware
zone: zone-a
zones:
- name: zone-a
numberOfBrokers: 3
numberOfReplicas: 3
priority: 100

# Keep the numbered brokers during the migration.
keepUnzonedBrokers: true

Check the dual-region networking​

A dual-region cluster already has cross-cluster networking in place, as described in dual-region setup. The zone-aware brokers reuse the existing CAMUNDA_CLUSTER_INITIALCONTACTPOINTS value and DNS setup, so you don't need to configure new networking. Single-region clusters can skip this section.

Before you start the migration, check how the existing values address the brokers:

  • Initial contact points: Addresses of the shared headless Service, such as camunda-zeebe.camunda-london.svc.cluster.local:26502, keep working during and after the migration, because the Service selects both the numbered and the zone-aware brokers. Addresses of individual numbered pods, such as camunda-zeebe-0.camunda-zeebe.camunda-london.svc.cluster.local:26502, stop resolving when you remove the numbered brokers. Replace them with the shared Service address.
  • Advertised host: The chart applies orchestration.env to every broker pod in the release, including both the numbered and the zone-aware StatefulSet. If the values set CAMUNDA_CLUSTER_NETWORK_ADVERTISEDHOST to a fixed value, several brokers advertise the same address. Derive the advertised host from each pod instead.

For example, if brokers use host networking, advertise the node IP of each pod, and configure pod anti-affinity so that no two broker pods, numbered or zone-aware, run on the same node:

orchestration:
hostNetwork: true
env:
- name: CAMUNDA_CLUSTER_NETWORK_ADVERTISEDHOST
valueFrom:
fieldRef:
fieldPath: status.hostIP

Start the migration​

Upgrade each release with the zone-aware values.

helm upgrade "$RELEASE" "$CHART" \
--version "$CHART_VERSION" \
--namespace "$NAMESPACE" \
--values "$VALUES" \
--wait \
--timeout 15m

Each release now contains both the existing numbered StatefulSet and a new zone-specific StatefulSet. Wait for the zone-specific brokers to become ready before you continue. Don't remove the numbered brokers yet. The new brokers haven't joined the logical cluster, so the numbered brokers are still the only active members.

Check that the numbered brokers are still healthy. The new zone-aware brokers report unhealthy until you migrate their zone.

Update the partitioning configuration​

Use the partitioning API once to update the persisted partitioning configuration. This is the point after which you can't revert the migration.

The order of the zones is significant during this one-time migration. The first zone must be the zone of the numbered brokers with zoneIndex: 0, the second zone the one with zoneIndex: 1.

The sum of the zones' numberOfReplicas values must equal the cluster's current replication factor. Otherwise, the request fails with an error similar to:

Sum of zone replicas [2] must equal the current replication factor [1] before zone migration starts.

If the target topology needs a higher factor, first increase the numbered cluster's replication factor with the cluster scaling API, and monitor the change until it completes. Then set orchestration.replicationFactor in the values of every release to the new factor and upgrade the releases, so the retained numbered brokers use the updated configuration.

curl --fail --request PUT \
"$MANAGEMENT_URL/actuator/cluster/partitioning" \
--header 'Content-Type: application/json' \
--data @- <<'JSON'
{
"config": {
"scheme": "ZONE_AWARE",
"zones": [
{"name": "zone-a", "numberOfReplicas": 3, "priority": 100}
]
}
}
JSON

The response includes a changeId. Monitor the change until it completes.

Migrate each zone​

Migrate one zone at a time. For each zone, add its zone-aware brokers to the cluster with the management API, then remove the numbered brokers of that zone's release. Removing the numbered brokers right after each zone frees their CPU and memory before you migrate the next zone, which helps when cluster capacity is tight.

In a dual-region cluster, migrate the secondary zone (zoneIndex: 1) first, then the primary zone (zoneIndex: 0), as described in the migration overview. Don't migrate both zones concurrently.

Add the zone's brokers to the cluster​

Use the zone migration endpoint to add the zone's zone-aware brokers to the cluster. The new brokers take over the partitions of the zone's numbered brokers, and the numbered brokers leave the cluster. Before you send the request, check that every broker in the logical cluster is healthy. A partition that can't start blocks the migration, and the change stays IN_PROGRESS.

Set LOCAL_ZONE to the zone to migrate:

curl --fail --request PUT \
"$MANAGEMENT_URL/actuator/cluster/zones" \
--header 'Content-Type: application/json' \
--data "{\"zone\":\"$LOCAL_ZONE\"}"

The response includes a changeId. Monitor the change until it completes.

After the change completes, use GET /actuator/cluster in the Cluster API to check the topology:

curl --fail "$MANAGEMENT_URL/actuator/cluster"

In the response, confirm that:

  • The zone-aware brokers of the migrated zone, with IDs such as zone-a_0, are listed with "state": "ACTIVE" and host the expected partitions.
  • The numbered brokers of the migrated zone, with numeric IDs such as 0, are no longer listed.
  • The zone-aware brokers of the migrated zone report healthy.

Leave keepUnzonedBrokers: true if the zone migration is incomplete. Don't remove the numbered Kubernetes resources while a numbered broker of the zone still owns a partition or remains in cluster membership.

Remove the numbered brokers of the migrated zone​

After the zone migration completes and the zone's numbered brokers no longer own partitions or belong to the logical cluster, remove them from the release that owns the zone. Releases whose zone you haven't migrated yet keep keepUnzonedBrokers: true and their migration values.

Update the values of the release as follows:

  • Set keepUnzonedBrokers: false.
  • Remove numberOfZones and zoneIndex.
  • Remove orchestration.clusterSize and orchestration.replicationFactor, or set them to the totals of the zones list.
  • Keep the same complete zones list and the local zone value.

With the zone-aware scheme and without numbered brokers, the chart derives the cluster size and replication factor from the zones list. If the values still contain conflicting settings, the upgrade fails with errors similar to these:

[camunda][error] orchestration.partitioning.numberOfZones and orchestration.partitioning.zoneIndex cannot be used with the zone-aware scheme; the zone list describes the topology instead.
[camunda][error] orchestration.clusterSize is <size> but orchestration.partitioning.zones sums to <total> brokers. With the zone-aware scheme the zone list is authoritative; remove the key or make it agree.
[camunda][error] orchestration.replicationFactor is <factor> but orchestration.partitioning.zones sums to <total> replicas. With the zone-aware scheme the zone list is authoritative; remove the key or make it agree.

For example, the secondary region (zoneIndex: 1) of the dual-region cluster uses these values:

orchestration:
partitioning:
scheme: zone-aware
zone: zone-b
zones:
- name: zone-a
numberOfBrokers: 4
numberOfReplicas: 2
priority: 100
- name: zone-b
numberOfBrokers: 4
numberOfReplicas: 2
priority: 90
keepUnzonedBrokers: false

Upgrade the release:

helm upgrade "$RELEASE" "$CHART" \
--version "$CHART_VERSION" \
--namespace "$NAMESPACE" \
--values "$VALUES" \
--wait \
--timeout 15m

The upgrade removes the numbered StatefulSet and pods. The zone-aware StatefulSet and shared Services remain managed by Helm. Helm doesn't delete the numbered PVCs, so they remain bound.

Repeat the steps in Migrate each zone for each remaining zone.

Verify the migration​

After you migrate every zone and remove the numbered brokers from every release, query the cluster topology through the management API:

curl --fail "$MANAGEMENT_URL/actuator/cluster"

Confirm that:

  • The partitioning object in the response reports "scheme": "ZONE_AWARE" and lists every zone.
  • Every zone-aware broker, such as zone-a_0 and zone-b_0, is ACTIVE and hosts the expected partitions.
  • No numbered broker is listed.
  • Every zone-aware broker reports healthy.
  • No numbered StatefulSet or pod remains in any release, for example with kubectl get statefulsets,pods --namespace "$NAMESPACE".

Don't delete the numbered PVCs until you have confirmed these checks. Then delete them explicitly according to your storage-retention policy.

Recover from an incomplete migration​

You can't revert the migration after you update the persisted partitioning configuration. If a step fails, keep keepUnzonedBrokers: true and both broker generations running, fix the cause, and complete the migration. Don't delete the numbered PVCs until you have verified the migration.

A zone migration stays in progress​

If the change started by PUT /actuator/cluster/zones stays IN_PROGRESS, a partition usually can't start on one of the brokers involved:

  1. Monitor the change and note the operations in pending, and the brokers they target.
  2. Check the health of those brokers and of every other broker in the logical cluster, and inspect the logs of the brokers that report unhealthy.
  3. Fix the cause, for example missing CPU, memory, or storage, a pod that can't be scheduled, or a networking problem. The change continues after the brokers can apply the pending operations.

Only one configuration change can run at a time, so don't send another zone migration request while the change is IN_PROGRESS.

You can cancel a change with DELETE /actuator/cluster/changes/<changeId>, but canceling doesn't revert the operations already applied and leaves the cluster in an intermediate state that needs manual intervention. Only cancel a change that can't make progress, and contact Camunda support before you do.

Numbered brokers were removed too early​

If you removed the numbered brokers of a zone while they still owned partitions or belonged to the logical cluster, restore them before you continue:

  1. In the values of the release, restore the migration values: keepUnzonedBrokers: true, numberOfZones, zoneIndex, orchestration.clusterSize, and orchestration.replicationFactor.
  2. Upgrade the release. The chart recreates the numbered StatefulSet with the same name, so its pods reattach to the retained numbered PVCs and rejoin the cluster with their data.
  3. Check that the numbered brokers are healthy, then continue with Add the zone's brokers to the cluster.

If the numbered PVCs were deleted, you can't restore these brokers. Try to complete the migration without them first: if every partition still has a quorum of replicas on the remaining brokers, the zone migration can finish, and the zone-aware brokers replicate the data from the other replicas. Monitor the change, or start it with Add the zone's brokers to the cluster if you haven't sent the request yet. Restore the cluster from the backup you took before the migration only if the migration can't complete.