For the complete documentation index, see llms.txt.
Skip to main content
Version: 8.10 (unreleased)

Restore a backup with the Restore API (RDBMS)

Restore Zeebe partition data through the Orchestration Cluster Restore API without restarting the brokers, when using a relational database management system (RDBMS) as secondary storage.

This page is part of the RDBMS restore procedure. To compare it with the legacy Restore Application, see choosing a restore approach.

About the Restore API​

With Camunda 8.10 and later, a Restore API recovery runs during a downtime window while the cluster is in recovery mode. It runs in four phases, driven by two API requests:

  1. Entering recovery mode: every broker deactivates its partitions and switches to a restricted partition manager. While the cluster is in recovery mode it processes no work, and only read-only operations and restore remain available.
  2. Restoring secondary storage: while the cluster is in recovery mode, restore the RDBMS to the intended point that the primary storage backup aligns to.
  3. Restoring the partitions: the cluster plans a single change that, for every broker and partition, first drops the local partition data and then restores that partition from the selected backups. The steps of that plan run one at a time across the cluster.
  4. Returning to processing: once every partition is restored, the same change switches all brokers back to PROCESSING and the partitions become active again.

Both requests are non-blocking. Each is acknowledged as soon as the cluster accepts the change and returns the changeId of the cluster configuration change that carries it out.

Prerequisites​

The Restore API requires the following:

PrerequisiteDescription
Camunda versionCamunda 8.10 or later, restored with the exact version the backup was created with.
Backup storeEvery broker is configured with the same backup store that holds the backup, as described in the RDBMS backup prerequisites.
Completed backupA completed backup exists for every partition. List the available backups with list runtime backups.
Partition countThe partition count of the cluster matches the partition count of the backup. Brokers can be scaled between backup and restore as long as the partition count is unchanged.
API accessAuthenticated access to the Orchestration Cluster REST API. See authentication.
AuthorizationsIf authorizations are enabled, the caller needs the RESTORE permission on the BACKUP resource.

Restoring an RDBMS-backed cluster​

The examples below use the following variables:

export ORCHESTRATION_CLUSTER_API=http://localhost:8080/v2
export ORCHESTRATION_CLUSTER_MANAGEMENT_API=http://localhost:9600

Before you start, be aware of the following. Entering recovery mode stops all processing in the cluster, so plan the restore as a downtime window. The Restore API then deletes the local partition data on every broker before it writes the data from the backup, and this cannot be undone. Run the Restore API only against a cluster whose current primary storage data you intend to replace.

1. Switch the cluster into recovery mode​

Change the cluster mode to RECOVERING:

curl -X PATCH "${ORCHESTRATION_CLUSTER_API}/mode?mode=RECOVERING"

The response returns the ID of the cluster change and the operations it will apply. The plan contains one ModeChangeOperation and one AwaitModeChangeOperation per broker:

Example response
{
"changeId": "7",
"plannedChanges": [
{
"physicalTenantId": "default",
"operations": [
{ "operation": "ModeChangeOperation", "mode": "RECOVERING" },
{ "operation": "AwaitModeChangeOperation", "mode": "RECOVERING" }
]
}
]
}

Wait until this change has completed before you trigger the restore. A restore request is only accepted while every broker of the cluster is in recovery mode. Requests sent earlier are rejected with 409. Verify that all brokers are in recovery mode using either the Cluster API or the Management API.

You can use the topology API to verify the state for all partitions and brokers:

curl "${ORCHESTRATION_CLUSTER_API}/topology"

The response shows the current state of all brokers and partitions. In recovery mode, every partition should have the state recovering.

Example response
{
"brokers": [
{
"nodeId": 0,
"brokerId": "0",
"host": "192.168.1.51",
"port": 26501,
"partitions": [
{
"partitionId": 1,
"role": "inactive",
"health": "healthy",
"state": "recovering"
},
{
"partitionId": 2,
"role": "inactive",
"health": "healthy",
"state": "recovering"
}
]
}
]
}

2. Restore RDBMS​

With the cluster in recovery mode, nothing is exported to secondary storage, so restore the RDBMS now. Skip this step if the RDBMS was already restored another way, for example as part of a wider disaster recovery procedure.

Restore the RDBMS to the point in time you intend to restore the primary storage to. Camunda aligns the Zeebe and RDBMS restore points automatically.

Complete this step before you trigger the Zeebe restore. The Restore API switches the brokers back to PROCESSING as soon as the last partition is restored, and processing then resumes against whatever secondary storage is in place.

3. Trigger the restore​

There are multiple ways to define your restore point objective depending on your backup strategy. Requests that combine backupIds with from or to, that specify a time range without continuous backups enabled, or that reference a backup with no completed state in the store, are rejected with 400.

Provide the restore parameters. Camunda validates the request, resolves the backups for every partition, and acknowledges the request with 202 before the restore itself runs:

Camunda resolves the best available restore point automatically for each partition when you omit the request body.

curl -X POST "${ORCHESTRATION_CLUSTER_API}/restore"

The response returns the changeId of the restore, along with the planned operations. The plan drops and restores every partition of every broker, switches all brokers back to PROCESSING, and ends with an incarnation number update:

Example response
{
"changeId": "8",
"plannedChanges": [
{
"physicalTenantId": "default",
"operations": [
{
"operation": "SchemaInitializationOperation",
"brokerId": "0"
},
{
"operation": "PartitionPreRestoreOperation",
"brokerId": "0",
"partitionId": 1
},
{
"operation": "PartitionRestoreOperation",
"brokerId": "0",
"partitionId": 1,
"backupIds": [1748937221]
},
{
"operation": "ModeChangeOperation",
"brokerId": "0",
"mode": "PROCESSING"
},
{
"operation": "AwaitModeChangeOperation",
"brokerId": "0",
"mode": "PROCESSING"
},
{ "operation": "UpdateIncarnationNumberOperation", "brokerId": "0" }
]
}
]
}

4. Track the Restore API operation​

While a restore is in flight, query the restore status to track progress per broker and per partition:

curl "${ORCHESTRATION_CLUSTER_API}/restore"
Example response
{
"status": "IN_PROGRESS",
"changeId": "8",
"startedAt": "2026-01-01T10:00:00Z",
"brokers": [
{
"brokerId": "1",
"partitionsRestored": 1,
"partitionsToRestore": 2,
"partitions": [
{
"partitionId": 1,
"state": "RESTORED",
"backupIds": [1748937221],
"completedAt": "2026-01-01T10:02:00Z"
},
{
"partitionId": 2,
"state": "RESTORING",
"backupIds": [1748937221],
"completedAt": null
}
]
}
]
}

The overall status reports the state of the cluster change that performs the restore:

StatusMeaning
IN_PROGRESSThe restore is running.
COMPLETEDEvery partition was restored and the brokers returned to processing.
FAILEDThe restore change failed and did not complete.
CANCELLEDThe restore change was canceled.

Each partition entry reports the progress of a single broker's copy of that partition:

StateMeaning
PENDINGThe partition is queued and its restore has not started yet.
RESTORINGThe partition is being restored from its backups.
RESTOREDThe partition was restored and validated, and completedAt is set.

At most one restore is in flight at any time. Once the restore has finished, this endpoint returns 404 and the per-partition detail is no longer retained, so use the cluster monitoring API to confirm that the restore's changeId completed.

5. Confirm the cluster state after Restore API recovery​

The cluster leaves recovery mode as part of the restore, so no further action is required. Use either the topology API or the cluster monitoring API to verify that the state of all partitions and brokers has returned to normal operational status after the restore.

curl "${ORCHESTRATION_CLUSTER_API}/topology"

Restoring a cluster with multiple Physical Tenants​

Self-Managed only

In a cluster running multiple Physical Tenants, the /v2/mode and /v2/restore endpoints used above are scoped to whichever Physical Tenant your credentials belong to. There is no way to target a different tenant from these self-service endpoints, because the tenant is resolved from the caller's identity, not from the request path.

To restore a specific tenant other than your own, or every tenant at once, use the cluster-wide endpoints under /cluster/v2/.... These require cluster admin access instead of an Orchestration Cluster user's credentials:

StepTenant-scoped (your own tenant)Cluster-wide (cluster admin)
Recovery modePATCH /v2/modePATCH /cluster/v2/mode
TriggerPOST /v2/restorePOST /cluster/v2/restore
TrackGET /v2/restoreNo cluster-wide status endpoint exists. Check each tenant's own restore status, or confirm recovery through cluster-wide topology below.
ConfirmGET /v2/topologyGET /cluster/v2/topology

Choosing the Restore API scope​

Use a tenant-scoped restore when one Physical Tenant has corrupted or missing data and the other tenants should keep processing. Use a cluster-wide restore when several tenants need recovery, or when the whole cluster must be returned to a coordinated state.

The cluster-wide endpoints accept an optional physicalTenantId query parameter. Naming a tenant restores only that tenant; omitting the parameter restores every configured tenant.

export CLUSTER_ADMIN_API=http://localhost:8080/cluster/v2

curl -X POST "${CLUSTER_ADMIN_API}/restore" \
-H 'Content-Type: application/json' \
-d '{ "backupIds": [1748937221] }'

To restore tenants from different backups in a single request, supply per-tenant restore arguments in the overrides field of the request body. A request that both names a single tenant and supplies overrides is rejected, because the two express conflicting targets.

Cross-tenant safety​

A backup created for one Physical Tenant is not reachable from another tenant's restore. This is enforced by configuration rather than by a runtime check: every Physical Tenant must resolve to a distinct backup store location, and Camunda fails startup if two tenants resolve to the same one. See storage isolation.

Before returning a restored tenant to normal traffic, confirm through tenant-scoped topology that its partitions are healthy, that the expected process definitions, instances, variables, and history are present, and that exporting has resumed.

Validating a Restore API request without applying it​

Both endpoints accept the dryRun query parameter. With dryRun=true, the request is validated and the resulting plan is returned, but nothing is applied to the cluster. Use this to check a backup selection before the downtime window starts:

curl -X POST "${ORCHESTRATION_CLUSTER_API}/restore?dryRun=true" \
-H 'Content-Type: application/json' \
-d '{ "backupIds": [1748937221] }'

A dry run of a restore covers the same validation as the real request. It rejects invalid parameter combinations, checks that a completed backup exists for every partition, and, for an RDBMS time range or an empty request body, resolves the restore point from the backup metadata. A request that passes the dry run is accepted as a real request as long as the cluster and the backup store do not change in between.

The dry run does not report which backups it resolved. The response only contains the changeId and the planned operations, in the same shape as a real request, so the concrete backup ID per partition is not part of it. To confirm the selection, list the available backups with the Zeebe backup management API before the restore, or pass explicit backupIds instead of relying on automatic resolution.

Handling a failed Restore API operation​

If a single partition fails to restore, for example because its backup is corrupted or the backup store is temporarily unreachable, the partial data of that partition is dropped and the failed step is retried automatically with a backoff. The restore change stays pending, and the restore status keeps reporting the partition as RESTORING.

Because the retry is automatic, first try to fix the root cause instead of sending a new restore request. Once the cause is resolved, the pending change continues on its own and completes.

Automatic retries can't help if the problem is the backup itself, for example if the selected backup is corrupted or turns out to be the wrong restore point. In that case, retry from the outside:

  1. Cancel the pending restore change on the management API, using the changeId the restore returned:

    curl -X DELETE "${ORCHESTRATION_CLUSTER_MANAGEMENT_API}/actuator/cluster/changes/8"

    The restore status reports the change as CANCELLED, and the cluster stays in recovery mode.

  2. Send a new restore request. Because each restore drops the local partition data before it writes the backup data, the new attempt does not build on the partial result of the canceled one, and you can select a different backup target.

warning

Don't leave a partially failed restore unfinished. Between canceling a restore and completing a new one, Zeebe's internal data is a mix of restored and pre-restore state and cannot be trusted. Keep the cluster in recovery mode and retry until every partition reaches RESTORED. If you switch the cluster back to PROCESSING in that state, treat it as unrecoverable and restore again from a clean state.

(Optional) Restoring Optimize data​

If you previously backed up Optimize data, restore it independently using the standalone Optimize restore procedure. Optimize can be restored while the Orchestration Cluster restore is in progress or after it completes; the restore procedures are independent.

See back up and restore Optimize independently for the complete procedure.

(Optional) Restoring Camunda Hub data​

If you previously backed up Camunda Hub data, restore it using the same database tools.

See back up and restore Camunda Hub data for the complete procedure.