Restore a backup with the Restore API
Restore Zeebe partition data through the Orchestration Cluster Restore API without restarting the brokers, when using Elasticsearch or OpenSearch as secondary storage.
This page is part of the Elasticsearch/OpenSearch restore procedure. To compare it with the legacy Restore Application, see choosing a restore approach.
About the Restore API
With Camunda 8.10 and later, a Restore API recovery runs during a downtime window while the cluster is in recovery mode. It runs in four phases, driven by two API requests:
- Entering recovery mode: every broker deactivates its partitions and switches to a restricted partition manager. While the cluster is in recovery mode it processes no work, and only read-only operations and restore remain available.
- Restoring secondary storage: while the cluster is in recovery mode, restore the Elasticsearch/OpenSearch snapshots for the intended backup ID.
- Restoring the partitions: the cluster plans a single change that, for every broker and partition, first drops the local partition data and then restores that partition from the selected backup. The steps of that plan run one at a time across the cluster.
- Returning to processing: once every partition is restored, the same change switches all brokers back to
PROCESSINGand the partitions become active again.
Both requests are non-blocking. Each is acknowledged as soon as the cluster accepts the change and returns the changeId of the cluster configuration change that carries it out.
Prerequisites
In addition to the general restore prerequisites, the Restore API requires the following:
| Prerequisite | Description |
|---|---|
| Camunda version | Camunda 8.10 or later, restored with the exact version the backup was created with. |
| Backup store | Every broker is configured with the same backup store that holds the Zeebe backup, and Elasticsearch/OpenSearch is configured with the same snapshot repository as the backup. See prerequisites. |
| Sizing | Elasticsearch/OpenSearch should be sized the same or larger than the original cluster; a smaller cluster can prevent shards from being assigned and fail the restore. |
| Optimize stopped | Optimize must be stopped before you restore the Elasticsearch/OpenSearch snapshots in step 3; every other component keeps running in recovery mode. |
| Partition count | The partition count of the cluster matches the partition count of the backup. Brokers can be scaled between backup and restore as long as the partition count is unchanged. |
| API access | Authenticated access to the Orchestration Cluster REST API. See authentication. |
| Authorizations | If authorizations are enabled, the caller needs the RESTORE permission on the BACKUP resource. |
Restoring an Elasticsearch/OpenSearch-backed cluster
The examples below use the following variables:
export ORCHESTRATION_CLUSTER_API=http://localhost:8080/v2
export ORCHESTRATION_CLUSTER_MANAGEMENT_API=http://localhost:9600
Before you start, be aware of the following. Entering recovery mode stops all processing in the cluster, so plan the restore as a downtime window. The Restore API then deletes the local partition data on every broker before it writes the data from the backup, and this cannot be undone. Run the Restore API only against a cluster whose current primary storage data you intend to replace.
1. Switch the cluster into recovery mode
Change the cluster mode to RECOVERING:
curl -X PATCH "${ORCHESTRATION_CLUSTER_API}/mode?mode=RECOVERING"
Wait until the mode change has completed. Query the cluster monitoring API and check that lastChange.id matches the returned changeId and that no pendingChange is reported:
curl "${ORCHESTRATION_CLUSTER_MANAGEMENT_API}/actuator/cluster"
A Restore API request is only accepted while every broker of the cluster is in recovery mode. Requests sent earlier are rejected with 409.
2. Find available backup IDs
With the cluster in recovery mode, use the Orchestration Cluster REST API to list the available runtime and history backups for the current Physical Tenant. Both endpoints require the BACKUP:READ permission. Use the returned backup ID to select the matching Elasticsearch/OpenSearch snapshots and Zeebe primary storage backup.
- Runtime backups
- History backups
Use list runtime backups to list available Zeebe primary storage backups. Omit prefix to list all backups, or use a numeric prefix followed by * to narrow the results.
curl "${ORCHESTRATION_CLUSTER_API}/backups/runtime"
To list backups matching a prefix:
curl "${ORCHESTRATION_CLUSTER_API}/backups/runtime?prefix=1748937*"
Use list history backups to list available Operate, Tasklist, and Optimize history backups. This endpoint is available because Elasticsearch/OpenSearch is the secondary storage. Use verbose=false when snapshot-level details are not needed.
curl "${ORCHESTRATION_CLUSTER_API}/backups/history"
To list backups matching a prefix without snapshot-level details:
curl "${ORCHESTRATION_CLUSTER_API}/backups/history?prefix=1748937*&verbose=false"
The runtime and history listings are scoped to the Physical Tenant associated with the caller's credentials. For other tenants or all tenants, use the cluster-admin endpoints described in restoring a cluster with multiple Physical Tenants. Ensure that the runtime and history backups you select use the same backup ID before continuing.
3. Restore Elasticsearch/OpenSearch snapshots
With the cluster in recovery mode, restore the Elasticsearch/OpenSearch snapshots to the intended point in time, using the same backup ID that you pass to the Restore API. A mismatched backup ID produces an inconsistent restore point.
Keep the Orchestration Cluster running in recovery mode while you restore the snapshots.
1. Restore templates
This step is only required for restoring an Elasticsearch/OpenSearch snapshot on a fresh cluster.
This step includes restoring index and component templates crucial for Camunda 8 to function properly on continuous use.
These templates are automatically applied on newly created indices. These templates are only created on the initial start of the components and the first seeding of the secondary datastore, due to which you have to temporarily restore them before you can restore all Elasticsearch/OpenSearch snapshots.
Start Camunda 8 configured with your secondary datastore endpoint:
- For example, deploy the Camunda Helm chart.
- For manual context, start Camunda 8 components manually.
- Depending on your setup this can mean Orchestration Cluster (Operate, Tasklist, Zeebe), Optimize and the required secondary datastore.
The templates are created by the Web Applications (Operate, Tasklist), and Optimize on startup on the first seeding of the datastore. Zeebe creates this whenever it is required, and isn't limited to the initial start. We recommend starting your full required Camunda 8 stack for the applications to show up as healthy.
You can confirm the successful creation of the index templates by using the Elasticsearch/OpenSearch API. The index templates rely on the component templates, so it also confirms these were successfully recreated.
- Elasticsearch
- OpenSearch
The following uses the Elasticsearch Index API to list all index templates.
curl -s "$ELASTIC_ENDPOINT/_index_template" \
| jq -r '.index_templates[].name' \
| grep -E 'operate|tasklist|optimize|zeebe' \
| sort
Example Output
operate-batch-operation-1.0.0_template
operate-decision-instance-8.3.0_template
operate-event-8.3.0_template
operate-flownode-instance-8.3.1_template
operate-incident-8.3.1_template
operate-job-8.6.0_template
operate-list-view-8.3.0_template
operate-message-8.5.0_template
operate-operation-8.4.1_template
operate-post-importer-queue-8.3.0_template
operate-sequence-flow-8.3.0_template
operate-variable-8.3.0_template
tasklist-draft-task-variable-8.3.0_template
tasklist-task-8.5.0_template
tasklist-task-variable-8.3.0_template
...
The following uses the OpenSearch Index API to list all index templates.
curl -s "$OPENSEARCH_ENDPOINT/_index_template" \
| jq -r '.index_templates[].name' \
| grep -E 'operate|tasklist|optimize|zeebe' \
| sort
Example Output
operate-batch-operation-1.0.0_template
operate-decision-instance-8.3.0_template
operate-event-8.3.0_template
operate-flownode-instance-8.3.1_template
operate-incident-8.3.1_template
operate-job-8.6.0_template
operate-list-view-8.3.0_template
operate-message-8.5.0_template
operate-operation-8.4.1_template
operate-post-importer-queue-8.3.0_template
operate-sequence-flow-8.3.0_template
operate-user-task-8.5.0_template
operate-variable-8.3.0_template
tasklist-draft-task-variable-8.3.0_template
tasklist-task-8.5.0_template
tasklist-task-variable-8.3.0_template
...
2. Stop Optimize
The Restore API keeps the brokers running in recovery mode and the web applications up, so only Optimize needs to be stopped before you restore the Elasticsearch/OpenSearch snapshots.
If you are using the Camunda Helm chart, disable Optimize in the values.yml:
optimize:
enabled: false
3. Delete all indices
Now that you have successfully restored the templates and stopped the components adding more indices, you must delete the existing indices to be able to successfully restore the snapshots (otherwise these will block a successful restore).
If multiple physical tenants are configured, make sure to delete the indices corresponding to the required tenant by specifying the proper prefix.
- Elasticsearch
- OpenSearch
The following uses the Elasticsearch CAT API to list all indices. It also uses the Elasticsearch Index API to delete an index.
for index in $(curl -s "$ELASTIC_ENDPOINT/_cat/indices?h=index" \
| grep -E 'camunda|operate|tasklist|optimize|zeebe'); do
echo "Deleting index: $index"
curl -X DELETE "$ELASTIC_ENDPOINT/$index"
done
Example Output
Deleting index: operate-import-position-8.3.0_
{"acknowledged":true}Deleting index: operate-migration-steps-repository-1.1.0_
{"acknowledged":true}Deleting index: operate-flownode-instance-8.3.1_
{"acknowledged":true}Deleting index: operate-event-8.3.0_
{"acknowledged":true}Deleting index: operate-incident-8.3.1_
{"acknowledged":true}Deleting index: tasklist-web-session-1.1.0_
{"acknowledged":true}Deleting index: tasklist-variable-8.3.0_
{"acknowledged":true}Deleting index: operate-user-task-8.5.0_
{"acknowledged":true}Deleting index: tasklist-import-position-8.2.0_
{"acknowledged":true}Deleting index: tasklist-task-variable-8.3.0_
{"acknowledged":true}Deleting index: tasklist-flownode-instance-8.3.0_
{"acknowledged":true}Deleting index: operate-process-8.3.0_
{"acknowledged":true}Deleting index: tasklist-process-instance-8.3.0_
{"acknowledged":true}Deleting index: operate-operation-8.4.1_
{"acknowledged":true}Deleting index: operate-job-8.6.0_
{"acknowledged":true}Deleting index: operate-metric-8.3.0_
{"acknowledged":true}Deleting index: tasklist-migration-steps-repository-1.1.0_
{"acknowledged":true}Deleting index: operate-decision-8.3.0_
{"acknowledged":true}Deleting index: tasklist-process-8.4.0_
{"acknowledged":true}Deleting index: operate-variable-8.3.0_
{"acknowledged":true}Deleting index: operate-message-8.5.0_
{"acknowledged":true}Deleting index: operate-decision-requirements-8.3.0_
{"acknowledged":true}Deleting index: operate-batch-operation-1.0.0_
{"acknowledged":true}Deleting index: operate-web-session-1.1.0_
{"acknowledged":true}Deleting index: tasklist-user-1.4.0_
{"acknowledged":true}Deleting index: operate-list-view-8.3.0_
{"acknowledged":true}Deleting index: tasklist-metric-8.3.0_
{"acknowledged":true}Deleting index: operate-post-importer-queue-8.3.0_
{"acknowledged":true}Deleting index: tasklist-task-8.5.0_
{"acknowledged":true}Deleting index: tasklist-form-8.4.0_
{"acknowledged":true}Deleting index: operate-user-1.2.0_
{"acknowledged":true}Deleting index: tasklist-draft-task-variable-8.3.0_
{"acknowledged":true}Deleting index: operate-decision-instance-8.3.0_
{"acknowledged":true}Deleting index: operate-sequence-flow-8.3.0_
{"acknowledged":true}
The following uses the OpenSearch CAT API to list all indices. It also uses the OpenSearch Index API to delete an index.
for index in $(curl -s "$OPENSEARCH_ENDPOINT/_cat/indices?h=index" \
| grep -E 'camunda|operate|tasklist|optimize|zeebe'); do
echo "Deleting index: $index"
curl -X DELETE "$OPENSEARCH_ENDPOINT/$index"
done
Example Output
Deleting index: operate-import-position-8.3.0_
{"acknowledged":true}Deleting index: operate-migration-steps-repository-1.1.0_
{"acknowledged":true}Deleting index: operate-flownode-instance-8.3.1_
{"acknowledged":true}Deleting index: operate-event-8.3.0_
{"acknowledged":true}Deleting index: operate-incident-8.3.1_
{"acknowledged":true}Deleting index: tasklist-web-session-1.1.0_
{"acknowledged":true}Deleting index: tasklist-variable-8.3.0_
{"acknowledged":true}Deleting index: operate-user-task-8.5.0_
{"acknowledged":true}Deleting index: tasklist-import-position-8.2.0_
{"acknowledged":true}Deleting index: tasklist-task-variable-8.3.0_
{"acknowledged":true}Deleting index: tasklist-flownode-instance-8.3.0_
{"acknowledged":true}Deleting index: operate-process-8.3.0_
{"acknowledged":true}Deleting index: tasklist-process-instance-8.3.0_
{"acknowledged":true}Deleting index: operate-operation-8.4.1_
{"acknowledged":true}Deleting index: operate-job-8.6.0_
{"acknowledged":true}Deleting index: operate-metric-8.3.0_
{"acknowledged":true}Deleting index: tasklist-migration-steps-repository-1.1.0_
{"acknowledged":true}Deleting index: operate-decision-8.3.0_
{"acknowledged":true}Deleting index: tasklist-process-8.4.0_
{"acknowledged":true}Deleting index: operate-variable-8.3.0_
{"acknowledged":true}Deleting index: operate-message-8.5.0_
{"acknowledged":true}Deleting index: operate-decision-requirements-8.3.0_
{"acknowledged":true}Deleting index: operate-batch-operation-1.0.0_
{"acknowledged":true}Deleting index: operate-web-session-1.1.0_
{"acknowledged":true}Deleting index: tasklist-user-1.4.0_
{"acknowledged":true}Deleting index: operate-list-view-8.3.0_
{"acknowledged":true}Deleting index: tasklist-metric-8.3.0_
{"acknowledged":true}Deleting index: operate-post-importer-queue-8.3.0_
{"acknowledged":true}Deleting index: tasklist-task-8.5.0_
{"acknowledged":true}Deleting index: tasklist-form-8.4.0_
{"acknowledged":true}Deleting index: operate-user-1.2.0_
{"acknowledged":true}Deleting index: tasklist-draft-task-variable-8.3.0_
{"acknowledged":true}Deleting index: operate-decision-instance-8.3.0_
{"acknowledged":true}Deleting index: operate-sequence-flow-8.3.0_
{"acknowledged":true}
4. Restore the snapshots
Although the backup order was important so far to ensure consistent backups, you can restore the backed up indices in any order.
As the components do not have an endpoint to restore the backup in Elasticsearch, you will need to restore it yourself directly in your selected datastore.
Using your chosen backup ID from the previous step, restore the snapshots in Elasticsearch/OpenSearch for each available backup under the same backup ID.
- Elasticsearch
- OpenSearch
The following uses the Elasticsearch snapshot API to restore a snapshot.
curl -XPOST "$ELASTIC_ENDPOINT/_snapshot/$ELASTIC_SNAPSHOT_REPOSITORY/$SNAPSHOT_NAME/_restore?wait_for_completion=true"
The following uses the OpenSearch snapshot API to restore a snapshot.
curl -XPOST "$OPENSEARCH_ENDPOINT/_snapshot/$OPENSEARCH_SNAPSHOT_REPOSITORY/$SNAPSHOT_NAME/_restore?wait_for_completion=true"
Where $SNAPSHOT_NAME would be any of the following, based on the backup ID you found earlier:
camunda_optimize_1748937221_8.8.0_part_1_of_2
camunda_optimize_1748937221_8.8.0_part_2_of_2
camunda_webapps_1748937221_8.8.0_part_1_of_5
camunda_webapps_1748937221_8.8.0_part_2_of_5
camunda_webapps_1748937221_8.8.0_part_3_of_5
camunda_webapps_1748937221_8.8.0_part_4_of_5
camunda_webapps_1748937221_8.8.0_part_5_of_5
camunda_zeebe_records_backup_1748937221
Ensure that all your backups correspond to the same backup ID and that each one is restored one-by-one.
Complete this step before you trigger the Zeebe restore. The Restore API switches the brokers back to PROCESSING as soon as the last partition is restored, and processing then resumes against whatever secondary storage is in place.
4. Trigger the restore
Provide the restore parameters. Camunda validates the request, resolves the backups for every partition, and acknowledges the request with 202 before the restore itself runs:
curl -X POST "${ORCHESTRATION_CLUSTER_API}/restore" \
-H 'Content-Type: application/json' \
-d '{ "backupIds": [1748937221] }'
The response returns the changeId of the restore, along with the planned operations. The plan drops and restores every partition of every broker, switches all brokers back to PROCESSING, and ends with an incarnation number update:
Example response
{
"changeId": "8",
"plannedChanges": [
{
"physicalTenantId": "default",
"operations": [
{
"operation": "SchemaInitializationOperation",
"brokerId": "0"
},
{
"operation": "PartitionPreRestoreOperation",
"brokerId": "0",
"partitionId": 1
},
{
"operation": "PartitionRestoreOperation",
"brokerId": "0",
"partitionId": 1,
"backupIds": [1748937221]
},
{
"operation": "ModeChangeOperation",
"brokerId": "0",
"mode": "PROCESSING"
},
{
"operation": "AwaitModeChangeOperation",
"brokerId": "0",
"mode": "PROCESSING"
},
{ "operation": "UpdateIncarnationNumberOperation", "brokerId": "0" }
]
}
]
}
5. Track the Restore API operation
While a restore is in flight, query the restore status to track progress per broker and per partition:
curl "${ORCHESTRATION_CLUSTER_API}/restore"
Example response
{
"status": "IN_PROGRESS",
"changeId": "8",
"startedAt": "2026-01-01T10:00:00Z",
"brokers": [
{
"brokerId": "1",
"partitionsRestored": 1,
"partitionsToRestore": 2,
"partitions": [
{
"partitionId": 1,
"state": "RESTORED",
"backupIds": [1748937221],
"completedAt": "2026-01-01T10:02:00Z"
},
{
"partitionId": 2,
"state": "RESTORING",
"backupIds": [1748937221],
"completedAt": null
}
]
}
]
}
The overall status reports the state of the cluster change that performs the restore:
| Status | Meaning |
|---|---|
IN_PROGRESS | The restore is running. |
COMPLETED | Every partition was restored and the brokers returned to processing. |
FAILED | The restore change failed and did not complete. |
CANCELLED | The restore change was canceled. |
Each partition entry reports the progress of a single broker's copy of that partition:
| State | Meaning |
|---|---|
PENDING | The partition is queued and its restore has not started yet. |
RESTORING | The partition is being restored from its backups. |
RESTORED | The partition was restored and validated, and completedAt is set. |
At most one restore is in flight at any time. Once the restore has finished, this endpoint returns 404 and the per-partition detail is no longer retained, so use the cluster monitoring API to confirm that the restore's changeId completed.
6. Confirm the cluster state after Restore API recovery
Check that every partition is active and healthy again using the topology:
curl "${ORCHESTRATION_CLUSTER_API}/topology"
The cluster leaves recovery mode as part of the restore, so no further action is required.
Restoring a cluster with multiple Physical Tenants
Self-Managed onlyFor multiple Physical Tenants, tenant-scoped Restore API calls target the Physical Tenant associated with the caller's credentials. To restore another tenant or all tenants, use the cluster-wide endpoints with cluster admin access.
| Step | Tenant-scoped (your own tenant) | Cluster-wide (cluster admin) |
|---|---|---|
| Recovery mode | PATCH /v2/mode | PATCH /cluster/v2/mode |
| Trigger | POST /v2/restore | POST /cluster/v2/restore |
| Track | GET /v2/restore | No cluster-wide status endpoint exists. Check each tenant's own restore status, or confirm recovery through cluster-wide topology below. |
| Confirm | GET /v2/topology | GET /cluster/v2/topology |
Use a tenant-scoped restore when one Physical Tenant has corrupted or missing data and the other tenants should keep processing. Use a cluster-wide restore when several tenants need recovery, or when the whole cluster must be returned to a coordinated state.
The cluster-wide endpoints accept an optional physicalTenantId query parameter. Naming a tenant restores only that tenant; omitting the parameter restores every configured tenant. Each tenant must have its own non-overlapping backup location, and the same backup ID must refer to compatible snapshots for every tenant included in the restore.
export CLUSTER_ADMIN_API=http://localhost:8080/cluster/v2
curl -X POST "${CLUSTER_ADMIN_API}/restore" \
-H 'Content-Type: application/json' \
-d '{ "backupIds": [1748937221] }'
Before returning a restored tenant to normal traffic, confirm through tenant-scoped topology that its partitions are healthy, that the expected data is present, and that exporting has resumed.
Validating a Restore API request without applying it
The Restore API accepts the dryRun query parameter. With dryRun=true, the request is validated and the resulting plan is returned, but nothing is applied to the cluster. Use this to check a backup selection before the downtime window starts:
curl -X POST "${ORCHESTRATION_CLUSTER_API}/restore?dryRun=true" \
-H 'Content-Type: application/json' \
-d '{ "backupIds": [1748937221] }'
A dry run rejects requests without a backup ID, with multiple backup IDs, or with a time range. It also checks that a completed backup exists for every partition. A request that passes the dry run is accepted as a real request as long as the cluster and the backup store do not change in between.
Handling a failed Restore API operation
If a single partition fails to restore, for example because its backup is corrupted or the backup store is temporarily unreachable, the partial data of that partition is dropped and the failed step is retried automatically with a backoff. The restore change stays pending, and the restore status keeps reporting the partition as RESTORING.
Because the retry is automatic, first try to fix the root cause instead of sending a new restore request. Once the cause is resolved, the pending change continues on its own and completes.
Automatic retries can't help if the problem is the backup itself, for example if the selected backup is corrupted or turns out to be the wrong restore point. In that case, retry from the outside:
-
Cancel the pending restore change on the management API, using the
changeIdthe restore returned:curl -X DELETE "${ORCHESTRATION_CLUSTER_MANAGEMENT_API}/actuator/cluster/changes/8"The restore status reports the change as
CANCELLED, and the cluster stays in recovery mode. -
Send a new restore request. Because each restore drops the local partition data before it writes the backup data, the new attempt does not build on the partial result of the canceled one, and you can select a different backup target.
Don't leave a partially failed restore unfinished. Between canceling a restore and completing a new one, Zeebe's internal data is a mix of restored and pre-restore state and cannot be trusted. Keep the cluster in recovery mode and retry until every partition reaches RESTORED. If you switch the cluster back to PROCESSING in that state, treat it as unrecoverable and restore again from a clean state.
(Optional) Restoring Camunda Hub data
If you previously backed up Camunda Hub data, restore it using the same database tools.
See back up and restore Camunda Hub data for the complete procedure.