For the complete documentation index, see llms.txt.
Skip to main content
Version: 8.10 (unreleased)

Multi-region resilience

Learn about multi-region deployment and choose the right strategy for your recovery and resilience needs.

About​

Camunda provides a structured multi-region resilience framework for Self-Managed Orchestration Cluster deployments.

Comparison of Cold Recovery, Dual-Region, and three-region active-active architectures with shared RDBMS secondary storage
  • Cold Recovery: Camunda's lowest-cost multi-region configuration uses scheduled cross-region backups and a manual restore procedure to recover from complete primary-region loss. Recovery measured in hours is operationally acceptable.

  • Dual-Region: Dual-region deployment with continuous replication. A full Camunda Orchestration Cluster runs continuously in both a primary and secondary region.

  • Multi-Region RDBMS: One Orchestration Cluster runs active-active across two or more regions, and survives a region loss with three or more. A relational database (RDBMS) with cross-region replication holds the secondary storage. Losing one region preserves the cluster quorum. You must fail over the database writer if the lost region held the writer.

Get started: choose your strategy​

Choosing the right recovery strategy is determined by how critical your process automation is to your business. How much downtime and data loss can you tolerate, and what compliance obligations do you have?

First, determine how critical your workload is:

If your business can accept the following outcome:Choose this option
Recovery measured in hours, and minutes to hours of data loss.Cold Recovery
Recovery in ~15 minutes, with no data loss, and audit-ready posture.Dual-Region
Processing continues after a single region loss, and you can run without Optimize.Multi-Region RDBMS

What each strategy asks of you:

  • Cold Recovery is a manual procedure built on the backup and restore guide. There is no reference architecture. Validate the procedure in your own environment.
  • Dual-Region includes a reference architecture and an operational runbook, with documented Recovery Time Objective (RTO) and Recovery Point Objective (RPO) targets. A region loss stops processing until an operator runs the failover.
  • Multi-Region RDBMS, with three or more regions, keeps processing through a region loss with no Zeebe operator step. The database handles secondary-storage replication. In exchange, it costs a third region of capacity. Optimize is unavailable, because Optimize requires Elasticsearch or OpenSearch instead of a relational secondary storage.

Comparison of multi-region resilience​

The following table provides a detailed comparison of the available multi-region deployment options:

ConsiderationCold RecoveryDual-Region (Elasticsearch)Multi-Region RDBMS
RegionsOne, plus cross-region object storageExactly twoTwo or more. Three or more to keep processing through a region loss
Recovery time (RTO)~1–4 hours~15 minutesWith three or more regions, no Zeebe recovery procedure. The window is a Raft re-election plus your own client and traffic failover, plus the database writer promotion if the lost region held the writer
Data loss (RPO)15 min – 4 hours (backup-interval dependent)0 minutes0 minutes for engine state, and for secondary storage once replay completes under the documented replication monitoring
Failover modeManual, operator-initiatedManual, operator-initiatedWith three or more regions, automatic for Zeebe. Manual for the database writer, and only if it was in the lost region
Secondary storageElasticsearch or OpenSearch, restored from backupElasticsearch, one cluster per regionRDBMS, one database replicated by the database itself
ArchitectureScheduled backup to cross-region object storage. Manual restore into a secondary regionOrchestration Cluster running in both regions. Dual-region exporters. Manual failoverOne Orchestration Cluster across every region. One exporter. Zone-aware partition placement, where this reference layout maps one zone to one region. A zone can also be an availability zone
Typical use caseLow-criticality production; environments where hours-long recovery is acceptableEnterprise production workloads that must survive a region failureWorkloads that cannot pause while an operator runs a failover
OptimizeSupportedSupportedNot available, Optimize requires Elasticsearch or OpenSearch
Compliance fitBasic business continuity management (BCM) requirementsCertified, auditable region-recovery posture with a published runbookAuditable region-recovery posture with a published runbook, for organizations that also require continuous processing
Relative cost$ (lower cost): Object storage only; no standing second region$$$ (higher cost): Orchestration Cluster running across both regions with extra capacity to sustain load in case of Region failure, plus cross-region traffic$$$$ (highest cost): Orchestration Cluster running across three or more regions, plus cross-region traffic
important

Cold Recovery RTO and RPO targets are bounded by data volume, backup frequency, and operator restore speed. Treat published ranges as planning targets, not contractual commitments.

Dual-Region RTO is based on internal operational tests. Actual times may vary depending on your environment, level of automation and the specific manual steps performed during recovery. See Dual-Region for a phase-by-phase breakdown.

With three or more regions, Multi-Region RDBMS removes the recovery procedure, not the recovery window. A published RTO figure would not be meaningful here. Most of the elapsed time comes from your client timeouts, your traffic routing, and your database failover, not from the architecture. Measure the actual recovery window, including client reconnection, with a real failover test in your environment.

Multi-Region RDBMS reaches RPO 0 for engine state, and for secondary storage only under the replication conditions you configure. See recovery objectives.