For the complete documentation index, see llms.txt.
Skip to main content
Version: 8.10 (unreleased)

Multi-Region RDBMS operational procedure

This runbook covers the day-2 operations of a Multi-Region RDBMS setup: losing a region, bringing it back, and adding a region.

caution

Develop, test, and rehearse these procedures in a non-production environment before you need them. The commands below are examples from the reference implementation. Adapt them to your environment.

What is different from dual-region​

In a dual-region setup, losing a region costs the Zeebe quorum. Processing stops, and the failover procedure exists to restore it. That procedure removes the lost brokers, disables the exporter to the lost region, and later restores secondary storage from a snapshot.

With three or more zones and no zone holding half the replicas or more, none of that applies. Every partition keeps a majority of its replicas. Zeebe keeps processing, and you need no Zeebe action to restore service. The dry run confirms this before you act. The failover procedure mostly reports. It only acts on the database writer, and only when the writer was in the lost region.

Side-by-side timelines of the same zone loss. A triangle marks an incident, a play icon an operator action, a check a healthy state, and a dash a step that does not exist. In a Dual-Region cluster with Elasticsearch, Zeebe loses quorum and processing stops until an operator force-removes the lost brokers and disables the exporter. Failback also requires a secondary storage snapshot and restore, for four operator steps in total. In a three-zone Multi-Region RDBMS cluster, quorum holds and processing continues. Three operator steps remain: promoting the database writer if it was in the lost zone, removing the lost zone, which is recommended but not needed for quorum, and redeploying the zone at failback.What a region loss costs youThe same event, in a two-zone and in a three-zone cluster. Time runs downward.Triangle: incident. Play: operator action. Check: healthy. Dash: no such step here.Dual-region, two zonesMulti-Region RDBMS, three zonesZone B is lostZeebe loses quorum, processing STOPSevery partition is down to 1 replica of 2Force-remove the lost brokersprocessing only resumes after thisDisable the exporter to zone BProcessing resumesFailback: redeploy zone BSnapshot and restore the secondary storagezone B has no copy of the exported dataCluster is whole againZone C is lostQuorum holds, processing CONTINUESevery partition still has 3 replicas of 5Promote the database writeronly if the writer was in zone CRemove the lost zonerecommended, not needed for quorumNothing to disablethere is one exporter, and one databaseFailback: redeploy zone CNothing to restorethe database already holds the exported dataCluster is whole again4 operator steps.Processing stays down for the first two.3 operator steps.Processing never stopped.!!!
StepDual-regionMulti-Region RDBMS
Restore processingForce-remove the lost brokersNothing, processing never stopped
Secondary storage after failoverDisable the exporter to the lost regionNothing, there is one exporter and one database
Promote the databasen/aOnly if the writer was in the lost region
Remove the lost zoneSame step as restoring processingRecommended, not needed for quorum
FailbackSnapshot and restore secondary storageRedeploy the region

The dual-region procedure takes 10 operator steps: two to fail over and eight to fail back. The diagram above counts three operator actions here: promote the writer if needed, remove the lost zone, and redeploy the region at failback. The runbook below adds confirmations around them, for five steps in total.

Use this runbook only for Multi-Region RDBMS

This runbook applies only to a zone-aware cluster with RDBMS secondary storage. Its region-loss procedures assume three or more zones. A cluster that starts on two zones uses only Add a region until it runs three. If it loses a zone before then, processing stops. Bring the lost zone back before you add the third region, because the add-zone change needs a quorum. Don't run the dual-region procedure on it: force-removing brokers or restoring secondary storage from a snapshot is unnecessary here and can lose data. For a two-region cluster with Elasticsearch, use the dual-region procedure instead.

Terminology​

TermMeaning
SlotA position in the region list, numbered from 0. Fixed when the cluster is bootstrapped.
ZoneThe Camunda-level name of a region, for example london. One zone per region.
Active regionA slot that is actually deployed.
WriterThe single database instance accepting writes from every region.

Prerequisites​

You need a local copy of the aws/kubernetes/eks-multi-region-rdbms reference architecture, from the camunda-deployment-references repository. It holds the Terraform modules, the Helm values, and every procedure script this documentation refers to.

The following clones the repository and changes into the architecture directory. Every command in this documentation runs from there.

aws/kubernetes/eks-multi-region-rdbms/procedure/get-your-copy.sh
loading...

The reference architecture is a starting point you own and extend, not a module you consume, so the workflow is to copy it into your own repository rather than reference it remotely.

Source the environment before running any procedure. The scripts derive everything from the Terraform state, and refuse to run against an inconsistent topology:

cd procedure
. ./export-terraform-outputs.sh
. ./export_environment_prerequisites.sh

Source them with the leading dot (. ./script.sh). These scripts export variables into your current shell, not into a subshell. For what each variable means, see prepare the environment in the deployment guide.

You also need the credentials and the CLI tools the deployment used: kubectl contexts for every active region, helm, jq, and your cloud provider's CLI. The deployment guide lists them.

Confirm the cluster is healthy before you start, so you can tell what the procedure changed:

./check-cluster-topology.sh

Handle a region loss​

1. Confirm the quorum is intact​

The surviving zones keep processing if they hold a majority of each partition's replicas. The concept page explains when this holds.

A cluster with only two zones, such as 2-2 before you add the third region, has no such margin. Losing either zone leaves two replicas of four, and processing stops.

Confirm this rather than assuming it. The script takes one lost slot and computes the surviving replicas without it. Its verdict only covers a single lost zone. If more than one zone is affected, the reference procedures don't cover the situation. Read the partition health of the surviving brokers from GET /actuator/cluster on a surviving region instead. Don't use ./check-cluster-topology.sh here: it expects every active region to be up.

./failover.sh <lost-region-slot> --dry-run

With --dry-run, the script reports the quorum state, prints the current cluster view, and warns if the surviving zones no longer hold a majority. It changes nothing, so you can read the verdict before deciding to act.

2. Promote the database writer if needed​

If the writer was in the lost region, promote a surviving member. The mode depends on whether the lost region is still reachable:

If the writer was not in the lost region, no database action is needed. If it was, a reachable region allows a planned switchover with ./failover.sh and no data loss, while a lost region needs the AWS global database recovery, where the database loses its replication lag and Zeebe replays those records from its retained log. Either way, the JDBC URL resolves to the new writer, and Camunda needs no reconfiguration and no restart. Then raise the priority of the zone with the new writer, wait for COMPLETED, and run POST /cluster/v2/rebalance.Promote the database writer after a region losswriter in thelost region?nono database actionyesplanned: switchover./failover.sh <slot>still reachable, no data lostunplanned: AWS globaldatabase recoveryDB loses its lag, Zeebe replays itJDBC URL resolvesto the new writerCamunda: no reconfiguration,no restartThen move the Raft leaders next to the new writer1. raise the priority of the zone with the new writer2. wait for COMPLETED3. POST /cluster/v2/rebalance

The region is still reachable, for example during a scheduled evacuation. A switchover completes replication before promoting, so no data is lost in the RDBMS. It also takes considerably less time than an unplanned failover.

Run the same script as in step 1, without --dry-run. It repeats the quorum report, then promotes a surviving member if the writer was in the lost region:

./failover.sh <lost-region-slot>

Camunda needs no reconfiguration and no restart, as long as the JDBC URL keeps resolving to the current writer. The reference implementation gets that from the AWS Advanced JDBC Wrapper. Its failover plugin follows the writer on established connections, and on brokers that start after the promotion. This is not general JDBC behavior. With your own database, whether connections re-resolve the writer depends on your driver and endpoint. Confirm it or plan a restart.

If the writer was not in the lost region, you need no database action.

Move the Raft leaders to the new writer region​

Once the writer moves, the zone priorities still favor the region that hosted the old one. Partition leaders keep exporting across regions and pay the inter-region round trip on every flush. Move the leaders next to the new writer:

  1. Raise the priority of the zone that now hosts the writer. See zone-aware clusters for the priority property, and the Partitioning API for applying it to a running cluster.
  2. Wait until the change reports COMPLETED. The cluster rejects a new change while one is still in progress.
  3. Run a rebalance with POST /cluster/v2/rebalance. Priorities apply at the next election and don't move existing leaders on their own.

3. Route client traffic away from the lost region​

Zeebe keeps processing, but the gateway in the lost region is unreachable. Update your DNS or load balancer to stop sending client traffic there. Traffic routing sits outside Camunda's control and depends on your own setup.

4. Remove the lost zone​

Remove the brokers of the lost zone. One atomic change evicts them. It also drops the zone from the persisted partition distribution, so quorum stops counting replicas that cannot answer:

./failover.sh <lost-region-slot> --drain-brokers

This issues DELETE /actuator/cluster/zones/<zone>?force=true against a surviving region. Without force=true, the API tries a graceful drain, which fails when the zone is down. Only do this for a zone that is down and unreachable, and for one zone at a time. See the cluster management API.

In a planned evacuation, the zone is still reachable, so don't force-remove it. Drain it gracefully instead: send DELETE /actuator/cluster/zones/<zone> without force=true through the Remove a zone API. The engine moves the zone's partitions to the remaining zones before it removes the brokers. The request is asynchronous. Wait until GET /actuator/cluster reports the change as COMPLETED before you shut down the zone's brokers.

This step applies to a cluster with three or more zones. You must remove the zone when it held half the replicas or more. The replica count decides this, not the number of zones. See step 1.

Don't remove a zone from a cluster that still runs on its two-zone bootstrap. Bring the lost zone back instead, then add the third region. An evenly split Dual-Region cluster does need a force-removal, which is why it has its own failover runbook.

The trade-off is failback cost. You must add a removed zone back when you bring the region back, and its brokers start empty.

5. Verify the degraded cluster​

./verify-degraded-cluster.sh <lost-region-slot>

The cluster should report the surviving brokers, all partitions healthy, and processing continuing.

Bring a region back​

Failback is short by design. It has no secondary storage snapshot and restore step. The database holds a single copy of the exported data and replicates it itself. A returning region has nothing to catch up on at the Camunda level.

./failback.sh <recovered-region-slot>
./failback.sh redeploys Camunda in the region and re-exports its services. If the zone was not force-removed during failover, its brokers rejoin and catch up from the Raft log with no membership change. If it was removed, the script adds the zone back with POST /actuator/cluster/zones/<zone> and waits for COMPLETED while the brokers rebuild. Then run ./check-cluster-topology.sh. The --switch-writer option moves the writer back.Bring a region back: ./failback.sh <slot>No secondary storage snapshot, no restore. The database already holds the only copy.redeploy Camundain the regionre-export itsserviceszone force-removed duringfailover?nobrokers rejoin and catchup from the Raft logno membership changeyesadd the zone backPOST /actuator/cluster/zones/<zone>wait for COMPLETED, brokers rebuildThen: ./check-cluster-topology.sh Optional: --switch-writer moves the writer back

The procedure does four things:

  1. Redeploys Camunda in the recovered region: namespace, database secret, Helm values, and chart.
  2. Re-exports the region's services to the ClusterSet, so brokers in other regions can resolve them again.
  3. Re-adds the zone if you force-removed it during failover. If you left the zone in place, its brokers rejoin and catch up from the Raft log with no membership change at all.
  4. Reports the database state, and stops if an unplanned recovery left the global topology incomplete.

Move the writer back to the recovered region if the other regions are further from the current writer:

./failback.sh <recovered-region-slot> --switch-writer

Leaving the writer where it is costs nothing but cross-region latency for the regions furthest from it.

After an unplanned failover

An unplanned recovery can leave the promoted member detached from the global database. Restore a complete Aurora Global Database topology with the AWS recovery procedure before running failback.sh. The script refuses to continue while the global cluster has only one member.

Confirm the topology when done:

./check-cluster-topology.sh

Add a region​

Adding a region to a running cluster is an online operation. The regions already running keep processing and are not restarted.

The new region's brokers start first. Then activate-region.sh adds its zone with POST /actuator/cluster/zones/<zone> and waits for the change to report COMPLETED. The engine places the zone's replicas and raises the replication factor in one change. It does not renumber any broker.

Three stages of the same cluster. First, two zones, london and paris, hold two replicas each, for a replication factor of four. The zurich slot exists but is not in the zone list. Losing either zone leaves two of four replicas, so processing stops. Second, the operator deploys the zurich brokers, adds the zone with POST /actuator/cluster/zones/zurich, one replica and priority 800, and waits for COMPLETED. Third, three zones in a 2-2-1 layout at replication factor five, where losing a database zone leaves three of five replicas and processing continues. The engine renumbers no broker and restarts no running region.Grow a Multi-Region RDBMS cluster by adding a zone1. Bootstrap on two zoneslondon2 replicasparis2 replicaszurich slotprovisioned, not declaredreplicationFactor 4 (2-2)Losing a zone leaves2 of 4 replicas:no majority,processing stops2. Add the zone onlinea. Deploy the zurich brokersb. Add the zone to the clusterc. Wait for COMPLETEDPOST /actuator/cluster/ zones/zurich{ "numberOfReplicas": 1, "priority": 800, "brokers": [ "zurich_0", "zurich_1" ]}3. Three zoneslondon2 replicasparis2 replicaszurich1 replica, addedreplicationFactor 5 (2-2-1)Losing a database zoneleaves 3 of 5 replicas:quorum kept,processing continuesThe engine renumbers no broker and restarts no running region.Adding a zone leaves the partition count unchanged.

This section applies to a region slot that you provisioned but never ran. A zone that you removed during failover comes back through Bring a region back instead.

1. Provision the infrastructure​

Raise active_region_count by one in terraform-cluster.tfvars, the variable file of the initial deployment, so the new region's cluster, Transit Gateway attachments, and security group rules exist. Then apply the file:

cd ../terraform/clusters
terraform apply -var-file=terraform-cluster.tfvars

Keep the new value in the file. A later terraform apply with a lower active_region_count destroys the region's infrastructure.

2. Update the environment​

Re-source the environment so CAMUNDA_ACTIVE_REGIONS, the cluster size, and the replication factor reflect the new count, and register a kubectl context for the new cluster:

cd ../../procedure
unset CAMUNDA_ACTIVE_REGIONS CAMUNDA_CLUSTER_SIZE CAMUNDA_REPLICATION_FACTOR
. ./export-terraform-outputs.sh
. ./export_environment_prerequisites.sh
./register-kubecontexts.sh

3. Add the region​

./activate-region.sh <slot>

The procedure does the following:

  1. Joins the new cluster to the ClusterSet.
  2. Prepares its storage class, namespace, and database secret.
  3. Renders the Helm values with the longer contact point and zone lists.
  4. Installs only the new region.
  5. Exports its services.
  6. Adds the zone to the cluster.
  7. Waits for the change to complete.

The regions already running keep their shorter contact point list, and they don't restart. The contact point list matters at bootstrap. Once a cluster forms, a newcomer only has to reach one member, and the rest learn about it by gossip. The running regions pick up the longer list on their next upgrade.

warning

activate-region.sh only adds the zone of a slot that was in regions when you bootstrapped the cluster. The reference implementation provisions its infrastructure from that slot list. List every region you may ever run in regions before the first deployment. The script rejects any slot outside the provisioned range, 0 to CAMUNDA_REGION_SLOTS - 1.

Upgrade the cluster​

Upgrade one region at a time, and wait for the cluster to report healthy before starting the next:

./check-cluster-topology.sh

Upgrading several regions at the same time risks losing quorum.

Follow the general upgrade guidance and create a backup first.

Diagnose problems​

SymptomStart here
Brokers do not reach the expected count./submariner/verify-submariner.sh, then ./submariner/diagnose-submariner.sh
Cross-region traffic is dropped./verify-cross-region-connectivity.sh
Export latency is higher than expected./measure-rdbms-latency.sh
Partition distribution looks wrong./check-cluster-topology.sh

For the underlying causes and the AWS commands that confirm them, see troubleshooting in the EKS guide.