Skip to main content

Site Recovery

Enterprise Feature

Site Recovery is available exclusively with an Enterprise license.

Site Recovery provides disaster recovery (DR) capabilities for your Proxmox environments. It manages data replication between nodes or clusters, orchestrates recovery plans, and supports failover, failback, and emergency DR operations -- giving you confidence that critical workloads can be restored quickly when disaster strikes.

Overview​

Site Recovery is built around two core concepts:

  1. Replication Jobs -- Continuous or scheduled data replication from a source node/cluster to a target, ensuring an up-to-date copy of your VMs is always available.
  2. Recovery Plans -- Predefined sequences of actions that describe how to restore a set of VMs on a target cluster in case of failure.

Together, these allow you to protect workloads, test your DR strategy regularly, and execute real failovers with minimal downtime.

Replication runs on Ceph RBD or on ZFS

Site Recovery mirrors block storage between clusters, so it needs at least two Proxmox connections carrying Ceph RBD or ZFS storage. With fewer, the page invites you to "Configure Ceph RBD or ZFS storage on two Proxmox connections to create replication jobs." and the replication, recovery plan and emergency features stay unavailable until a second suitable connection is configured in Settings > Connections. Which of the two engines a job uses is chosen at creation and never changes afterwards, see Storage Engines. Replication also needs passwordless SSH between the source and the target cluster, which the job creation dialog verifies before it lets you save. Only QEMU virtual machines are replicated; LXC containers are not part of Site Recovery.

tip

Use Site Recovery for planned DR workflows with defined source and target clusters. For one-off VM movement between clusters, use Migration instead.

Storage Engines​

A replication job runs on one of two engines, picked in Create Replication Job under Replication Engine: Ceph RBD or ZFS. Both work the same way in outline -- a snapshot taken on the source, an incremental stream pushed over SSH to the target, a rolling set of snapshots kept on each side -- but what they replicate, and where a replica can be started, differ.

Ceph RBDZFS
What is replicatedThe RBD images of the guest's disksThe zvols backing the guest's disks
Where the job landsA Target Pool, a Ceph pool reachable from every node of the DR clusterA Target storage on one Target node, because a ZFS pool belongs to the node that imports it
Where a replica can startAny node of the DR clusterThe job's target node only
Test failoverStarts the replica images themselves, and the cleanup rolls them forward againClones the newest snapshot both sides share and points the test guest at the clones
Real failoverRolls the image back to the chosen restore pointRolls the dataset back recursively, after checking that nothing depends on the snapshot

The engine is part of a job's identity: Clusters, pool, engine and VMID prefix are immutable, as the edit dialog says. To change one, delete the job and create it again.

Choosing ZFS​

Picking ZFS replaces the pool picker with Select storage and node and adds a Target node, under the hint "ZFS is local to a node: replicas live on this node and DR VMs start there." That is the whole operational difference: a Ceph replica can be brought up on any node of the DR cluster, a ZFS replica only on the node holding the pool. Pick a node that can carry the guests of that job on its own.

The target has to be a zfspool storage that is enabled and active on the chosen node; a storage that is none of those is refused by name when you save. On the source side each disk is resolved to its zvol, and the guest picker greys out what the engine cannot carry:

RefusedWhat the dialog says
A guest with disks outside ZFS"This VM has disks outside ZFS and cannot be replicated with ZFS."
A disk reference the engine cannot resolve, a passthrough disk for instance"This VM has an unsupported disk reference, such as a passthrough disk."
A guest another job already replicates"Already replicated by ...", see One Job per Guest

Two further cases are refused when the job runs rather than in the picker, because they are read from the volumes themselves: two volumes carrying the same name in different pools, which would collide on the target, and a volblocksize above 128 KiB, a geometry the replication cannot preserve.

A Ceph job is more forgiving on mixed storage: it warns that "Only Ceph RBD disks will be replicated; other disks are skipped." and replicates what it can.

Test Failover on ZFS Clones​

A ZFS test failover never touches the replicated dataset. It clones the newest snapshot the two sides share, rewrites the test guest's configuration onto the clones, and records them in a manifest written before anything else is changed. Cleanup test destroys exactly what that manifest lists, so a test interrupted halfway still leaves the cleanup something to act on.

Those clones are also why the Snapshots tab can refuse a deletion with "snapshot has dependent clones": the snapshot a live test is built on cannot be removed until that test is cleaned up. The tab refuses on the same grounds to delete what a plan references or what a job is streaming right now, and names the reason instead of failing silently.

Mixed Pairs and Where the Engine Shows​

One pair of clusters can carry jobs of both engines. The Replication tab then splits that pair into one labelled section per engine, Ceph first; a pair running a single engine stays flat. Each job, each recovery plan and both ends of their route carry the glyph of the engine behind them, and the Snapshots tab carries an Engine column.

Interface Tabs​

The Site Recovery page is organized into six tabs: Dashboard, Replication, Snapshots, Recovery Plans, Emergency DR and Simulation. The sections below cover the ones that make up a replication and recovery workflow.

Dashboard​

The Dashboard tab provides a high-level view of your replication health:

  • Overall replication status -- healthy, degraded, or critical
  • Active replication job count and their current states
  • Error count -- jobs in an error state are flagged immediately
  • Job Status Distribution -- a stacked bar splitting every job by status: synced, syncing, pending, error, paused, Partially synced for a job whose last run left at least one VM in error, and No matching VMs for a tag-based job whose tags currently match nothing
  • Recovery plan status overview

Use this tab as a daily check-in to verify that your DR posture is healthy.

Replication​

The Replication tab manages replication jobs. Each job defines what data is replicated, from where, and to where.

Creating a Replication Job

Click Create Replication Job to open the creation dialog. You can configure:

  • Replication Engine -- Ceph RBD or ZFS, see Storage Engines. The rest of the dialog follows the choice
  • Source connection -- the Proxmox cluster containing the VMs to protect
  • What to replicate -- a VMs / Tags toggle. In VMs mode you tick individual guests from the source cluster; in Tags mode you tick tags, and the job protects whatever carries them, re-resolved at every sync
  • Target connection -- the destination cluster for replicated data. ProxCenter checks passwordless SSH to it before letting you save
  • Target pool, on a Ceph job -- the Ceph pool on the target cluster that receives the replicated RBD images, listed with its usage so you can see what room is left
  • Target storage and Target node, on a ZFS job -- the zfspool storage that receives the replicated zvols, and the single node the replicas live and start on
  • Schedule -- either an RPO target, an Interval stated directly in minutes, or a calendar schedule (hourly, daily, weekly or monthly) with its own timezone. See Choosing the Schedule
  • Snapshot retention -- how many mirror snapshots are kept on each side, described below
  • VMID prefix -- the numbering the replicas get on the DR cluster
  • Replica name -- an optional prefix and suffix added to the replicated guest's own name, so production and DR are not same-named twins in the inventory
Power state does not matter

A guest does not have to be running to be replicated. Replication works on the block images themselves, whichever engine carries them, and the one step that does need a running guest, the filesystem freeze that makes a snapshot application consistent, is already skipped per VM when the guest is stopped. Stopped and paused guests are therefore listed in the picker alongside running ones, with their state shown as a coloured dot.

Templates are the exception and stay out of both modes, in the picker and in tag resolution: a replica of a template could not be started at failover, so protecting one would buy nothing.

One Job per Guest​

A guest belongs to one replication job and one only. Snapshot housekeeping works per image, not per job: it keeps the newest few mirror snapshots of a source image whoever wrote them, so two jobs on the same guest prune each other's base, and the one left without a common snapshot cannot re-seed on its own.

The picker therefore greys out a guest another job already replicates and names the job holding it, "Already replicated by ...", so you see it while choosing rather than on a refusal after submitting. The server refuses the overlap in any case, the API being reachable without this dialog. The check is made on the source cluster only: the same VMID on another cluster is a different disk and stays selectable.

Choosing the Schedule​

The schedule builder offers two modes, Continuous (RPO) and Scheduled.

ModeWhat you stateWhat it produces
Continuous (RPO)An RPO TargetA cadence derived from that target. The orchestrator schedules at a third of the target for margin, so a job whose target reads 30 minutes replicates every 10
Scheduled > IntervalRun every, a cadence in minutes or hoursExactly that cadence. The picker offers the regular values only: 1, 2, 3, 4, 5, 6, 10, 12, 15, 20 and 30 minutes, then 1, 2, 3, 4, 6, 8, 12 and 24 hours
Scheduled > Hourly, Daily, Weekly, MonthlyA calendar, with its own timezoneThe occurrences of that calendar. The RPO target is derived from the interval between two runs, and it is that derived target the RPO alerts measure against

Interval exists because an RPO target does not state a cadence. Asking for a 30 minute RPO and getting a sync every ten is a threefold cost in snapshots for the same retention window: 7 days of restore points at that real cadence does not fit under the 500 snapshot cap, and nothing used to say so. Pick Interval when what you actually want to control is how often the job runs.

Both the create and the edit dialog spell out the result under the schedule, as "Replicates every ..., so these points cover ...": the cadence the chosen schedule really produces, and how far back the kept restore points reach at that cadence. Read it before saving rather than discovering the window in production.

Tag-Based Jobs and the Protected Set

A tag-based job resolves its tags against the source cluster before every sync, so the protected set follows the tags rather than a list frozen at creation. Both directions are recorded in the job's execution log:

  • A guest that has just gained one of the tags is added, and the log reads "Tag re-resolution added N VM(s) to protection".
  • A guest that lost the tag, or that was deleted, is removed with a warning line: "Tag re-resolution removed N VM(s) from protection ... Existing replicas on the target are kept but will no longer be updated." Its per-VM row is dropped from the job so it stops being displayed as protected.
Read the removals

A removal is a protection gap, not housekeeping. The replica already on the DR cluster is left in place, so the plan still finds something to boot, but it is frozen at the last sync that included the guest. Check the tag before assuming the guest is still covered.

A Tag-Based Job That Matches Nothing

If a job's tags stop matching any guest, the job does not fail. It lands in a status of its own, No matching VMs, and:

  • No failure notification is sent, and the retry cascade reserved for real errors is not entered.
  • The next sync stays scheduled and the tags are resolved again at that point, so the job recovers on its own the moment a guest is tagged again.
  • The status is shown wherever a job status appears: the status chip and the All statuses filter of this tab, the chip on the Emergency DR tab, and the Job Status Distribution bar of the Dashboard tab.
  • RPO alerts keep firing. This is deliberate: a DR job that protects nothing has to keep nagging until somebody looks at it. See Replication and Ceph Alerts below.

This status only appears on a job that has already been created. At creation time a tag selection that matches no guest is refused outright, with "no VMs found matching tags".

A Job With One VM in Error

One VM that cannot be replicated does not stop the job. The sync records the failure, moves on to the remaining VMs, and the job ends in a status of its own, Partially synced:

  • The VMs that synced are protected as usual, and their per-VM row shows the time of the sync. The failed VMs keep a red Error row with the reason, and the job's message lists them: "1 of 6 VMs failed: VM 279: ...".
  • The job stays on its schedule. The failed VMs are retried at the next sync, alongside the others, rather than through the retry cascade reserved for a job that fails as a whole. That cascade would re-sync every VM three times for one broken one, and would then leave the whole job parked in error.
  • A failure notification is sent when a VM enters the error set, and after the automatic retry of a job that had failed as a whole, since that failure had not been reported yet. A VM that stays in error from one sync to the next is not notified again, whether the next run is scheduled or started with Sync Now.
  • The failure alert is raised as a warning, with the same per-VM summary.
  • The status is shown wherever a job status appears: the status chip and the All statuses filter of this tab, the chip on the Emergency DR tab, and the Job Status Distribution bar of the Dashboard tab. The job's card shows the per-VM summary as a warning.

A job whose VMs all fail is a whole-job failure and keeps the previous behaviour: Error, three automatic retries, then manual intervention.

Proxmox snapshots on a replicated VM

A replicated guest may carry Proxmox snapshots. Earlier versions read the snapshot sections of the guest configuration as extra disks and tried to take the mirror snapshot twice on the same image, which failed with "rbd: failed to create snapshot: (17) File exists" and left the VM, and every VM listed after it, unsynced until the snapshot was deleted. Only the current configuration is read now, so a Proxmox snapshot no longer interferes with replication. The replica itself does not carry the guest's Proxmox snapshots: they live on the source images only.

Managing Replication Jobs

For each job, the following actions are available:

ActionDescription
SyncTrigger an immediate replication sync
PauseTemporarily suspend replication
ResumeResume a paused replication job
DeleteRemove the replication job entirely

Selecting a job opens its details dialog: the source and the target side by side, joined by a connector that carries a stream while a sync runs and a plain rule otherwise, one row per protected guest with the volume and duration of its last run, the bandwidth chart, and the tail of the execution log with the history of sync operations, timestamps and results. The per-guest table is paginated and is shown for a single-guest job too, where the figures of the last run appear nowhere else.

A pause takes effect even mid-run

A Pause issued one second into a scheduled run used to be undone by that run's own completion, which set the status and the next occurrence unconditionally: the job went back to synced with a fresh schedule and kept replicating while the interface, and the operator, believed it paused. A run now records only what it measured, and never moves the status or the next occurrence of a job that has been paused or failed over in the meantime.

tip

Run a manual sync after making significant changes to a protected VM to ensure the latest state is replicated before relying on it for recovery.

Naming the Replicas

A replica carries the source guest's name verbatim unless you say otherwise, so production and DR show up twice under the same name in the inventory and the replica is what gets started by mistake. Replica name, in Create Replication Job and in Edit replication job, takes an optional prefix and an optional suffix for the replicated guest's own name on the DR cluster: web01 becomes DR-web01, web01-DR or both. Leave both empty and the replica keeps the source name. A guest Proxmox never named keeps no name, since a prefix on its own is not a name.

The dialog previews the result on a guest you have actually selected, so you see the real name before saving.

Both fields accept letters, digits and hyphens only, up to 24 characters, and the composed name may not start or end on a hyphen. That is not arbitrary: Proxmox stores a guest name as a DNS name, and ProxCenter writes the replica's configuration to the target node directly, so a name Proxmox would refuse is not caught on the way in and would surface later as a broken replica configuration. The rule guarantees the composed name is always one Proxmox accepts.

Unlike the VMID prefix, the replica name can be changed on a job that is already running. The replica's configuration is rewritten from the source at every sync, so a new prefix or suffix is applied on the next cycle, without re-copying a single disk.

Snapshot Retention

Each job keeps a rolling set of mirror snapshots on both sides, and how many it keeps is configurable in Create Replication Job and in Edit replication job, under Snapshot retention:

FieldMeaning
Keep on sourceMirror snapshots kept per disk on the source cluster after each sync
Keep on target (DR)Mirror snapshots kept per disk on the DR cluster after each sync

The unit is a count of snapshots per disk, not a duration. Both sides default to 3, the minimum is 2 and the maximum is 500. Each field pairs a number box with a slider; the slider stops at 50, so type the value in the box if you want more than that.

Retention is one of the settings you can change after creation, along with the replica name, the schedule, the rate limit and the guest list. The clusters, the target pool, the engine and the VMID prefix are immutable: to change one of those, delete the job and recreate it. On a tag-based job the list of tags is immutable too, but that is about the list of tags, not about the guests: the protected set follows whatever carries those tags and is re-resolved at every sync, which is what makes tagging a guest enough to protect it.

Editing the guest list. A job in VMs mode can gain or lose guests without being recreated, from Edit replication job. The picker offers only the guests whose disks this job's engine can replicate, since one without them syncs nothing and the run merely logs that it skipped it. Five cases are refused, and said so rather than silently dropped: a guest whose replica is currently started, a guest a recovery plan lists, a guest another job already replicates, an empty list, and a job that is mid-sync or failed over. A guest that leaves keeps its replica and its mirror snapshots on the DR side, exactly as deleting a job does, so removing a guest from a job is not a way to clean up the DR cluster.

Target retention sets how far back you can recover

The restore points offered when you fail over are the mirror snapshots that still exist on the DR side, so Keep on target (DR) is what decides how far back in time you can recover. Raise it when you want more than the last few sync points, keeping in mind that every snapshot holds Ceph capacity.

How pruning behaves

Pruning runs at the end of each sync, per disk, and a floor of two snapshots is always re-applied so a common incremental base survives on both sides. It also removes at most ten snapshots per image per run, so lowering retention from a large value does not free the space at once: the backlog drains over the following syncs and the job log reports how many are still pending. During a failback the two roles are swapped, since the DR side is the one being replicated from.

What a Sync Transfers

The first sync of a disk carries its allocated data only, never its virtual size. A thin 500 GB disk that holds 96 MiB of data costs a 96 MiB transfer, and an empty disk added to a protected VM is replicated in seconds. The job log states both figures: "full sync (96.0 MB allocated of 500.0 GB)". Every later sync is incremental and carries the changes since the last common snapshot, logged as "incremental sync from base ... (up to 36.0 MB to send)", an upper bound counted in whole objects. The bytes credited to a sync, in the per-VM rows and in the volume figures of the Dashboard tab, are the bytes that actually crossed the wire, so a 20-second incremental on a 60 GB disk is no longer counted as 60 GB.

The live progress percentage and the rate limit of a job both rely on the pv tool being installed on the source cluster's nodes (apt install pv). Without it the sync runs the same, the progress bar simply stays at zero until the disk completes and the bytes credited are the estimate computed before the transfer rather than a measurement. That estimate is read from the image's object map and rounded up to whole objects, 4 MiB each by default, so it can overstate a sync made of many small scattered writes.

An interrupted first sync heals itself

A first sync stopped halfway, by an orchestrator restart for instance, leaves a target image without any snapshot on the DR cluster. Site Recovery stamps every image it creates with the job that created it, so the next run recognises its own leftover, logs "target image left by an interrupted initial sync, replacing it", and starts over. Anything else that bears the destination name is treated as somebody's data and never removed: a disk of a DR-site VM that happens to use the same VMID, or an image that carries snapshots of its own, none of them shared with the source. The run refuses to touch it and says so, and you have to remove or rename that image, or its snapshots, on the DR cluster before the disk can be replicated. Images left by versions before this one carry no stamp and fall in that second group.

Replication Network

By default the replication stream travels to the management address of the target connection, the one ProxCenter uses for the Proxmox API. A connection can carry a Replication network instead, in Settings > Connections, in the SSH section of a Proxmox VE connection: a CIDR such as 10.10.50.0/24. When it is set, Site Recovery looks up the address of the DR node inside that network and sends the stream there, the way Proxmox itself uses its migration network. The job log records the choice: "Replication stream target: 10.10.50.11 (replication network 10.10.50.0/24)".

The setting belongs to the connection whose nodes receive the stream. For a normal sync that is the target connection; during a failback, the reverse replication flows towards the source connection, so it is that connection's replication network that applies. Set it on both connections when both directions should stay off the management network.

A few things to know before setting it:

  • The nodes of the other cluster must have a route to that network, and the same SSH trust applies: the keys that let the source cluster reach the target cluster's management address let it reach the replication address too, since Proxmox authorises them cluster-wide.
  • The SSH check that the job dialog runs before it lets you save tests the replication address, not the management one, so a missing route shows up before the job exists.
  • The setting is checked at every sync. A node with no address in that network fails the run with an explicit message that lists the addresses it does carry, rather than falling back to the management network without telling you. A network that is not a valid CIDR is refused when you save the connection.
  • The Proxmox API traffic is not affected: it keeps using the connection address.

Recovery Plans​

Recovery Plans define the procedure for restoring services on a target cluster. The tab lists all existing plans and lets you create new ones.

Creating a Recovery Plan

Click Create Plan to define:

  • Plan name and description
  • Source and target clusters
  • Associated replication jobs -- which replication jobs feed into this plan
  • VM startup order and dependencies

An existing plan is edited from the same form, prefilled. Editing is disabled while the plan is executing, failing back or failed over.

Recovery Plan Operations

Each recovery plan supports three operations:

OperationDescription
Test FailoverExecutes the recovery plan on the DR cluster, by default with every network interface shut down. Production workloads are not affected, and the source VMs are left alone. Each recovered VM's console is captured while the test runs. Two options are offered before the run starts, see Test Failover Options. Use this to validate your DR strategy regularly.
FailoverActivates the recovery plan for real. The source VMs are fenced, then the replicas are started on the DR cluster from the chosen restore point. Use this during an actual disaster.
FailbackAvailable once a plan is failed over. It brings the workloads home in two phases: a reverse incremental sync from the DR site to the source, then an operator-driven cutover. See Failback.

When any operation is executed, ProxCenter tracks its progress in real time, polling the execution status every 3 seconds and displaying step-by-step updates. During a real failover and during a failback cutover, each VM also shows the step it is on: Stopping the source VM, Rolling back to the selected restore point, Starting on the DR site, and for the failback cutover Stopping the DR VM, Applying the final delta, Starting on the source site and Re-protecting. The step is shown while the VM is running that step; a VM that has finished shows its result, and a VM that failed shows its error instead.

warning

Failover is a disruptive operation. Ensure the source site is truly unavailable before initiating a production failover, as running the same VMs on both sites simultaneously can cause data corruption.

No step detail on a test failover

A test failover does not report per-VM steps. It reports progress per VM and, in its final phase, a three-step checklist for the console captures: Starting VMs, Letting guests settle and Capturing boot screenshots.

What a Test Failover Does to Replication

A guest whose replica is under test cannot be replicated onto at the same time: the test runs on the very images the next sync would overwrite. Site Recovery therefore claims the guests of the plan, not the jobs behind them:

  • A job whose guests are all covered by the plan is paused outright for the length of the test, and the cleanup resumes it.
  • A job that also carries guests outside the plan keeps running and skips only the claimed ones. Its other guests keep replicating and their RPO is unaffected.

A claimed guest is reported as Suspended (test in progress) on its row in the job's details, so a row that stops advancing says why instead of merely looking stale. The claim is read off the plan rather than recorded on the job, so the cleanup releases it by clearing the active execution, with nothing to resume.

Why this matters most on a tag-based job

A tag-based job routinely covers far more than a plan does, so testing one guest of such a job used to stop protecting every other guest it carried for the whole length of the test. Claiming per guest is also the only mechanism that can cover a tag-based job at all, since its own VMID list is written at the end of a run and therefore always lags.

A real failover still pauses the jobs unconditionally. It promotes the plan, and nothing skips its guests afterwards.

Restore Points

By default a failover boots each VM from the latest replicated state. Both Test Failover and Failover let you pick an older point instead, per VM, from a Choose the restore point selector on each VM row of the dialog. The default entry is Latest (default); the other entries are the replicated snapshots that still exist on the DR side, newest first, shown with their timestamp.

The list is read from the live state of the DR cluster each time the dialog opens, and only snapshots present on every disk of a VM are offered, so a partially failed sync cannot boot a torn VM. A VM with nothing to offer shows No restore points. When the list cannot be read at all, the dialog says "Restore points could not be loaded, the latest replicated state will be used." and runs on the latest state.

An older restore point is destructive on a real failover

Choosing an older point on a real failover re-bases the DR image: "Failing over to an older restore point permanently deletes the newer DR snapshots of that VM." The dialog shows that warning as soon as you select a point. A test failover does not re-base, and its cleanup rolls the image forward again, so testing an older point costs you nothing.

Test Failover Options​

The Test Failover dialog carries two settings, in a panel above the VM list. They are editable until you press the button that starts the run, and they apply to that run only: nothing is stored on the plan.

OptionDefaultWhat it does
Network isolationOnEvery network interface of every recovered VM gets link_down=1 before the VM is started, so the replica boots with its cards administratively down. The interfaces themselves are kept, so nothing about the guest's networking is lost, and the flag disappears on its own at the next sync, which rewrites the replica configuration from the source.
Boot screenshot delay45 sHow long ProxCenter waits after the last VM has started before capturing the consoles. Accepts a whole number of seconds between 5 and 600.

Turning network isolation off. The switch is labelled "DR VMs boot with every network interface shut down. Turn it off to test with the interfaces connected to the mapped target bridges." Turning it off raises a warning in the dialog, on the spot:

Address collision

"The DR VMs will boot with their production addresses on the mapped target bridges. Only run this where the DR site is not bridged to production, otherwise addresses will collide."

A replica carries the source guest's network configuration verbatim, MAC address and bridge name included, so it comes up on the DR cluster's bridge of that name with the production addresses of the guest it copies. If that bridge reaches the production network, two machines with the same address land on the same wire. Run a connected test only on a DR site that is genuinely separated from production, and use it for what isolation cannot answer: whether the application actually serves, not merely whether the guest reaches a login prompt.

Once the run has started, the option it used is no longer editable but stays visible: a chip on the dialog reads Network Isolated or Network connected, so an execution you come back to never leaves you guessing which mode it ran in.

Choosing a delay. 45 seconds is enough for a guest that reaches a login prompt quickly, and too short for a large Windows guest or anything that runs filesystem checks at boot: those get photographed mid-boot. Raise the delay to whatever your slowest guest in the plan needs. The value applies to the whole pass, not per VM, since the wait happens once after all the VMs have started. An empty field, a decimal, or a value outside the 5 to 600 range is refused and the run is blocked until you fix it.

A real failover is never network isolated and takes no such options: it starts the replicas with their interfaces connected, which is the point.

Boot Screenshots on a Test Failover​

During a test failover, ProxCenter captures the console of each recovered VM and keeps the image with the execution, so you can see what a guest actually booted into instead of trusting a power state. The capture itself is automatic: there is no switch to turn it on, only the delay described above.

The capture pass runs once, after every VM of the plan has been started. ProxCenter waits for the boot screenshot delay of that execution, 45 seconds unless you changed it, then takes exactly one capture per VM. The dialog shows the wait as a countdown next to Letting guests settle, and the countdown follows the delay you chose.

Executions from before v1.4.8

The delay is stored on the execution. Test executions recorded before the option existed carry no value and are read back as the 45 second default, so their history stays consistent.

The captures appear as small thumbnails in two places, both on the Recovery Plans tab: under the corresponding test execution in the plan's Execution History, and as a camera button on each VM row of the test failover dialog, labelled View boot screenshot. Clicking a thumbnail opens the full-size image with the VM name and the time it was captured.

What can prevent a capture

Captures need the orchestrator and SSH access to the DR cluster, they only work for QEMU virtual machines, and they need a guest with an emulated display: a serial-console-only guest produces nothing. A capture that does not succeed is skipped silently. It does not fail the test and it produces no error message, so the only sign is a missing thumbnail. A capture is also only attempted when at least one VM of the plan actually reached a completed state.

Captures survive the test cleanup

Running Cleanup test does not delete the captures: they are kept as the archived evidence of the test. They are removed only when you clear the plan's execution history or delete the plan itself.

Source Fencing on a Real Failover

A real failover fences each source VM before starting its replica, so the same guest cannot end up running on both sides. Fencing means, in order: a graceful shutdown request with a 60 second deadline, a hard stop if the guest has not stopped by then, and finally clearing onboot in the source VM's configuration so a recovering source site cannot boot the guest back into a split brain. This is unconditional on a real failover and never happens on a test failover. There is no option to turn it off.

Fencing is best effort on a real failover

A real failover is a disaster procedure, so it does not stop when fencing cannot be carried out. If the source cluster is unreachable, fencing is skipped for the whole failover and the DR replicas are started anyway; the orchestrator log records that it skipped it. Treat fencing as a safety net that usually works, not as a guarantee: if the source site is only partly down, confirm the source guests really are stopped before you let the site come back.

Locked After a Failover

Once a plan is failed over, ProxCenter locks it and its replication jobs so nothing can overwrite the copy that is now production. The plan carries a Failed Over status and the tooltip "Plan is failed over -- only failback is possible.", with Test Failover and Failover disabled. Each replication job of that plan carries a Failed over chip, and its tooltip explains that the job's DR copy is now production and that resuming replication would overwrite it, so a failback has to come first. Sync Now, Resume and Edit are disabled on that job. Deleting is still allowed.

The lock is released per job by the failback, once every VM of that job has completed its cutover.

Test Cleanup

After running a test failover, use the Cleanup test action to tear down the test environment and release resources on the target cluster. This ensures that test artifacts do not consume storage or interfere with future tests. While a test is awaiting cleanup, the plan row shows a Cleanup pending chip, a second test failover is refused with "A test failover is already active for this plan. Clean it up first.", and reopening the dialog shows a banner naming the date the pending test was started.

A cleanup that did not finish says so

The dialog reads the orchestrator's own verdict, not the absence of errors. A cleanup cut off halfway comes back as Cleanup did not complete, with "The cleanup could not be completed, the DR guests may still be running on the replica images", and it keeps its retry button until everything really is cleaned. Only a run the orchestrator reports as complete shows Cleanup completed. Press retry until it does: an unfinished cleanup means DR guests are still up on the replica images, and replication is about to resume onto them.

A cleanup handles the guests one after another, around twenty seconds each for a single-disk guest, and is given 55 seconds per attempt. That budget sits deliberately under the 60 second read timeout a packaged install's reverse proxy applies, so an expiry comes back as a ProxCenter error you can act on rather than as a proxy error page. A plan of several guests may therefore need more than one pass, which is what the retry button is for.

DR replicas are no longer enrolled in automatic HA either. A replica reaches its target cluster as a brand new guest and used to be handed to the cluster resource manager in the started state, which overrode the onboot: 0 its configuration carries on purpose and undid every stop the cleanup issued.

Execution History

Select a recovery plan to view its execution history -- a chronological list of all test, failover, and failback operations with their outcomes, timestamps, and any errors encountered. Test executions also show their boot screenshot thumbnails.

A Clear history action removes the past executions of the plan. It asks for confirmation first, because "Boot screenshots are deleted with them."; the active test execution is kept.

Emergency DR​

The Emergency DR tab is designed for critical situations where you need to act fast without going through the full recovery plan workflow. It aggregates every replication job and recovery plan into one list, giving operators a single view to manage a crisis.

What sits where. A guest's row carries only the actions that really are per guest. Anything that acts on a whole recovery plan sits on that plan's card header, where its blast radius is legible:

LevelActionWhat it does
Guest rowStart VM on DR siteBoots that one replica on the DR cluster, from a restore point you pick
Guest rowStop VM on DR sitePowers that replica off again, optionally resuming its replication
Plan headerEmergency FailoverFails the whole plan over immediately
Plan headerPlan failbackOpens the plan's failback, the same dialog as Failback

Plan failback is disabled unless the plan has failed over, under the tooltip "Failback covers the whole plan, and only once the plan has failed over", and reads Open plan failback while one is already running. Emergency Failover is disabled for a plan that is executing, failing back, or already failed over.

The Replica column. Each row states the live state of the replica it acts on, Started or Stopped, read from the DR cluster's inventory. The chip says started, never running: Proxmox reports a paused guest as running too, and an operator reading "running" on a DR row mid-incident must not be told the replica is serving. The two actions follow that state, so Start is disabled on a replica already up and Stop on one already down, and either action keeps refreshing the lists it changes for a minute afterwards. A guest the orchestrator skips because its replica is up says so on its own line instead of repeating its job's healthy status.

Starting a replica. The dialog offers the restore points of that one guest, newest first, with Latest (default) at the top. The rule is the one a plan failover uses: only a snapshot present on every disk of the guest is a restore point, since booting from one missing on a disk would tear the guest across disk epochs.

Two things the dialog states before you confirm:

  • Only that guest stops replicating. Starting a replica claims the guest on its own row; the job keeps running for every other guest it carries, and only a job left with nothing to replicate is paused. The wording follows the count: "Only this guest stops replicating while its replica runs. The N other guests of its job keep going.", or, for a job carrying that guest alone, "This guest is all its replication job carries, so the job pauses while the replica runs."
  • Picking an older point keeps the newer ones. Unlike a real failover, an emergency start does not re-base: "The replica boots from that snapshot and the newer restore points are kept: stopping it and resuming replication rolls the image forward to the newest one." An emergency start is meant to be reversible, which is exactly why it must not destroy restore points.

A start is refused while that guest's job is syncing, rather than queued behind it: a transfer is writing to the very disks the start would roll back. The button says so, "This guest's replication job is syncing right now. Starting the replica is refused while a transfer writes to its disks: wait for the sync to end, this button unlocks on its own."

Stopping a replica. The stop dialog powers the replica off and carries a Resume replication for this job switch. Resuming is not free and the dialog says what it costs: "Anything the replica wrote since the last sync is discarded when replication resumes." The replication guard rolls the replica image back to its last mirror snapshot before the next stream. Leave the switch off to keep the job paused and the replica's disks as they are; either way, "This guest rejoins its replication job as soon as the replica is stopped."

warning

Emergency operations bypass the normal validation steps. Use them only when time is critical and you understand the implications of starting replicated VMs without a full plan execution.

Replication and Ceph Alerts​

Until v1.4.8 an alert threshold could only describe how full something was, so a replication job that fell behind, or failed outright, was visible on the Site Recovery page and nowhere else. The orchestrator now raises three alert types that concern this page, evaluated once a minute and surfaced on Alerts like any other. They are configured in Settings > Alert thresholds, under Performance & replication.

AlertShipsRaised when
Replication RPOEnabled, 25% graceA job's last successful sync is older than its own RPO target plus the grace margin
Replication failuresEnabledA job is in the error state
Ceph OSD latencyDisabledAn OSD answers slower than the threshold you set

How the RPO Alert Judges a Job​

The comparison is made against the RPO target of that job, not against a figure shared by the fleet, because a job set to 15 minutes and a job set to 24 hours have nothing in common. The grace margin is therefore a percentage of each job's own target, not a fixed delay: ten minutes late is meaningless against a 24 hour target and unacceptable against a 15 minute one.

With the default 25% grace, a job whose target is 15 minutes warns once its last successful sync is 18 minutes and 45 seconds old, and turns critical at twice that. A margin is needed at all because a job that is perfectly on time mechanically reaches its RPO just before each run: with no grace, every healthy job would alert.

A few cases are treated on purpose:

  • A job that has never completed a sync is measured from the date it was created, so the most broken job in the fleet is not the only one that never alerts.
  • A paused job and a failed-over job are exempt, and any RPO alert already raised on them is cleared. One was stopped by an operator, the other replicates in the opposite direction, so the original target no longer means anything.
  • A job in the No matching VMs state keeps alerting. Its clock keeps running because a DR job protecting nothing is exactly what somebody needs to be told about.
  • A calendar job is measured against the target derived from its schedule. A schedule from which no interval can be derived leaves the check with nothing to compare and it is skipped.

The Failure Alert​

Raised as soon as a job enters the error state, and independent from the RPO alert: a job can fail without having missed its RPO yet, and can miss its RPO without ever being marked failed, for instance an unreachable source that leaves the job sitting in pending. It escalates with the retries: warning while attempts remain, critical once the three automatic retries are spent, since at that point nothing will restart on its own. The job's own error message is appended to the alert.

A job in the Partially synced state raises the same alert as a warning: the job keeps running and retries the failed VMs at every sync, so there is no retry budget to exhaust, but at least one VM is not protected and the alert carries the per-VM summary. It clears once every VM syncs again.

While a job is syncing, or pending between a manual Sync Now and its start, the failure alert is neither raised nor cleared: a run in progress proves nothing yet, and counting it as a recovery would clear the alert of any rerun longer than a few minutes only to raise it again at the end.

When these alerts clear

Deleting a job clears the alerts it had raised, and so does turning the corresponding family off in Settings > Alert thresholds, so a setting you switch off does not leave stale alerts active for ever. Recovery notifications are covered in Alerts.

These three alerts are produced by the orchestrator and read on the Alerts page, both of which are Enterprise. (Enterprise)

Failback​

Failback brings a failed-over plan home. It is started from the Failback action in the plan's detail panel on the Recovery Plans tab, or from the shortcut on a VM row of the Emergency DR tab. Like a real failover, it asks you to type the plan name to confirm, under the title Execute Failback.

Failback runs in two phases, and the moment you switch from one to the other is yours to choose. Nothing cuts over on its own.

Phase 1: Reverse Replication​

Starting the failback puts the plan in the Failing back state and opens a monitor headed "Reverse replication in progress". From then on ProxCenter replicates incrementally from the DR site back to the source, at the cadence of the job, while the DR VMs stay online and keep serving.

Before its first reverse pass, each VM's source disks are rolled back onto the most recent snapshot that both sides still have in common, so the reverse delta has a valid base. That rollback is automatic and per VM.

The monitor is a per-VM table:

ColumnMeaning
VMThe protected virtual machine
Last reverse syncWhen the last reverse pass completed, or No reverse sync yet
TransferredHow much data the last pass moved

The panel states the rule to follow: "Reverse replication is running while the DR VMs stay online. Trigger the cutover when the last sync is fresh enough."

You can close the dialog; the reverse replication keeps running on the orchestrator. While a plan is failing back, its detail panel offers only Open failback to come back to the monitor, and the other plan actions are gone. An orchestrator restart resumes the reverse replication by itself.

Transferred is a best-effort figure

The byte counter reads the throughput of the transfer, which requires the pv package on the DR node. Without it the column stays empty. It does not affect the sync itself.

Phase 2: Cutover​

When the last sync is fresh enough for the downtime you can afford, press Cutover. ProxCenter asks to confirm with "Stop the DR VMs, apply the final delta and start the source VMs?", then works through the VMs in the plan's boot order. Each one reports its step: Stopping the DR VM, Applying the final delta, Starting on the source site, Re-protecting.

When every VM is through, the dialog reports "Failback complete" and states that the source VMs are running and that replication protects them again in the original direction. The plan returns to Ready, and each replication job goes back to pending and resumes forward replication incrementally from the final delta. There is no full re-seed, and the lock on a job is released only once every VM of that job has completed its cutover.

Cutover readiness. The Cutover button stays disabled until every VM has completed at least one reverse sync, with the tooltip "Waiting for the first reverse sync of every VM". Beyond that first pass there is no automatic freshness threshold: judging when the remaining delta is small enough is the operator's call.

Cancelling. During phase 1, Cancel failback stops the reverse replication and returns the plan to failed over. It confirms first, and it is explicit that nothing is changed on the VMs. Once the cutover has started it can no longer be cancelled, only re-run.

Re-running an interrupted cutover. The cutover is safe to re-run: VMs that already completed it are skipped. If a VM's cutover fails after its DR VM was confirmed stopped, ProxCenter restarts that DR VM so it is not left down, marks the VM as failed with the step it died on, keeps the plan in Failing back, and resumes reverse replication so you can fix the cause and cut over again. The one case where the DR VM is deliberately left stopped is when the source VM was found unexpectedly running, because restarting the DR VM would create a split brain.

What a failback requires
  • The plan must be failed over. The Failback action is not disabled on a plan in any other state, so on a plan that has not failed over the dialog simply returns the error plan must be failed over to start failback.
  • Each VM needs a mirror snapshot that still exists on both the DR and the source images of every one of its disks. When none survives, that VM reports that a full re-seed would be needed, which this version does not support.
  • Every DR disk must still have its counterpart in the source VM's configuration on the production cluster: that configuration is what tells ProxCenter which production pool each disk goes back to, so the two sites do not need to name their Ceph pools alike. A disk that exists only on the DR side (added while running there) makes the VM fail with a message naming both images.
  • A disk that carries a newer common snapshot than the plan-wide base has to have that stale snapshot removed first; the VM's error names it.
Locked while failing back

While a plan is failing back you cannot edit it, delete it ("Cancel the failback before deleting this plan."), run a test failover or a failover on it, clean up a pending test, or start a second failback. Each of those is refused with an explicit message rather than silently ignored.

No notification is sent when a failback finishes

Test failovers and real failovers send a notification; a failback does not. It does appear in the activity feed and in the plan's execution history.

HA Failback​

A different feature with the same name

The Failback switch described here belongs to Proxmox HA resources and has nothing to do with the disaster recovery failback above. One returns an HA resource to its preferred node inside a single cluster; the other brings a failed-over recovery plan back from the DR site.

The HA tab lets operators toggle Failback on individual HA resources. Use this when you want a resource to automatically return to its preferred node after the node becomes healthy again.

Failback is configured per resource rather than globally, so you can keep conservative behavior for stateful workloads while enabling automatic return for services that are safe to move back.

Orchestrator Resilience​

Site Recovery, like every orchestrated feature, reaches Proxmox through the address configured on the connection. Since v1.4.7 the orchestrator no longer goes down with the node that address points at. After two consecutive network failures on a connection it probes the other nodes of the same cluster and continues against the first one that answers, then returns to the configured address as soon as that address answers again. The connection you configured is never rewritten. An HTTP error from a node that did answer is not a reason to move; only a genuinely unreachable node is.

It also addresses guest commands to the node that actually owns the guest, looked up from the cluster instead of assumed. Before v1.4.7, a replication job whose guest did not live on the configured node could not freeze that guest's filesystem, and the snapshot came out crash consistent instead of application consistent, with nothing in the log to say so.

Node addresses are discovered from the cluster and stored per node. A node that stops answering now keeps its recorded address: a failed discovery used to erase it, which removed that node from the failover candidates at the exact moment they mattered.

The SSH sessions Site Recovery opens follow the same rule, including the target of the RBD transfer itself. A replication run, a recovery run or a preflight opens its sessions against the configured connection address, and when that node does not answer it proceeds against the first other node of the cluster that does. The choice is made once, when the session opens: a transfer already streaming to a node that goes down fails that run, and the next scheduled run picks a live node. A node that answered and refused the session, because of its credentials for instance, is not a reason to move, since every node of the cluster would refuse the same credentials. Nothing is spread across nodes on purpose: a healthy cluster keeps every SSH session on the configured node, so its logs stay in one place and a fault on one node does not turn into an intermittent problem.

The "Behind reverse proxy" option turns the failover off

A connection marked Behind reverse proxy in Settings > Connections is deliberately excluded, because the point of that option is to reach the cluster through a DNS name and a certificate rather than through node IPs. Its helper text says so: "Disables failover to node IPs (use when accessing via DNS + reverse proxy with SSL)".

There is nothing to configure for any of this and no indicator in the interface: the orchestrator log is where a node failover is recorded.

Workflow Example​

A typical Site Recovery workflow looks like this:

  1. Set up replication: Create replication jobs for your critical VMs, pointing to a secondary Proxmox cluster
  2. Create a recovery plan: Group the replication jobs into a recovery plan with the correct startup order
  3. Test regularly: Run test failovers monthly to validate that recovery works as expected, then clean up
  4. Respond to incidents: If the primary site fails, execute a failover from the Emergency DR tab or the Recovery Plans tab, picking a restore point per VM if the latest state is not the one you want
  5. Restore normal operations: Once the primary site is back, start a failback on the plan, let the reverse replication bring the source up to date, then trigger the cutover in a window you choose

Permissions​

PermissionDescription
automation.viewRequired to open Site Recovery and read its jobs, plans, snapshots, executions, restore points and boot screenshots
automation.manageRequired to act: create, edit, pause, resume, sync or delete a job, create or delete a plan, run a test failover, a failover, a cleanup or a failback

Users without automation.view will not see the Site Recovery entry in the navigation sidebar. The entry is also shown only in the provider view, not in a tenant context, and it requires the Ceph replication feature of the Enterprise license.