News and analysis
A backup becomes protection when a restore has been tested
How to set RPO and RTO, test recovery of an application and its data, and choose a backup strategy that fits business operations.
A successful backup answers a limited question: did the system write a copy of the data? The business needs to know whether the team can restore order processing, staff access and customer services. Downloading the archive, rebuilding the environment, obtaining credentials and checking application compatibility all sit between those outcomes. Recovery planning should cover that entire path.
Decide how much work can be lost
RPO describes acceptable data loss in time. RTO describes the target time to restore operations. AWS advises setting recovery objectives from business needs. They are design objectives, not guarantees created by switching on a backup schedule.
Consider a hypothetical service that takes a database copy overnight and fails late the following day. The useful questions concern the intervening work: can those transactions be entered again, where are the confirmations, and who will reconcile them? If that information cannot be reconstructed, a daily copy does not fit the process. If it can be reconstructed reliably, improving the recovery procedure may deserve attention before increasing backup frequency.
Applying one requirement to every dataset makes that discussion harder. Documents, customer requests and technical logs have different operational value. For each critical process, identify the authoritative source, the consequences of loss and an acceptable recovery window. Process owners and engineers should agree the actual numbers together, including any contractual requirements that need separate review.
Include the difficult steps in the exercise
Run the restore in an isolated environment. Give the team a selected backup and the written procedure. Avoid starting with a conveniently prepared server when the real plan depends on building a replacement. Record time spent obtaining access, transferring data, starting components and validating the working process.
A restored database needs both integrity checks and application checks. Open an agreed sample order, inspect an attachment and verify user permissions and related records. Choose examples beforehand and keep customer information out of the exercise report. PostgreSQL documents several backup approaches, but selecting an approach does not validate an individual application.
Configuration, component versions and a way to retrieve secrets also belong in the plan. Keeping the only decryption key on the failed server creates a shared point of failure. Copying every secret into a broadly accessible instruction creates another problem. The procedure should identify an authorised access method and the person responsible for it.
Define who accepts the recovered service
Once the software starts, a process owner should confirm that the recovered records support real work. An order system may need reconciliation against an external payment record. An internal portal may need document and permission checks. A functioning sign-in page provides insufficient evidence for either decision.
Plan how users move back to the recovered system. The old instance must not silently accept changes that will later be discarded. Decide how writes stop, where incoming requests go and how the team announces the return to service. Where systems depend on each other, include their startup order and the treatment of work already in progress.
Turn the measurement into a practical decision
The exercise should produce a measured duration, a list of gaps and named owners for fixes. If transferring the archive consumes most of the recovery window, investigate its size and delivery path. If credentials cause the delay, a faster server will not address it. If the data restores but the application cannot use it, align versions and release checks.
Choose the next exercise based on failure consequences and system changes. A different schema, encryption method or substantially larger dataset can invalidate an earlier measurement. Automating repeatable steps helps, while acceptance of the restored process still needs explicit criteria.
Start with one critical scenario and complete it. Infrastructure work can then focus on observable gaps and the cost of addressing them. A successful exercise demonstrates recovery under the conditions tested; it does not promise the same result for every possible failure.
My archived Debian infrastructure description connects application hosting with recovery. Its scope included persistent disks, application and database backups, and boot-disk snapshots. The page records requirements and a proposed setup; it does not establish a completed restore or a measured RTO. It can serve as a dependency checklist when planning your own exercise.
The historical Debian remote-desktop review also separates administrator access from service readiness. Reaching the machine leaves application startup, usable records and the team's return to work to verify. List the steps after access is restored and name the person who accepts the recovered process.
Sources
Sources checked on 7 October 2026.