Restore drills or fantasy DR: Why untested backups fail the first real outage

Backup dashboards love green checkmarks. The overnight job finished, the retention policy looks tidy, and the DR binder still impresses whoever last flipped through it for an audit. Then the first real outage arrives, someone asks for a restore under time pressure, and the story falls apart. The backup exists. The restore does not.
Mid-size teams rarely fail because they forgot to buy backup software. They fail because nobody owns a credible restore, RTO and RPO live in a slide, and drills never touch production-like data in a usable environment. This post is about making DR real through restore drills, clear ownership, and evidence you can trust when the clock is running.
Green backups are not a restore
A successful backup job proves that bytes were copied somewhere according to a schedule. It does not prove that those bytes can become a working system again. A working restore needs consistent database recovery, object-store keys and permissions, and application stacks with identity, networking, secrets, and a place to land that looks enough like production to matter.
In practice, the gap shows up as a familiar sequence. Restore starts, credentials are stale, the target cluster is too small, and the dump format is wrong for the current engine version. Encryption keys sit with a person who is on holiday. The runbook assumes a tool that was replaced two migrations ago.
That's why treating backup success as DR success is a category error. Backup is intake. Restore is the product. If you only measure intake, you are measuring the wrong thing.
Who owns RTO and RPO when the pager rings
RTO and RPO are not decorations for a compliance appendix. They are product commitments. Someone has to decide what "back" means for each critical service, who may declare a restore, and who signs off that the restored environment is safe to put in front of users.
However, mid-size org charts often split the pieces. Platform owns the backup job, a product team owns the data model, and security owns the keys. Nobody owns the end-to-end restore under a clock. When the outage starts, ownership becomes a hallway negotiation while customers wait.
Write the ownership down in operational language. Name the service, the recovery objective, the person who can start a restore, the person who can accept the restored state, and the environment where that restore is allowed to land. If those names change every quarter, refresh the list the same way you refresh on-call rotations. Folklore is not an owner.
Why drills miss the failure mode
Many drills never leave fantasy. They restore a tiny sample into an empty sandbox, skip identity and DNS, use synthetic data that never exercised the ugly production edge cases, and stop when "the file came back" instead of when "the service can take traffic with known gaps."
Still, a calm drill that never stresses time, dependencies, or people will not prepare you for a loud night. The failure mode you care about is restore under pressure into something resembling the real blast radius. That includes stale runbooks, missing break-glass access, and the first surprise dependency that only production traffic reveals.
Design drills the way you design incidents you hope never to have. Pick a service that matters and an objective that hurts if missed. Restore into an environment that shares identity, networking patterns, and data shape with production. Time the work, record where humans had to invent steps, and fix those steps before the next drill, not during the next outage.
Evidence that survives contact with an auditor and an outage
Auditors like policies. Operators need artifacts. Good restore evidence is boring and specific. It includes timestamped drill records, named owners, measured duration against the stated RTO, notes on what could not be restored and why, and links to the runbook version that was actually used, not the one nobody opened.
Meanwhile, avoid theater. A signed PDF that says "DR tested annually" without a restore log is not evidence of recovery. A green backup chart without a successful application-level smoke check is not evidence either. Prefer a short drill report that a tired on-call engineer could reuse at 02:00 over a polished narrative written for a quarterly review.
In reality, the best evidence also improves the system. If the drill found a missing IAM role, the fix is the point. If the restore needed three undocumented Slack messages, those messages become runbook steps or automation. Evidence without backlog is just paperwork that ages.
A restore drill mid-size teams can actually run
You do not need a war-game studio. You need a repeatable habit permanent staff can run without the person who wrote the first backup policy sitting in the room.
In practice, a workable pattern looks like this.
- Scope one critical service and state the RTO and RPO in plain numbers with an owner.
- Restore into a production-like target with real identity and networking patterns, not a toy namespace that skips half the stack.
- Exercise the data path that matters. Prefer a recent production-shaped restore over a decade-old sample that flatters the tool.
- Run application smoke checks after the bytes land, covering login, write path, read path, and one dependency that usually breaks first.
- Capture elapsed time, blockers, and runbook gaps in a short report with an owner and a follow-up date.
- Rotate who leads the drill so the knowledge is not trapped in one head.
Yet keep the blast radius honest. A drill that risks customer data is not bravery. Use isolation, scrubbing, or a dedicated recovery account where needed, and document those constraints in the same place as the happy path.
When fantasy DR becomes expensive
Fantasy DR feels cheap until the first real outage. Then you pay in extended downtime, rushed decisions, and trust you cannot buy back with a status page. Mid-size teams feel it faster because the same few people own backups, restores, and the customer conversation.
Instead of waiting for that night, treat restore drills as a product surface. Own the objectives, practice the path, and keep evidence that operators can use. In the end, ask the uncomfortable question while everyone is calm enough to answer it.
If the primary region for your most important service disappeared before lunch tomorrow, would you restore from a tested path with a named owner and a clock, or would you discover that the green backup job was the easy part?
TL;DR
- Green backup jobs prove intake. They do not prove a working restore under time pressure.
- RTO and RPO need named owners who can start a restore and accept the restored state.
- Drills fail when they use toy data, skip identity and networking, or stop when files reappear instead of when the service can take traffic.
- Keep short, timestamped drill evidence with duration, gaps, runbook version, and follow-ups.
- Run scoped, production-like restore drills mid-size teams can repeat without heroics.
- Ask whether tomorrow's outage meets a tested path with an owner, or a binder that never left the shelf.