Back to Insights
InfrastructureYesterdayJustin Pennington

Restore Day: Prove Your Backups Work Before You Actually Need Them

Restore Day: Prove Your Backups Work Before You Actually Need Them

Almost every business we talk to has backups. Very few have restores. The difference matters, because a backup is a promise and a restore is the proof. Green checkmarks in a backup dashboard tell you a job completed. They tell you nothing about whether the data inside is usable, whether anyone knows the decryption key, or how many hours it would take to get your ERP serving orders again.

The day you find out is always the worst possible day. So pick a better one. Schedule a restore drill — call it Restore Day — and treat it like a fire drill with a calendar invite.

Two Numbers to Agree On First

Before you test anything, leadership needs to answer two questions in plain language.

How much data can we afford to lose? If your last clean copy is from midnight and the system dies at 4 p.m., you've lost a day of orders, invoices, and support tickets. That tolerance is your recovery point objective.

How long can we be down? An hour? A shift? Three days? That's your recovery time objective.

These aren't IT numbers. They're business decisions with cost attached — tighter targets mean more expensive architecture. Get an owner or GM to say the numbers out loud, then check whether your current setup can actually hit them. In our experience, the gap between the assumed answer and the real one is where most of the risk lives.

What Restore Day Actually Looks Like

Block half a day. Pick a recovery date at random rather than the most recent backup — you want to know whether last Tuesday is recoverable, not just last night. Restore into an isolated environment, never on top of production.

Then validate like an operator, not an engineer. "The server booted" is not success. Success is a person from finance opening the restored ERP and confirming that last month's closed invoices are there, that inventory counts match, and that the reports run. Success is someone from sales confirming the CRM has the right contacts attached to the right accounts.

Time everything. Write down when you started, when the data was available, and when someone confirmed it was correct. That elapsed time is your real recovery time objective, and it is almost always longer than the one in the plan.

The Things Restore Day Usually Exposes

Drills fail in boring, fixable ways. The most common findings:

  • Credentials and keys nobody has. The backup is encrypted and the key lives in a former employee's password manager, or the hosting account has one owner who's on vacation.
  • Half the stack isn't backed up at all. The database is covered, but not file attachments, custom code, environment variables, DNS records, or integration configurations. You restore the data and discover none of the connections work.
  • The SaaS blind spot. Your cloud vendors protect their infrastructure. They generally do not protect you from your own bad import, a deleted record, or a departing admin. Hosted ERP, email, CRM, and file storage often need their own export or third-party backup — check the terms rather than assuming.
  • No clean copy. If every backup is reachable from the same account with the same credentials as production, ransomware reaches all of them. You want at least one copy that is immutable or offline.
  • Nobody knows the order. Systems have dependencies. Bringing the website up before the ERP or the identity provider can create more mess than downtime.

Write the Runbook While You're in It

The most valuable output of Restore Day isn't the restored data — it's the document you produce while doing it. A short runbook: which systems come back in what order, where the credentials live, who to call at each vendor, and the specific checks that confirm a system is genuinely healthy.

Keep it somewhere reachable when your systems are down. A runbook stored only in the file server you're trying to restore is a joke you only get to hear once. Print it, or keep a copy in a separate service.

Make It a Rhythm, Not a Heroic Event

Once a year is a start. Twice a year is better. And any time you make a major change — a new ERP module, a migration, a new integration, a hosting move — the drill should follow within a month, because that's when coverage gaps get introduced silently.

Rotate who runs it, too. If your recovery depends on one person's memory, you haven't built resilience; you've built a single point of failure with a pulse. The drill should be survivable by whoever is in the building.

None of this is glamorous work. It doesn't ship a feature or launch a campaign. But it's the difference between a bad afternoon and an existential event, and it's cheap compared to what it protects.

If you're not sure what your real recovery time would be — or you suspect the answer is "we'd find out the hard way" — that's a conversation worth having before the next incident. Infraxio helps operators pressure-test the infrastructure underneath their ERP, integrations, and web platforms, and fix the gaps a drill uncovers. Reach out and we'll walk through your setup with you.