Restore or It Didn't Happen
When I sat down to write this post, I tried to think of a time in my career that illustrated the point well. I grew a tad frustrated. I could not think of a situation where things went badly and forced me to change my perspective, my process, or my tools, which is how most of the stories on this blog go.
Then I realized I finally have an example of what I am trying to achieve with this blog. I listened to advice from someone who had already lived it, so I never made this particular mistake myself. Now it is my turn to pass that on to you, my dear reader.
I have touched on this many times, usually as a side comment while writing about something else, but I never addressed it directly. So this post is about backups. And, of course, restores.
Many moons ago, an older guy helped me with a few consulting jobs. At some point he told me a horror story.
His company had a policy. Before upgrading any system, take two tape backups of everything. Two copies. Belt and suspenders.
On one engagement, he decided to take a third. Not because the policy asked for it. Because he did not have much experience yet, and he wanted to practice restoring. So he made one extra backup for his own education, and he tested the restore.
You can probably guess where this is going.
Both official tapes got mangled. The only good copy was the one he had made for practice. His extra backup saved the company.
That story stuck with me for the rest of my career, though the lesson was never about the tapes. It was about what the word “success” means.
When a backup job reports success, it is telling you exactly one thing: the process ran and did not complain. It is not telling you the data is there, or that it is complete, readable, or restorable in a reasonable amount of time. Backup software has bugs too. Tapes get mangled. Disks fail. Someone changes a path and the job happily backs up an empty directory every night.
A backup is a hypothesis. The restore is the experiment.
Nobody wants backups. They want restores. The backup is the cost you pay. The restore is the product.
A while ago I wrote about a client whose backup was optimized for upload speed. It was a solid script and a reasonable choice. But the backup runs at night, unattended, and nobody watches the progress bar. What matters is how long it takes to get customers back online, and how that time feels to them. You only find that out by restoring.
I am a systems guy at heart, so this is advice I take myself.
At home, I often browse the snapshots on my ZFS system to check that things are where they should be. Every few months, I pull my cloud backups down into a temporary directory and check their structure. And every backup reports to a monitoring system, so if a job runs but does not succeed, I get an alert instead of a surprise.
To clients, I say the same thing every time. Restore your backups monthly. Quarterly at the very least. Know when you last restored, how long it took, and who besides you knows how to do it.
Now, the exceptions.
A full restore drill is not free. It takes time, storage, and sometimes a maintenance window. It makes no sense to spend a million dollars protecting a system whose failure would cost a hundred thousand. Restoring one customer, one database, or one folder catches most problems. Automated checks catch more. And some data is simply not worth the trouble.
But “we never tested it” is also a decision. It has a cost. Make it on purpose.
That older consultant made a third copy because he did not trust himself yet. Ironically, that humility is what saved the day.
A backup is a promise. A restore is proof.
If you are not sure your backups would come back when you need them, I can help you find out before it matters. Here’s what I offer.