Lessons from a Successfully Failed Disaster Recovery and Failover Test
Conducted during a busy release weekend, the failover test exposed gaps not in the technology itself, but in coordination and communication. While production ultimately stayed unaffected, the situation quickly escalated as subcontractors werenโt aligned, assumptions didnโt match reality, and information didnโt flow when it mattered most.
We unpack how a well-intentioned test turned into a coordination challenge, where timing, dependencies, and unclear responsibilities created confusion across teams. Itโs a story about how resilience isnโt just about systems and infrastructure, but also about people, processes, and making sure everyone is on the same page โ especially when things are supposed to โjust be a test.โ
Listen now on Apple Music, Spotify, Deezer, Youtube or where-ever you get your panic attacks.

Welcome to another episode of IT Horror Storiesโtoday, we bring you an epic saga from the trenches of enterprise IT. This is not the story of a smoothly executed disaster recovery test. Nope, this is about the failover that, wellโฆ technically succeeded and failed at the same time. Grab your coffee, your stress ball, and some popcornโletโs get into what happens when everyone does things โby the bookโ and it all still goes sideways.
Setting the Scene
Jack Smith: Welcome back, everyone. This is Horror Stories with Jack Smith. Iโm Jack Smith and, across the table, I have Bob.
Bob: Hi guys and girls and everybody in between, however you like to roll. Everybody is welcome. Every single oneโฆ
Letโs get comfy. This time, we’re talking about a release and failover process that went so โby the bookโ that it left us laughing, crying, and desperately searching for someone to blame other than ourselvesโฆ and possibly our supplierโs suppliersโ suppliers.
Corporate Environments: Big Fish, Bigger Pond
It all happened about a decade ago (give or takeโwe try not to count those years too closely). Our protagonist, Bob, was working for a large enterpriseโthink way larger than your average small business. We’re talking about organizations that don’t just have suppliers, they have suppliers with suppliers, and sometimes you need a full map to know who called who.
The Cast of Characters:
- The Company: Handles its own app management and development. Architects, analysts, developers, testersโall present.
- Third-Party Devs and Testers: Because, you know, offshore is cheaper.
- An Infrastructure Supplier: Because who wants to own all those racks and blinking lights?
- Hosting/Data Center Provider: The suppliers’ supplier. You see where this is going.
This is a story from an environment where more people are involved in launches than a small nationโs moon program.
Quote Highlight:
โA tad bigger than a small business, indeed. And essentially the story here happened about 10 years ago. I tend to forget my age or I don’t want to remember my age.โ
Failover Planning: Herding Cats With Clipboards
1. Annual FailoverโโBecause Regulators Said Soโ
Every year, in this regulated industry, the company had to prove that its failover system, disaster recovery setup, and business continuity machines actually worked. If youโre just picturing flipping a switchโstop. Weโre talking :
- At least two data centers
- A planned failover to the backup center for a full week of real business work
- A โfail backโ to primary at the end (fingers crossed)
- Different server types: mainframe, AS/400, Windows boxes, a dash of cloud, and some SaaS
- Dozens (sometimes hundreds) of people babysitting releases becauseโฆ well, see above
Quote Highlight:
โSo, so far nothing out of, you know, ordinary. That sounds dangerous. No, indeed, indeed.โ
2. Release Weekends: The Ritual
- Four times a year: Rolling changes from UAT (User Acceptance Testing) to Prod
- Up to 300 people in the building on a Sunday
- Saturday = ITโs technical deploy, basic checks
- Sunday = Business teams swarm in and try to break stuff (er, validate)
- Every step mapped out, every checklist ticked
Agile? Kind of. Waterfall? Mostly. Budget? Letโs just say: big.
Double Disaster: Because One Isn’t Enough
So, this particular year, management had a bright idea:
โSince we already have all these people around for the release weekend, why not combine the disaster recovery test with one of those weekends? We get two birds with one stone. Cheaper for us!โ
What Could Go Wrong?
- Infrastructure provider: Letโs align their release weekend with ours
- Double activity weekend: Application failover and infrastructure partnerโs UAT-to-production moves at the same time
- Us: โIsnโt that asking for trouble?โ (Answer: oh yes.)
Critical Failure: โWorks for Me!โ
Start of the Weekend
Saturday morning. Caffeine. Cautious optimism. Everyone following the plan to the letter:
- UAT promotes are in motion
- Infrastructure partner is busy, but in their own โnon-overlappingโ world
- Failover to backup data center: check. All green lights.
And Thenโฆ Validation Time
Nothing. Works.
Not a single login screen. Not even the โCitrix is starting upโ spinner.
โValidation part, nothing worked. And when I say nothing, I mean absolutely nothing worked. We couldn’t even log in anymore. It was a Citrix environment. We could not log in into our own accounts anymore. Everything broke.โ
- 11pm. Everyoneโs tired
- Canโt login, canโt check servers, canโt run tests
- Incident call time: Us, our incident managers, infrastructure partners, a growing Zoom callโฆ you know the drill
The Realization Moment
Now, hereโs where the magic happens:
โLooks Fine to Us!โ
Infrastructure and hosting partners both look.
โAll our systems are green!โ
โWe can see the servers, network is up!โ
Butโฆ
We canโt even ping the backup data center.
No traffic gets through. Itโs as dead as a doorknob.
โLuckily, several people on our side do have some infrastructure knowledge. So we didn’t have the rights, but the infrastructure part gave us the rights so we could perform some network traces and stuff like that. And essentially we came to the conclusion that we had no network connectivity in the secondary data center.โ
Hosting Provider’s Turn
After two hours of hair-pulling conference calls, the hosting partnerโs tech eventually pipes up:โWhich data center are you trying to connect to?โ
- Us: โNumber twoโthe backup.โ
- Hosting provider: โOh. Uh. Hold on a second.โ
That, friends, is the sound of โOh shit.โ
Split Brain: When Systems Canโt Decide
So, it turns out:
- Hosting provider thought: Since we were active and infrastructure partner was active, why not ALSO do their disaster recovery test? They fail over the network from backup data center back to the primary.
So, in the end:
- One data center: All apps and servers ready to respondโbut no network.
- The other data center: Just networkโno apps, no services.
Both suppliers ran their failover scripts at the exact same timeโin opposite directions.
Quote Highlight:
โOuch. We had one data center with zero applications and services running with network, and we had another data center with everything running, throwing errors everywhere because we had no network. But still, you could claim that both failovers were a success.โ
The โNo Going Backโ Policy
To add insult to injury, both the company and the suppliers had a policy:
โA failover canโt fail. Once you start, thereโs no way to abort and roll back.โ
So, the backup had to be resurrectedโthere was no option to simply โtry againโ on another day.
The Recovery: No Rest for the Weary
Hot Potato: Who Goes First?
- Suppliersโ suppliers: โLet’s finish our disaster recovery from backup back to primary. When done, we’ll return the network and you can do your business.โ
- Us: Wait for them to finish, then start our mad dash through IT validations and business validation.
Time Lost
- Their tests wrapped up by early Sunday afternoon (~2pm). Only then could we start our actual work.
- That meant 12โ15 hours behind schedule. The precious sleep-time window for IT folks? Gone.
- IT validation rushed through in just a few hours.
- Business validation crammed in: 6pm Sunday evening.
Result:
- Final all-clear came at 6:30am Monday morning.
- Doors open for office at 7:30; business as usual starts at 8:00.
- Relief, exhaustion, and a sense of โthat could have gone a lot worse.โ
Lessons Learned: Communication Breakdown
Looking back, nothing in the internal plan was wrong:
- Roadbooks? โ๏ธ
- Checklists? โ๏ธ
- Stakeholder communications? โ๏ธ
But left-hand and right-hand at the supplier chain werenโt talking. And nobody had a holistic view.
Quote Highlight:
โIt was just a case of at the supplier side, left hand didn’t really talk to the right hand.โ
Main Causes
- Assumptions:
- Hosting provider assumed their disaster recovery was โlow impact.โ
- Infrastructure partner assumed the same.
- Missed Communication:
- Each group only told their next-in-line about big changes.
- Anything โnon-impactful to customersโ didnโt make it onto our change log.
- Size = Complexity:
- With big organizations, itโs easy for a critical memo to never reach the exact people who need it.
- Holy Change Boards:
- Yes, there were change boards. But if a change is considered low-risk, no notification is needed. Until it isn’t.
Failovers, Due Diligence, and Expensive Sleep Loss
The Magic Question
From that point on, Bob always demanded, before any major event:
โWho are your suppliers? Can I get confirmation from EVERYONE along the chainโprovider, infrastructure, hostingโthat nothing will be happening, no matter how low-impact it might seem?โ
Sure, most of the time, itโs just a few calls or emails. But itโs so much easier than losing a weekend, paying for a dozen hotel rooms, and living on vending machine food.
Resistance
Getting everyone to sign off isnโt always easy:
- Sometimes, risk acceptance comes from leadership: โWeโll just take the risk this time.โ
- In that case: โFine, just sign right here and say youโre aware.โ
9 out of 10 times nothing happens. The 10th? Well, it makes for an epic horror story.
Change Boards: Blessing and a Curse
- Not every little change should be communicatedโotherwise, everyone drowns in emails.
- But if it even touches the backup data center, network infrastructure, or overlaps with a failover weekend?
- Communicate. Please.
Conclusion: TL;DRs and Takeaways
So whatโs the real lesson of โThe Failover That Failed Successfullyโ?
- Talk. It doesnโt matter if you own the infrastructure or outsource it. If youโre tied to suppliers, and especially suppliers with suppliers, demand explicit โnothing will happenโ confirmation before doing anything critical.
- Donโt stack disaster recovery tests unless you have an absolute, air-tight chain of communication.
- Be wary of โlow impactโ changes. To you, itโs a checkbox. To someone else, it’s the trigger for a 14-hour crisis.
- Check your suppliersโ suppliers. And get them all in the (virtual) room.
- Never assume โno impactโ means โabsolutely no impact.โ
Final Quote Highlight:
โIf you have something extremely critical, validate the entire chain. Even if it’s just for your own peace of mind. It’s usually a few mails, a few calls and you can sleep tight and you can avoid these type of things.โ
Some Visuals and Checklists
Quick TDA template
graph TDA["You (Organization)"]B["Infrastructure Partner"]C["Hosting/Datacenter Provider"]D["Suppliers' Supplier"]A --> BB --> CC --> Dclick A href "mailto:your-it-team@example.com"click B href "https://supplier-team-portal.com"
Release Weekend To-Do List
- [x] Align release weekends and failover weekends
- [x] Confirm all change boards know about critical events
- [x] Send explicit โno activityโ requests to every supplier and sub-supplier
- [x] Schedule overlap check meeting for all parties
- [x] Build in buffer time (and sleep time)
- [x] Organize on-call rotations and hotel rooms as backup
- [x] Double-check assumptions, especially โlow riskโ ones
Wrap-Up: Why We Share These Stories
Some IT war stories are funny in hindsight. Some are painful. This oneโs a bit of bothโwith a solid โdonโt let this happen to youโ message for every ops engineer, project manager, or C-level exec who thinks, โLetโs just combine tasks to save money!โ
Final Words
You donโt always need a highly polished, expensive โlessons learnedโ write-up to stay safe. Sometimes all you need is to ask, and ask again. That, and maybe donโt do parallel failover drills unless someone signs in triplicate.
Thanks for reading. You are one of us.
โOn paper we did everything right. We had the road book. Everything was communicated, everything was validated. It was just a case of at the supplier side, left hand didn’t really talk to the right hand.โ
Stay safe. Communicate. And always check the backup data center.

Leave a Reply