5 Downtime Lessons from Outer Space
CIO Office Mike Hicks ThousandEyes Principal Solution Analyst CiscoI recently consulted on Splunk’s The Hidden Costs of Downtime report. Although I knew downtime was a problem for enterprise companies, I was shocked it was a $600 billion problem. Now I’ve translated my out-of-this-world experience into five lessons to help CIOs navigate the complexities of downtime. Let’s blast off.
Lesson 1: Each component operates in its own orbit
At the European Space Agency, I managed the integration of satellite and ground systems, ensuring reliable data flow to partners like NASA. I spent a lot of time fine-tuning our applications and the network so they worked together seamlessly.
But even on Earth, everything operates in its own orbit. No matter the sophistication of your AI, automation, or operational oversight, system failure is inevitable. According to The Hidden Costs of Downtime, organizations face an average of 60 downtime incidents each year, originating throughout the entire technology stack. These breaks are rarely clean or predictable. They can be triggered by subtle timing mismatches or integration friction. And much of the time, they stem from interactions between systems rather than isolated components.
That’s why successful resilience strategies will map the entire dependency chain. Doing so helps identify potential failure points and reveal critical links between systems and services. By understanding how these systems influence one another, CIOs can better anticipate and manage the complexities of their interconnected infrastructure.
Lesson 2: Minimize the blast radius
Left unchecked, failures rarely stay in the system they started in. A wide blast radius doesn't just mean more damage. It means more people, more alerts, and more competing priorities to sort through before you can even start making decisions.
I saw this firsthand at the European Space Agency, where our network was made up of autonomous systems connected to governments around the world. Adding a static route for a local purpose — a seemingly minor action — ended up getting picked up and redistributed into our Border Gateway Protocol (BGP).
The result was a classic case of asymmetric routing. Outbound traffic to NASA continued to follow the intended high-capacity path, but the return traffic was forced through a single, low-bandwidth link. Because the traffic appeared normal at the source, identifying the root cause was incredibly difficult. We had to trace the full end-to-end path to realize that a single configuration change had created a massive, system-wide bottleneck.
But if you build your systems with containment in mind, you can stop a minor glitch from turning into a major outage. Here are a few ways to keep things isolated:
- Avoid simultaneous rollouts across your environment. Staggering updates means that if a patch fails, the issue stays contained to a small segment instead of taking down your entire network.
- Always have an automated "undo" button. If a deployment goes wrong, being able to roll back in seconds is the fastest way to minimize downtime and keep your systems stable.
Lesson 3: Ready mission control
The "Houston, we’ve had a problem" scenario isn’t just for space exploration; it’s a reality for every enterprise. Effective incident management relies on having clear owners with the authority to make calls, escalate when necessary, and get things fixed. Just as Mission Control relies on a pre-established command structure, your organization needs a clear authority chain to manage incidents from start to finish.
At the European Space Agency, accountability was everything. Every domain — like propulsion, range safety, weather, telemetry, and communications — was assigned a designated owner. These leads would complete rigorous pre-checks and provide explicit approval before the countdown could begin. No one assumed silence meant clearance, and only one person, the launch director, had the authority to call a hold.
By defining these roles in advance, your teams can operate with the precision of Mission Control rather than the chaos of improvisation. This will allow for faster, more decisive action when it matters most.
Lesson 4: Have a disaster recovery plan that includes “do nothing”
Disaster recovery plans can’t cover every contingency. For example, in all my years reviewing these plans, I never saw one that prepared for a sudden, global shift to remote work like we saw during the pandemic.
Start by determining how much downtime your organization can realistically afford. Manage risk by identifying your most critical services and setting clear loss thresholds. Pinpoint the bottlenecks that could cause major outages so you can focus your defenses where they matter most.
Your disaster recovery plan should be built around business continuity. That might mean running a skeleton setup or shutting down non-essential services to keep core operations running. It’s all about balancing the cost of perfect uptime against what the business actually needs.
Good planning also means looking at a wide range of risks — from system bugs and vendor issues to single points of failure in the cloud. Security threats, particularly malware, require their own response strategies. Some institutions, like hospitals and universities, “go dark” as a formal recovery method.
I've previously written about NASA’s response to the 2023 ISS communications outage, which is a perfect example of this in action. A power failure during upgrades at Johnson Space Center cut all contact with the ISS, triggering the first-ever emergency use of backup comms. Even with a redundant facility available, flight controllers chose to remain at the primary Houston site because its power and cooling systems remained stable.
That was a judgment call, not a default. Relocating carries its own risks and delays. Once the team understood the issue, they realized staying put was the faster path to resolution. It's a useful reminder that a “do nothing” option in a disaster recovery plan isn't the absence of a decision. It's a decision that’s earned by genuinely understanding the failure in front of you.
As NASA showed, sometimes floating in space — even while repairs are underway — is the soundest option.
Lesson 5: Curb repeat offenses with post-incident reviews
At the European Space Agency, our active command link to a geostationary weather satellite would seemingly drop every Friday afternoon. Without an active command link, the satellite would inevitably drift. Recognizing this risk, our flight controllers moved quickly to reconnect and stabilize it before it lost its position.
Eventually, we traced the issue to the Fiber Distributed Data Interface (FDDI) network, where non-essential traffic was clogging the lines and blocking critical commands. Clearing that traffic restored the command link and our control of the satellite.
But restoring service isn’t the end of your downtime story. It’s the beginning of the real work. Running a thorough post-incident review is the best way to uncover the underlying issues that led to the problem in the first place.
So instead of treating failures as one-off events, use them as diagnostic tools to build preventative measures that strengthen the entire environment. This might include automating manual tasks that are prone to error, improving monitoring to catch early warning signs, or simplifying complex dependencies that create hidden risks.
In outer space, we couldn’t afford to make the same mistake twice. Improving your architecture is the only way to ensure that the same problem doesn’t happen again.
A new frontier for resilience
When viewed from space, the world appears as one seamless, interconnected system. CIOs should view their IT infrastructure through this same holistic lens. The lessons I learned from my time at the European Space Agency show that preparation, clear command structures, and rigorous post-incident reviews are the cornerstones of a robust infrastructure. But ultimately, true resilience is built when an organization uses its failures as catalysts for lasting improvements.
Ready to launch your digital resilience strategy? Download The Hidden Costs of Downtime report for expert insights on containing system failures and protecting your operations.