The Hidden Cost Of Downtime In Communications And Media
Security Gaurav GuptaDowntime in communications and media is no longer a narrow IT issue. It is a business risk that reaches revenue, regulatory exposure, customer experience, and brand reputation. For organizations that operate always-on services, even short disruptions are immediately visible to end users.
That pressure is getting harder to manage. Splunk partnered with Oxford Economics to survey 2,000 executives across Global 2000 companies in APAC, the Americas, LATAM, and EMEA, spanning sectors including financial services, public sector, manufacturing, and communications and media. The findings show a broad shift: downtime is now an organization-wide challenge, not a siloed engineering problem.
For communications and media leaders, the message is clear. The cost of downtime is rising fast, the operational burden is deeper than many teams measure, outage sources are increasingly unpredictable, and AI is changing both the opportunity and the risk profile. Here’s what this means and how organizations can respond.
Downtime Has Become A Bigger Business Crisis
The scale of the problem is growing quickly. Across the Global 2000, the aggregate cost of downtime rose from $400 billion to $600 billion in two years. That is a 50% increase in a short period, and it shows how outages have become a major drain on enterprise value.
Frequency is also a concern. Organizations now face an average of 60 downtime or service degradation incidents each year. That is more than one incident per week, which leaves many teams operating in a constant cycle of response and recovery.
For communications and media organizations, the trend is even sharper. Costs have increased 86% over the last two years. In an industry where "always-on" service is the baseline, even brief interruptions are immediately visible to users, impacting brand loyalty and the bottom line.
Why Communications And Media Feel The Impact First
The financial burden in communications and media shows up in several places, but lost revenue remains the largest line item. It nearly doubled from $49 million to $95 million. In practical terms, downtime can mean lost subscriptions, failed transactions, and interrupted service delivery.
There is also a broader shift in the threat landscape. Regulatory fines, legal settlements, and ransomware-related costs are becoming more prominent for telcos, especially where service-level obligations are breached. This means downtime is no longer just a technical failure, its a business risk. Business leaders are responding to that pressure. In the research, 86% of leaders say downtime is a higher priority today than it was 12 months ago. Another 91% cite rising costs and customer expectations as key drivers. That combination makes downtime a core business capability issue, not simply an uptime metric.
The Costs You See, And The Costs You Usually Miss
Most outage discussions focus on direct impact. Teams measure service interruption, delays in root cause analysis, and immediate revenue loss. Those factors matter, but they do not capture the full operational burden.
A large share of the damage is harder to quantify. Leaders report increased customer support demand, greater customer frustration, long-term brand damage, and the strain of ransomware and regulatory exposure. These hidden costs often sit outside traditional incident metrics, even though they can shape business performance over time.
Operationally, outages also pull organizations into an all-hands response model. Critical teams are redirected into firefighting, which creates bottlenecks and reduces time available for innovation. Productivity falls when teams spend too much time trying to identify what broke, who owns the issue, and how to restore service. Over time, constant pressure contributes to fatigue and burnout.
Treat Resilience As A Business Decision
If downtime affects revenue, compliance, and customer relationships, resilience needs to be managed at the same level. That starts with changing the language organizations use to describe the problem.
Instead of focusing only on server uptime or latency, teams need to connect outages to business outcomes such as revenue loss, customer churn, and regulatory exposure. That context helps stakeholders understand why resilience deserves broader attention and investment.
It also means bringing downtime risk into board-level and executive decision-making. Resilience cannot be an afterthought that appears only after an incident. It needs to be considered in major business decisions and new product launches.
Cross-functional alignment matters too. Security may focus on protection, IT on uptime, and finance on cost. When those teams work from a shared resilience perspective, organizations can move toward a more unified operating model.
Build Systems That Reduce Firefighting
A resilient organization starts by accepting a practical reality: downtime cannot be eliminated completely. It is not possible to prevent or plan for every scenario where systems may fail. That makes safer system design essential.
One useful approach is to build systems that are safe by default and supported by guardrails. Guardrails help prevent accidental misconfigurations from becoming major outages. They also reduce the need for manual intervention, which gives engineers more time to focus on innovation instead of emergency response.
Standardization is another key step. Many organizations now work across microservices-led or composable architectures, where teams choose their own tools, processes, and deployment methods. That flexibility can create heterogeneous environments where teams operate in isolation and work from different assumptions. Standardizing change management and deployment practices gives teams a common playbook and reduces confusion across the environment.
Faster Detection Changes The Economics Of Outages
Detection speed is one of the biggest factors in outage response. The most exhausting part of an incident is often the time spent figuring out what failed, where it failed, and who needs to respond.
That is why mean time to detect and mean time to respond remain critical. Systems that can quickly identify what has gone wrong, where the issue originated, and what the likely root cause is can reduce both MTTD and MTTR. Faster guidance also helps teams deploy new changes with more confidence.
Shared context matters here. Outages often persist because teams are looking at different versions of the truth. Time is lost reconciling fragmented information across systems and teams. Connecting data from different sources through a data fabric-style approach can help break down silos, give teams multiple perspectives from the same underlying data, and accelerate problem resolution.
Outages Can Come From Anywhere
Even with modernization, human error remains the most common cause of downtime. That alone makes resilience a broad operational challenge. But the issue does not stop there.
Communications and media organizations are also seeing a major increase in phishing attacks. Third-party dependencies, including cloud providers, cloud workflows, and source code dependencies, are another major contributor to costly incidents. In a highly interconnected digital ecosystem, the source of failure is increasingly difficult to predict.
That changes the defensive model. Traditional perimeter-based approaches are not enough when disruptions can originate across internal systems, external providers, and human workflows. Organizations need to assume that failures will happen and build systems that can absorb and respond to them more effectively.
AI Can Improve Resilience, But It Also Adds Risk
AI is now central to how many leaders think about resilience. Many leaders are turning to AI to help address slow outage response, fragmented data, and the need to improve MTTD and MTTR. Used well, AI can accelerate insight and help teams move faster.
At the same time, many of those same leaders report that AI is contributing to new incidents. AI-led failures are becoming a more frequent issue for teams to manage. That creates clear tension: the technology can improve resilience, but it can also introduce new complexity.
The practical takeaway is straightforward. AI should accelerate operations, not replace human judgment. Teams need to validate what AI systems are doing, understand why a recommended remediation is being made, and put guardrails in place to retract or correct errors quickly. Human oversight remains essential if AI is going to strengthen resilience instead of creating new failure paths.
Four Practical Takeaways For Leaders
Communications and media organizations are dealing with a more expensive, more frequent, and more complex outage environment. A resilience strategy needs to reflect that reality.
Here are four actions that stand out:
These steps do not remove every risk. They do help organizations move from reactive troubleshooting to a more predictable and collaborative resilience model.
- Treat downtime as a business risk, not just an IT issue.
- Design systems for people, with safe defaults and guardrails that reduce operational strain.
- Make detection and root cause analysis a shared effort, supported by common data context.
- Use AI to accelerate insight, while keeping humans in control of decisions and remediation.
Building A More Resilient Operating Model
The goal is not only to reduce outages after they happen. It is to change how organizations prepare for them, detect them, and recover from them. That shift matters in communications and media, where service continuity is tied directly to revenue, customer trust, and regulatory exposure.
A resilience-based strategy helps teams spend less time in crisis mode and more time building, improving, and innovating with confidence. It creates a stronger foundation for managing the unexpected without overwhelming the people responsible for keeping systems running.
The hear more on how Splunk is helping organizations improve resilience, reduce detection and response times, and bring shared context to outage management in Communications and Media organizations – join our upcoming webinar or read the report
Related Articles

Introducing Synthetic Adversarial Log Objects (SALO)

Staff Picks for Splunk Security Reading August 2022
