observability

AI SRE

Stop guessing, start resolving with an agentic teammate to troubleshoot issues faster.

Free edition Get Observability Cloud free for up to 15 hosts.

How it works

Embedded AI and agentic support across the entire incident lifecycle

An agentic teammate that detects, troubleshoots, and fix issues 

Get extra assistance to ensure your systems are running as planned. If something goes wrong, AI SRE helps you move beyond static thresholds by automatically detecting issues. Dependencies are mapped, alerts are correlated, and  probable root cause is delivered along with a remediation plan that provides a step-by-step guide to get everything back up and running. 

Meets you where you are

Whether you're new to observability or a veteran, our built-in agentic AI capabilities help you reduce mean time to resolution (MTTR) and gain actionable insights from day one.

Free your time for what matters most

AI SRE helps you spend less time firefighting issues and more time focused on what matters most, so you can build what’s next.

Agentic observability for the new speed of business

Explore the documentation

AI-driven detection

Proactive detection and automated incident grouping

Automatically correlate related alerts into unified incidents to reduce noise and eliminate redundant investigations. This dynamic grouping provides a single, actionable view of service health, accelerating resolution, and helping teams focus on what matters most. 

pd-o-ai-sre-features-automatically-detect

AI troubleshooting agent

Let our agent troubleshoot for you

The AI troubleshooting agent automatically sifts through all metrics, log, and trace data, identifying whether your application or infrastructure is at fault, and surfacing the most likely root causes — all in plain language, within your existing workflow.

pd-o-ai-sre-features-trouble-shooting-agent

Remediation plan

Get to resolution faster

Move from investigation to resolution. AI SRE doesn't just tell you what broke; it generates guided, step-by-step remediation plans to fix the root cause. Execute fixes confidently with human-in-the-loop oversight, reducing downtime and preventing future outages. 

pd-o-sre-agentic-teammate-ani

AI Assistant in Observability Cloud

Get insights and answers in plain English

Easily extract insights from Observability Cloud and accelerate investigations using natural language. If you need more help, just ask the AI assistant.

Splunk MCP Server & agentic AI

Use Splunk capabilities in one unified MCP server

Leverage a secure interface to connect your local AI agents, LLMs, tools, and data with Observability Cloud data to build custom AI workflows and debug performance issues in production without leaving your environment.

pd-o-ai-sre-features-mcp-agentic-ai

We work with amazing customers.

See why the world’s leading organizations rely on Splunk.

Repay customer story Repay customer story

CUSTOMER STORY

Repay Pays it Forward with AI Assistant in Observability Cloud

With so many different systems with various endpoints, knowing it all is impossible. So it’s not just about efficiency but also identifying the unknown anomalies and getting insights from the data like a subject matter expert

Van Wolfe, VP of Platform Engineering at Repay
50%
faster triage
30%
Transaction latency reduced by 30%
Resources
Explore more from Splunk

5 Big Myths of AI and Agentic AI

Separate AI fact from fiction. Learn how agentic AI reshapes observability and security while helping teams work smarter and faster.

Read the e-book

AI SRE FAQs

Bringing together agentic and embedded AI across Splunk Observability Cloud, the AI SRE is an agentic AI experience spanning the entire incident response lifecycle including detection, troubleshooting, and remediation. This is an AI-native user experience including the AI Assistant in Observability Cloud and Splunk MCP Server that delivers in-context insights and frees teams to focus on what matters most and building what’s next, rather than manual troubleshooting.

AI SRE acts as an automated extension of your engineering team. Instead of forcing you to hunt through dashboards, logs, and traces across multiple screens, it accelerates resolution through a three-step workflow: 

  • Detect: It continuously monitors your environment, moving beyond static thresholds to automatically detect anomalies and group related alerts.
  • Troubleshoot: It instantly sifts through telemetry data to pinpoint the exact root cause, summarizing the issue and its impact in plain English.
  • Remediate: It generates a guided, step-by-step remediation plan. Your team can review, execute, or undo these steps directly within their existing workflow to restore service quickly. 

AI SRE is built into Splunk Observability Cloud and works seamlessly across modern, complex environments, including Kubernetes, microservices, and hybrid cloud architectures. Because it leverages OpenTelemetry, it can ingest and analyze metrics, logs, and traces from a vast ecosystem of supported integrations. Whether you are troubleshooting a cloud-native application or legacy infrastructure, AI SRE correlates data to pinpoint root causes. 

The primary benefit of AI SRE is drastically reducing mean time to resolution (MTTR) by reducing manual toil.

Key benefits include:

  • Faster triage and resolution: Turn hours of manual investigation into minutes of actionable insight, helping teams achieve up to 50% faster triage. 
  • Reduced alert fatigue: Automatically correlate related alerts into unified incidents to cut through the noise. 
  • Guided remediation: Move beyond just finding the problem to actually fixing it with AI-generated, step-by-step action plans. Increased productivity:
  • Free up your DevOps and SRE teams to focus on innovation and customer experience rather than firefighting outages. 

No, you always remain in control. Splunk AI SRE is designed with a "human-in-the-loop" approach while AI SRE’s troubleshooting agent automatically analyzes telemetry data, identifies probable root cause, and builds a step-by-step remediation plan. With an easy-to-use UI, it relies on your engineering team to review, approve, and execute the final steps. This ensures you get the speed of machine-assisted investigation without sacrificing governance or control over your production environments. 

Related capabilities

Application Performance Monitoring

Solve problems faster in monoliths and microservices by immediately detecting problems from new changes, confidently troubleshooting the source of an issue, and optimizing service performance.

Explore APM

Infrastructure Monitoring

Improve hybrid cloud performance with instant visibility and real-time alerts.

Explore Infrastructure Monitoring

AI Assistant in Observability Cloud

Get expert guidance in plain English to find and fix issues faster.

Explore AI Assistant

AI Observability

Observe the performance, quality, security, and cost of your AI stack.

Explore AI Observability
Get started

Experience the embedded AI in Observability Cloud for free.

Contact sales
Free trial