Rapidly Control AI Agent Risk Before It Hits the Business

Observability Kashyap Merchant

Key takeaways

  1. An AI agent can complete every technical step successfully while still making a costly business mistake, like approving a $1,200 refund for a $699.98 purchase.
  2. Splunk Agent Observability tracks agent decisions and tool calls, helping teams identify exactly where a workflow failed and verify that fixes actually work.
  3. Cisco AI Defense scans AI systems for security weaknesses and blocks malicious prompts or manipulated tools before they cause harm in real time.

AI agents are rapidly moving into production across industries, taking on work that ranges from answering customer questions and processing claims to supporting employees and acting in business systems. The value is clear: faster service, greater scale and less manual work. The risk is harder to see: an agent can complete every technical step successfully and still make the wrong business decision.

Consider what this looks like in a retail support workflow. A shopper asks an AI agent to refund two headphones purchased for $699.98. The agent confidently initiates a $1,200 refund while every underlying service continues to report healthy. Technically, the workflow succeeded. For the retailer, it created a $500 over-refund, a reconciliation exception and a control failure that conventional health metrics may never flag.

risk-1.png

Figure 1: A refund agent workflow can complete successfully while producing the wrong business outcome: a $1,200 refund against a $699.98 purchase creates a $500 over-refund.

One bad refund is a customer problem. The same pattern repeated across orders, payments, tickets and inventory becomes an operational and financial risk. Agents can act across these systems faster than teams can manually review each decision. When a tool call completes, infrastructure success does not guarantee that the customer or the business received the right outcome.

This is no longer an edge case for a distant future. Cisco’s The State of AI Security Report 2026 found that 83% of organizations planned to deploy agentic AI, but only 29% felt prepared to do it securely. The State of Eval Engineering Report found that 84.9% of production AI teams had experienced an incident in the previous six months, with missing or failed guardrails accounting for 62% of those incidents.

The numbers make the urgency clear, but they do not explain what a team should do when an agent goes wrong. A technically successful workflow can still produce the wrong business result, and a manipulated workflow can still look complete. Teams need to connect what the agent did, whether outside influence changed the decision and what happened to the business process.

A Successful Transaction Can Still Be a Failure

An AI agent does not simply answer a question. It can retrieve policy, reason across company data, choose a tool and change a real system before a person sees the result. That autonomy is what makes agents useful. It also makes a bad production outcome harder to explain.

Conventional monitoring can tell the retailer that the support service was available, the tool calls completed and latency stayed within its target. Those facts matter, but none of them answer the question Finance and Customer Support will ask: why did a $699.98 purchase produce a $1,200 refund request?

The same symptom can come from different causes. The agent may have used stale context, reasoned incorrectly or reached a tool with no validation. It may also have encountered prompt injection, a poisoned tool description or another malicious input. The business sees one bad outcome; the teams responsible for reliability and security need different evidence to fix it.

For an unreliable business decision, teams must evaluate the outcome against a business expectation, preserve the agent session, isolate the failed model or tool step and verify the workflow after a change.

For malicious manipulation, teams must understand the AI assets and trust boundaries, validate realistic attack paths and enforce policy before an unsafe response or tool action completes.

The Observability and AI team use Splunk Agent Observability traces to follow the order lookup and refund tool calls. A Refund Amount Mismatches Signal groups similar failures and points to the missing validation. The team enables a refund compliance policy and reruns the same request. This time, the agent processes the correct $699.98 refund.

Splunk Agent Observability turns that investigation into a repeatable operating loop. The application emits sessions, traces and spans through the Agent Observability SDK, supported framework integrations or OpenTelemetry. Those records preserve prompts and responses, model and tool calls, timing, errors and context.

Evaluate each run against the expected outcome. The team defines evaluators and attaches the receipt total and proposed refund to the trace. The first $1,200 request fails because the purchase was $699.98. The Tracing view places that result beside the inputs, outputs and operational metrics.

risk-2.png

Figure 2: Splunk Agent Observability tracing view brings sessions, traces, spans, evaluator results and operational metrics into one place so teams can compare correct and incorrect runs.

Turn repeated failures into an investigation signal. Agent Observability groups the affected spans, traces and sessions into a Refund Amount Mismatches Signal. It summarizes the pattern, identifies the failed receipt comparison and recommends validation before the refund is processed.

risk-3.png

Figure 3: Splunk Agent Observability with refund amount mismatch signal connects the affected telemetry to a root-cause summary and a concrete recommendation for the workflow

Follow the agent through its tools and dependencies. The Agent graph shows the path through refund requests, support tickets and order lookup. The investigator can trace the mismatch to the step that selected the order or invoked the refund tool.

risk-4.png

Figure 4: Splunk Agent Observability Agent graph maps the refund workflow across the agent and its tool calls, including refund requests, tickets and order lookup.

Apply a targeted control and verify the next run. The team enables a refund compliance policy at the post-model stage so the proposed amount is checked before the action completes. When the same request is run again, the agent returns the correct $699.98 amount, giving the team evidence that the control changed the business outcome.

risk-5.png

Figure 5: Splunk Agent Observability Controls view shows the refund compliance policy enabled for the Agent Stream alongside other reusable controls.

The team now knows which expectation failed, which workflow step caused it, which control changed the behavior and whether the fix held.

How teams deploy it. Teams instrument agent workflows with supported SDKs, framework integrations or OpenTelemetry and route the telemetry to an Agent Stream. Splunk Agent Observability is available in Splunk Observability Cloud and as an on-premises deployment.

Cisco AI Defense Turns an Exposed AI Boundary Into Enforceable Policy

The retail refund scenario demonstrates an unreliable business outcome. A different risk appears when an agent follows a malicious instruction exactly as designed. Cisco AI Defense gives Security a control loop for that path by mapping the AI system, validating how its boundaries can be abused and enforcing runtime policy.

Consider an enterprise assistant that calls an MCP tool. The tool appears legitimate, but its metadata contains a hidden instruction to place private data into the tool arguments. The agent may follow that instruction faithfully. The problem is malicious influence, not a simple quality error.

Discover the AI system and its trust boundaries. AI Defense inventories applications, models, agents, MCP servers, tools and connected data, then maps how those assets relate.

Validate realistic attack paths. Supply-chain assessment and automated red teaming test models, applications, agents and tools for prompt injection, poisoned behavior, data exposure and unsafe actions. The result shows which technique worked and where a control is missing.

risk-6.png

Figure 6: Cisco AI Defense Validation view shows attack-blocking performance, top threats and attack techniques so Security can see where an AI application needs stronger guardrails before production

Protect the live interaction. AI Runtime Protection inspects prompts, responses and agent-tool interactions against policy. It can monitor, sanitize, deny or block based on the rule and enforcement design, while preserving the triggering content and action for investigation.

risk-7.png

Figure 7: Cisco AI Defense Events view connects blocked and monitored interactions with the triggering conversation, matched privacy and prompt-injection rules, enforcement point and event details.

The event tells Security which application and interaction triggered policy, what the control did and why. The team can refine the rule without treating every AI interaction as an undifferentiated alert.

How teams deploy it. Use the AI Defense Inspection API when an application needs direct control over what is inspected and how it handles the verdict. Use AI Defense Gateway or Multicloud Defense when policy should be applied centrally across applications and model endpoints. Deployment choices include cloud, customer VPC, on-premises and AI POD environments.

Bringing the Two Operating Loops Together

Splunk Agent Observability and Cisco AI Defense support distinct but complementary operating loops. AI product, platform and engineering teams use Agent Observability to define a correct business outcome, inspect the session that produced it and verify a targeted control. Security teams use AI Defense to discover AI assets, validate attack paths and enforce runtime policy.

Agent Observability needs instrumented workflow telemetry that preserves what the agent saw and did. AI Defense needs the protected application, policy and AI traffic passing through an inspection or enforcement point.

Product
What it receives
What teams get
Splunk Agent Observability
Instrumented sessions, traces and spans from the Agent Observability SDK, supported framework integrations or OpenTelemetry. Records can include prompts, responses, model and tool calls, timing, errors, token and cost metrics, and business context.
Evaluators score each run, Signals group recurring failures, and tracing plus the Agent graph isolate the responsible step. Teams apply a targeted control and rerun the workflow to verify the result.
Cisco AI Defense
A registered application, active policy, and prompts, responses or HTTP traffic sent through the Inspection API, AI Defense Gateway or Multicloud Defense. Discovery and validation add asset and attack-path context.
AI Defense returns a verdict and applies the configured action. A triggered rule creates an AI Event with the application, traffic direction, rule match and enforcement result.

For Agent Observability, telemetry is the evidence and evaluators turn it into measured outcomes. For AI Defense, inspected traffic and policy are the input, and an AI Event records a detected violation.

The same customer complaint can lead to different actions. If the refund agent used stale order context with no evidence of manipulation, the team should repair the workflow and validate the amount against the receipt. If a poisoned tool changed the agent's intent, Security needs to investigate the asset and enforce a runtime policy. If an attack exploited the missing refund validation, both teams have work to do.

When the same incident needs both teams, the evidence should answer four questions: what changed for the customer, which agent decision caused it, whether untrusted input influenced that decision and whether the next run behaved correctly. The demonstrations use separate applications and environments; they do not imply an automatic event handoff today.

A Practical First Step

Choose one production workflow where a wrong or manipulated action has a meaningful consequence. Define the expected outcome, instrument the session, map the AI boundary and select one realistic attack objective. Run a correct interaction, an unreliable decision and a manipulated interaction before scaling the approach.

Summary

AI agents create a production question that application monitoring and security alerts cannot answer on their own: did the agent fail at its job, or did someone manipulate it? Splunk Agent Observability shows where behavior departed from the intended business outcome. Cisco AI Defense shows whether an AI boundary was vulnerable or manipulated and whether policy stopped the interaction. Look for more news from Splunk and Cisco on how these capabilities will come together to support agent resilience end to end.

Resources

Related Articles

Staff Picks for Splunk Security Reading June 2021
Security
5 Minute Read

Staff Picks for Splunk Security Reading June 2021

Splunk Security Content for Threat Detection & Response: September Recap
Security
2 Minute Read

Splunk Security Content for Threat Detection & Response: September Recap

Splunk's September ESCU update: New security content & analytics for robust threat detection. Covers Cisco ASA, ArcaneDoor, diverse malware, and Office365 Copilot activity.
Supercharge Your SOC Investigations with Splunk SOAR 6.4
Security
4 Minute Read

Supercharge Your SOC Investigations with Splunk SOAR 6.4

Splunker Nick Hunter explains how to integrate Cisco Talos threat intelligence, leverage Azure scalability, and streamline investigations.