splunk background

evaluation

Agent and model evaluation

Score the correctness, safety, and tool use of every agent interaction in production, not a sample.

Free edition Get Observability Coud free for up to 15 hosts.
Take a guided tour Got 5 minutes? Get a look at how it works.
pd-ao-eval-hero-d

Find issues before they reach your customers

When the judge is a small model instead of a frontier LLM, grading every interaction becomes affordable. No sampling, no blind spots.

Score every agent interaction, not a sample

Rather than relying on a frontier LLM, Luna is a small model that makes it more affordable to score every request in production instead of a lucky slice. Grade correctness, safety, and tool use on 100% of traffic.

Tailor evaluators to your specific criteria

Out-of-the-box metrics rarely match a specific domain's bar for a good answer. With Autotune, the evaluator adapts to your standard instead of a generic one, so the scores reflect how your team actually judges the work.

Assess quality next to latency, errors, and cost

Every evaluation score sits on the same trace as latency, errors, and token cost, so you see not just that an answer was wrong, but also what it cost and how long it took. Get one view of quality and performance, not four disconnected tools.

.Conf 26 promo image

Get hands-on with Splunk

Join us September 14–17 in Denver, CO for an immersive learning and networking event.

Register for .conf26

features

Ensure agents are right before and after you ship

Explore the documentation
continuous-auditing continuous-auditing

Score the things agents get wrong

Grade action completion, tool selection, hallucination, and RAG adherence, and more. Use out-of-the-box metrics or write your own, and view scores next to latency, errors, and cost on one trace.

Blog Blog

Tune an evaluator inline without a data scientist

Correct a wrong score right on the trace and Autotune updates the evaluator from it. Eliminate the need for labeling projects or prompt engineering handoffs. From about 5 examples, accuracy lands around 95%, with lifts of roughly 10 to 17 F1 points.

real-time-performance real-time-performance

Run experiments and A/B tests

Compare prompts, models, and agent versions against the same evals before they reach production. See which change actually improved correctness and tool use, instead of guessing from spot checks.

reduce-cyber-security-threats reduce-cyber-security-threats

The evals you trust become guardrails

An evaluation that scores reliably in testing becomes a control in production. Promote it to a runtime check that blocks the bad action before the tool fires.

See guardrails

 

 

Resources
Explore more from Splunk

What Is Eval Engineering?

Read the blog

Agent evaluation FAQs

Agent evaluation scores what an agent actually did, from action completion and tool selection to hallucination and RAG adherence, so you can tell whether it behaved correctly instead of guessing from spot checks.

Sampling leaves most traffic ungraded, and the failure that reaches your customer is usually in the part nobody looked at. Because the judge is a small model rather than a frontier LLM, you can afford to score every interaction, not a lucky slice.

Yes. With Autotune you correct a handful of scores directly on the trace and the evaluator adapts to your standard, with no labeling project or prompt-engineering handoff.

An evaluation that scores reliably in testing can be promoted to a runtime check that blocks or steers the bad action before the tool fires, so there is no separate model of correctness to maintain.

Yes. Run experiments and A/B tests that compare prompts, models, and agent versions against the same evaluations, so you can see which change improved correctness.

Related products

Splunk Cloud Platform

Unify data, context, and action across every domain.

Learn more

Splunk Enterprise Security

Unified threat detection, investigation, and response for the agentic SOC.

Learn more

Splunk IT Service Intelligence

Predict and prevent IT issues with AI-driven service monitoring.

Learn more
Get started with Splunk

Evaluate every agent interaction and ship with confidence.

Request a demo
Explore free trials