splunk background

luna

Luna evaluation models

Purpose-built small language models that evaluate and guardrail every interaction, so you can watch 100% of traffic affordably.

Free edition Try it free for 14 days — no credit card required.
Take a guided tour Got 5 minutes? See how it works.
Luna evaluation models
bg-image

98% cheaper changes the economics

When a verdict costs about $0.15 per million tokens instead of $5.00, you stop rationing. Grade and guard every interaction, not a sample.

bg-image

The LLM-as-judge tax

Luna is built for one job, judging agent interactions. That single focus is why it can grade and guard 100% of traffic in real time, where a frontier judge is too expensive and too slow.

Too expensive to grade everything

A frontier judge can cost around $5.00 per million scanned tokens. At production volume that is too expensive to run on everything, so teams check 1 to 5% of interactions and hope the rest look the same. The ones you miss are the ones that hurt: the hallucination you never saw, the injection that slipped through.

Too slow to stop anything

A frontier judge can take about 3,200ms to return a verdict. By then the tool has already run, so the verdict can only describe what happened. It can audit, but it cannot guard. Luna returns a verdict in about 152ms, fast enough to block an action before the tool fires.

Deterministic, single-token scoring

Luna returns a verdict as a single token, so scores are fast and repeatable. Run the same input twice and you get the same answer, which is what makes a model trustworthy as a judge. And because Luna runs and trains inside your VPC, your data never leaves your firewall.

Tune Luna to your domain, no code required

Out-of-the-box metrics get you started, but your domain has its own definition of right. With Luna Studio, correct a handful of Luna's verdicts and it learns your standard, reaching around 95% accuracy on your own tasks. No labeling pipeline, no model training expertise. Review, correct, and ship a tuned metric, all from the studio.

.Conf 26 promo image

Get hands-on with Splunk

Join us September 14–17 in Denver, CO for an immersive learning and networking event.

Register for .conf26

features

Govern the cost of agentic AI

Explore the documentation
Evaluate Evaluate

Evaluate every interaction in production

Run always-on evaluations without the cost and latency of larger models. Luna-2’s fine-tuned SLMs deliver millisecond-level verdicts at pennies per million tokens, making high-scale production evaluation practical.

Catch Catch

Catch risky agent actions before they execute

Guardrail agentic workflows before mistakes become actions. Luna-2 evaluates tool selection and agent flows in real time, catching risky behavior before tools execute rather than only detecting problems afterward.

visibility-into-it-and-industrial-data visibility-into-it-and-industrial-data

Run more checks without slowing agents down

Evaluate multiple dimensions of agent behavior at once without adding seconds of delay. Luna-2 can run 10–20 checks simultaneously in under 200 milliseconds on L4 GPUs, enabling real-time guardrails at scale.

traffic traffic

Cut the cost of production evaluations

Get production-grade evaluation without paying for a frontier LLM judge on every interaction. Luna-2 delivers evaluations at $0.12 per million tokens while maintaining high accuracy and low latency.

priced-and-packaged-for-small-it-environments priced-and-packaged-for-small-it-environments

Know whether agents actually complete the job

Measure more than the final response. Luna-2 evaluates tool errors, tool selection quality, action advancement, and action completion so you can see whether agents successfully move toward and accomplish user goals.

optimize-incident-response optimize-incident-response

Scale the metrics that matter to your application

Evaluate the behaviors specific to your AI application. Luna-2 can power custom LLM metrics for production use cases, while lightweight adapters allow one base model to scale across hundreds of metrics with minimal infrastructure overhead.

 

 

Resources
Explore more from Splunk

Luna Evaluation Models FAQs

Luna is a family of purpose-built small language models for evaluation. They grade and guard agent interactions quickly and cheaply, which is what makes it affordable to watch 100% of traffic instead of a sample.

A frontier judge is expensive and slow, so teams sample 1 to 5% of traffic and can only review actions after they run. Luna is built for one job, judging agent interactions, so it is cheap enough to grade everything and fast enough to guard in real time.

Luna runs evaluations and guardrails at roughly 98% lower cost than an LLM-as-judge approach, which changes the economics from rationing to running on every interaction.

Yes. Luna returns a verdict as a single token, so scores are deterministic: run the same input twice and you get the same answer, which is what makes a model trustworthy as a judge.

Yes. With Luna Studio you correct a handful of verdicts and it learns your standard, reaching around 95% accuracy on your own tasks, with no labeling pipeline or model-training expertise.

Related products

Splunk Cloud Platform

Unify data, context, and action across every domain.

Learn more

Splunk Enterprise Security

Unified threat detection, investigation, and response for the agentic SOC.

Learn more

Splunk IT Service Intelligence

Predict and prevent IT issues with AI-driven service monitoring.

Learn more
Get started with Luna

Evaluate and guardrail every interaction, affordably.

Request a demo
Explore free trials