Pratik Bhavsar's Blog Posts

Pratik Bhavsar

Pratik Bhavsar is an AI Engineer at Splunk focusing on agent evaluation, reliability, and observability. He has spent a decade building across the full AI stack, from training transformers, building semantic search, and shipping ML systems to designing agentic architectures.

He joined Cisco through the acquisition of Galileo, where he led open-source evaluations and developer relations. He built the Agent Leaderboard, an open benchmark measuring AI agent performance on real-world tasks, and the Hallucination Index, a systematic study of factual reliability across foundation models.

He is the author of five technical books covering eval engineering, agentic systems, RAG, multi-agent architectures, and LLM-as-a-Judge methodologies. Prior to Galileo, Pratik was a founding NLP Scientist at Enterpret and Senior Data Scientist at Morningstar. Pratik holds an M.Tech from IIT Bombay and loves to share his thoughts on Substack.

LLM-as-Judge vs. Human Evaluation: When to Use Each (And Why Elite Teams Use Both)
Learn
7 Minute Read

LLM-as-Judge vs. Human Evaluation: When to Use Each (And Why Elite Teams Use Both)

Understand the tradeoffs between LLMs and humans for generative AI evaluation
LLM Judges vs. SLM Judges: When To Use Which
Learn
8 Minutes Read

LLM Judges vs. SLM Judges: When To Use Which

LLM judge vs SLM judge: an SLM costs 10-30x less and runs in 15-150ms, making 100% coverage affordable. See where each wins and when to make the switch.
SLM-as-Judge: How to Build and Deploy an SLM Judge
Learn
7 Minute Read

SLM-as-Judge: How to Build and Deploy an SLM Judge

How to build an SLM judge: See how to size the model, fine-tune with LoRA, validate it, and serve it in production.
Token Meter: A Live Cost Meter for Your Coding Agents
Artificial Intelligence
8 Minute Read

Token Meter: A Live Cost Meter for Your Coding Agents

See what your coding agents cost before the bill arrives. Token Meter tracks Claude Code, Codex, and Cursor usage live — free, open source, and 100% local.
How To Reduce Hallucinations in RAG Applications
Artificial Intelligence
15 Minute Read

How To Reduce Hallucinations in RAG Applications

Learn to diagnose the six primary hallucination patterns and implement a robust pipeline of retrieval, reranking, and guardrail enforcement to ensure reliability.
How Evals Become Guardrails
Learn
7 Minute Read

How Evals Become Guardrails

Learn the eval-to-guardrail lifecycle to transform offline evaluation criteria into runtime policies that intercept and block agent failures in production.
Character Error Rate (CER): Meaning, Formula, and How to Use It in 2026
Learn
5 Minute Read

Character Error Rate (CER): Meaning, Formula, and How to Use It in 2026

Character error rate (CER) is a an AI accuracy metric. Learn how to ntegrate CER into a quality stack for transcription, OCR, and extraction workflows.
How to Build an Agent Evaluation Framework for Production AI
Artificial Intelligence
8 minute read

How to Build an Agent Evaluation Framework for Production AI

Learn how to evaluate agents for production AI with this framework approach.
Evaluating AI Agents on Tool Calling and Planning
Artificial Intelligence
7 MINUTE READ

Evaluating AI Agents on Tool Calling and Planning

Evaluate AI agent performance on tool calling and planning by prioritizing consistency metrics, multi-turn state handling, and domain-specific assertions over public benchmark scores
4 Key RAG Metrics to Improve Retrieval and Generation
Artificial Intelligence
10 Minute Read

4 Key RAG Metrics to Improve Retrieval and Generation

Use these four key metrics—Context Relevance, Chunk Relevance, Context Adherence, and Completeness—to diagnose whether your RAG system is failing at retrieval or generation.
How To Choose a Vector Database Architecture
Artificial Intelligence
9 minute read

How To Choose a Vector Database Architecture

Choose the right vector database for your RAG systems. Compare database index structures, quantization methods, and deployment options to select the most suitable architecture.
Evaluating Agents for Manipulation, Deception and Adversarial Risk
Artificial Intelligence
10 Minute Read

Evaluating Agents for Manipulation, Deception and Adversarial Risk

How to evaluate AI agents for manipulation, deception and adversarial risk — where persuasion evals went, why prompt injection became the dominant threat, and what to measure.
Types of Multi-Agent System Failures: 7 Common Failures
Artificial Intelligence
7 Minute Read

Types of Multi-Agent System Failures: 7 Common Failures

Multi-agent systems face exponential coordination costs. This guide outlines common production failure modes and the architectural principles required to build reliable, scalable agent swarms.
Multi-Agent Coordination: 10 Strategies to Prevent System Failures
Artificial Intelligence
7 Minute Read

Multi-Agent Coordination: 10 Strategies to Prevent System Failures

Prevent failures in multi-agent AI systems with strategies like deterministic task allocation, hierarchical goal decomposition, and real-time observability.
Context Engineering for AI Agents in Production: Architecture, Failure Modes, and Metrics
Artificial Intelligence
15 Minute Read

Context Engineering for AI Agents in Production: Architecture, Failure Modes, and Metrics

Diagnose and resolve agent context failures by identifying common patterns, applying architectural strategies, and implementing a robust production observability loop.
Scaling Laws of AI Tokenomics
Artificial Intelligence
11 Minute Read

Scaling Laws of AI Tokenomics

Tokenomics isn’t about reducing tokens – it’s about understanding the marginal return of inference.
How To Continuously Improve Your LangGraph Multi-Agent System: A Tutorial
Artificial Intelligence
10 Minute Read

How To Continuously Improve Your LangGraph Multi-Agent System: A Tutorial

Improve the performance and reliability of your LangGraph multi-agent with observability to trace agent decisions, isolate failure patterns, and optimize workflows.
How to Reduce Agent Cost by Model Routing
Artificial Intelligence
8 Minute Read

How to Reduce Agent Cost by Model Routing

Learn how model routing helps AI agents reduce token costs by using the right model for each task while maintaining performance.