Pratik Bhavsar's Blog Posts
Pratik Bhavsar is an AI Engineer at Splunk focusing on agent evaluation, reliability, and observability. He has spent a decade building across the full AI stack, from training transformers, building semantic search, and shipping ML systems to designing agentic architectures.
He joined Cisco through the acquisition of Galileo, where he led open-source evaluations and developer relations. He built the Agent Leaderboard, an open benchmark measuring AI agent performance on real-world tasks, and the Hallucination Index, a systematic study of factual reliability across foundation models.
He is the author of five technical books covering eval engineering, agentic systems, RAG, multi-agent architectures, and LLM-as-a-Judge methodologies. Prior to Galileo, Pratik was a founding NLP Scientist at Enterpret and Senior Data Scientist at Morningstar. Pratik holds an M.Tech from IIT Bombay and loves to share his thoughts on Substack.

LLM-as-Judge vs. Human Evaluation: When to Use Each (And Why Elite Teams Use Both)

LLM Judges vs. SLM Judges: When To Use Which

SLM-as-Judge: How to Build and Deploy an SLM Judge

Token Meter: A Live Cost Meter for Your Coding Agents

How To Reduce Hallucinations in RAG Applications

How Evals Become Guardrails

Character Error Rate (CER): Meaning, Formula, and How to Use It in 2026

How to Build an Agent Evaluation Framework for Production AI

Evaluating AI Agents on Tool Calling and Planning

4 Key RAG Metrics to Improve Retrieval and Generation

How To Choose a Vector Database Architecture

Evaluating Agents for Manipulation, Deception and Adversarial Risk

Types of Multi-Agent System Failures: 7 Common Failures

Multi-Agent Coordination: 10 Strategies to Prevent System Failures

Context Engineering for AI Agents in Production: Architecture, Failure Modes, and Metrics

Scaling Laws of AI Tokenomics

How To Continuously Improve Your LangGraph Multi-Agent System: A Tutorial
