LLM-as-judge and evals
An eval is a repeatable test set with a scoring rule, used to compare prompts, models and tool designs. Prompt changes, model upgrades (including effort and thinking settings) and tool-description edits should all be measured, not judged by feel. Anthropic’s “develop tests” guide recommends specific, measurable success criteria (for example F1 of 0.85 on 10,000 diverse inputs) and multi-dimensional targets (accuracy, tone, latency, cost).
Grading methods (Anthropic docs)
Exact match for categorical tasks; embedding cosine similarity for consistency; ROUGE-L for summaries; LLM-based grading (Likert 1-5 or binary classification) for subjective or privacy/safety criteria. Design principles: mirror the real task distribution, include edge cases, prefer many automated tests over few hand-graded ones.
LLM-as-judge: what the research says
Zheng et al. (NeurIPS 2023, MT-Bench and Chatbot Arena) found strong judges such as GPT-4 reached over 80% agreement with human preferences, and documented biases: position bias, verbosity bias, self-enhancement bias and limited reasoning on hard items. Mitigations they propose include swapping answer order and checking agreement with humans on a sample.
Practice
- Give the judge a rubric and ask for a short rationale before the score; calibrate against human labels.
- Use a judge from a different model family where possible (self-enhancement bias).
- Keep judges for what code cannot check; use deterministic checks first, including schema validation (structured-outputs).
- Agent traces need tracing and platform support: see agent-evals-and-observability and rag-evaluation.
- Automatic prompt improvement loops depend on a trustworthy metric: prompt-optimization-dspy-gepa.
Sources
- https://platform.claude.com/docs/en/test-and-evaluate/develop-tests (2026-09-30)
- https://arxiv.org/abs/2306.05685 (2026-09-30)