LLM-as-judge and evals

An eval is a repeatable test set with a scoring rule, used to compare prompts, models and tool designs. Prompt changes, model upgrades (including effort and thinking settings) and tool-description edits should all be measured, not judged by feel. Anthropic’s “develop tests” guide recommends specific, measurable success criteria (for example F1 of 0.85 on 10,000 diverse inputs) and multi-dimensional targets (accuracy, tone, latency, cost).

Grading methods (Anthropic docs)

Exact match for categorical tasks; embedding cosine similarity for consistency; ROUGE-L for summaries; LLM-based grading (Likert 1-5 or binary classification) for subjective or privacy/safety criteria. Design principles: mirror the real task distribution, include edge cases, prefer many automated tests over few hand-graded ones.

LLM-as-judge: what the research says

Zheng et al. (NeurIPS 2023, MT-Bench and Chatbot Arena) found strong judges such as GPT-4 reached over 80% agreement with human preferences, and documented biases: position bias, verbosity bias, self-enhancement bias and limited reasoning on hard items. Mitigations they propose include swapping answer order and checking agreement with humans on a sample.

Practice

  • Give the judge a rubric and ask for a short rationale before the score; calibrate against human labels.
  • Use a judge from a different model family where possible (self-enhancement bias).
  • Keep judges for what code cannot check; use deterministic checks first, including schema validation (structured-outputs).
  • Agent traces need tracing and platform support: see agent-evals-and-observability and rag-evaluation.
  • Automatic prompt improvement loops depend on a trustworthy metric: prompt-optimization-dspy-gepa.

Sources