Evaluation
LLM as a Judge: Score AI Agent Outputs with Claude (2026)
A minimal, framework-free LLM-as-a-judge harness in Python on Claude: rubric design, pointwise and pairwise scoring, and fixes for position, verbosity, and self-preference bias.
7 min read7