
Agent Evaluation for Tool Use: Why pass@1 Lies, and How to Measure Reliability
Agent evaluation for tool use: why pass@1 hides unreliability, the pass^k metric, BFCL and tau-bench, state-based vs LLM-judge scoring, and building a harness.
Measuring what models actually do. Benchmark contamination and saturation, LLM-as-a-judge bias, pass@1 versus reliability under repetition, and the gap between a leaderboard number and behaviour you can depend on.

Agent evaluation for tool use: why pass@1 hides unreliability, the pass^k metric, BFCL and tau-bench, state-based vs LLM-judge scoring, and building a harness.

Effective context length is far shorter than the advertised window. What RULER and NoLiMa reveal about 1M-token models, why context rots, and when RAG still wins.

LLM evaluation is breaking down: benchmark saturation, contamination, and biased LLM-as-a-judge setups make leaderboard numbers misleading. Here is what to measure instead.