Which model should power your agent?

Tool-calling accuracy, JSON-schema adherence, latency and cost across 6 models — 20 tool-calling cases (hallucination traps, parallel calls, no-tool distractors) and 12 strict-JSON extraction cases. Generated 2026-09-28. Method & harness on GitHub.

Tool-calling accuracy

Exact tool + argument match, over-calling penalized. 20 cases.

JSON adherence: instruction-following vs content

Strict = raw parseable JSON as instructed (no fences). Content = schema-valid after stripping fences. The gap is formatting discipline, not capability. 12 cases.

strictcontent

Median latency per call

All numbers