Tool-calling accuracy, JSON-schema adherence, latency and cost across 6 models — 20 tool-calling cases (hallucination traps, parallel calls, no-tool distractors) and 12 strict-JSON extraction cases. Generated 2026-09-28. Method & harness on GitHub.
Exact tool + argument match, over-calling penalized. 20 cases.
Strict = raw parseable JSON as instructed (no fences). Content = schema-valid after stripping fences. The gap is formatting discipline, not capability. 12 cases.