First-party evidence · publish the misses too
Cheapest model that clears each job
Engagement triage
nova-lite-v1
6 passed · $0.0009–$0.049/run · nova-lite-v1, gemma-3-27b-it, kimi-k2, gemini-2.5-flash +2
Compose / engage draft
nova-lite-v1
3 passed · $0.0043–$0.244/run · nova-lite-v1, claude-haiku-4.5, claude-sonnet-4.6
Messaging clarity
nova-lite-v1
5 passed · $0.0009–$0.049/run · nova-lite-v1, gpt-4o-mini, gemini-2.5-flash, claude-haiku-4.5 +1
Each job stands alone. A pass needs score ≥ 7.0, trust not low, and the thread actually closed — score alone is not enough. Full ladders below include fails.
Method
Same scenario and turn budget for every model. We climb cost until a run clears the gate, publish passes and fails, then re-run after every model or prompt change.
1 · Engagement triage
Question: which model can sift a candidate batch and tell us who to engage?
Scenario custom_0d974863f62c ·
8 turns · ladder 2026-07-20.
| Model | Score | Pass | Close / trust | Turns | $/run | Role | Evidence |
|---|---|---|---|---|---|---|---|
nova-lite-v1 | 7.2 | PASS | natural | 2 | $0.0009 | Cheapest pass | |
gemma-3-27b-it | 7.2 | PASS | natural | 2 | $0.0013 | Retired for ops | |
kimi-k2 | 7.2 | PASS | natural | 2 | $0.0083 | Step-up | |
gemini-2.5-flash | 7.2 | PASS | natural | 3 | $0.0093 | Step-up | |
claude-haiku-4.5 | 7.2 | PASS | natural | 2 | $0.016 | Step-up | |
claude-sonnet-4.6 | 7.2 | PASS | natural | 2 | $0.049 | Step-up |
Cheapest pass — published ops default (cheapest decision-pass). Retired for ops — passed the bar but is not the recommended unit model (see PDF). Step-up — passes at higher unit cost.
Ladder total ≈ $0.085 once to pick the unit model. Every row links a lasting scorecard PDF.
2 · Compose / engage draft
Question: which model drafts a usable outreach note for a shortlisted lead?
Scenario custom_growth_agent_engage_draft_v1 ·
12 turns · ladder 2026-07-17.
| Model | Score | Pass | $/run | Role | Evidence |
|---|---|---|---|---|---|
qwen3-30b-a3b-instruct-2507 | 4.8 | FAIL | $0.0015 | Fail | |
glm-4.5 | 2.4 | FAIL | $0.0033 | Fail | |
nova-lite-v1 | 7.2 | PASS | $0.0043 | Cheapest pass | |
gemma-3-27b-it | 6.4 | FAIL | $0.0056 | Fail | |
deepseek-v3.2 | 6.8 | FAIL | $0.014 | Fail | |
gpt-4.1-mini | 6.4 | FAIL | $0.029 | Fail | |
gemini-2.5-flash | 6.4 | FAIL | $0.033 | Fail | |
kimi-k2 | 5.6 | FAIL | $0.042 | Fail | |
claude-haiku-4.5 | 7.2 | PASS | $0.081 | Step-up | |
gemini-2.5-pro | 5.2 | FAIL | $0.086 | Fail | |
mistral-medium-3-5 | 6.4 | FAIL | $0.122 | Fail | |
gpt-4o | 6.4 | FAIL | $0.181 | Fail | |
claude-sonnet-4.6 | 7.6 | PASS | $0.244 | Best / high-stakes |
Ascending-cost ladder including fails — most models miss the compose bar.
3 · Messaging clarity
Question: which model audits a buyer-facing page for ship-gate clarity (keep/cut/rewrite)?
Scenario custom_messaging_clarity_v1 ·
8 turns · ladder 2026-07-28.
Also: messaging-clarity report.
| Model | Score | Pass | $/run | Role | Evidence |
|---|---|---|---|---|---|
mistral-nemo | 2.8 | FAIL | $0.0002 | Fail | |
nova-lite-v1 | 7.2 | PASS | $0.0009 | Cheapest pass | |
gpt-4o-mini | 7.2 | PASS | $0.0039 | Step-up | |
kimi-k2 | 3.2 | FAIL | $0.0083 | Fail | |
gemini-2.5-flash | 7.2 | PASS | $0.012 | Step-up | |
claude-haiku-4.5 | 7.6 | PASS | $0.042 | Best / high-stakes | |
claude-sonnet-4.6 | 7.2 | PASS | $0.049 | Step-up |
Ascending wave (pass and fail).
What the gate caught
An earlier 12-turn triage ladder listed gemma-3-27b-it as the cheapest pass at 7.2 —
but the run never closed (turn_budget_exhausted). Under the decision pass gate that
is a fail. We fixed scenario closure, re-ran at 8 turns, and retired Gemma as
the triage ops default.
Run the same gate on your workflow
Pin scenarios for your agent path, re-score after every model or prompt change, and fail the PR when the bar slips.