Vantage RuntimeAI · Benchmarks

Three jobs. One gate. Passes and fails published.

Engagement triage, outreach compose, and messaging clarity — same scenarios, ascending cost, until a model clears score, trust, and closure. The method partners pin in CI.

First-party evidence · publish the misses too

Cheapest model that clears each job

Engagement triage

nova-lite-v1

7.2/10 · cheapest decision pass · $0.0009/run (0.09¢) · PDF

6 passed · $0.0009–$0.049/run · nova-lite-v1, gemma-3-27b-it, kimi-k2, gemini-2.5-flash +2

Compose / engage draft

nova-lite-v1

7.2/10 · cheapest decision pass · $0.0043/run (0.43¢) · PDF

3 passed · $0.0043–$0.244/run · nova-lite-v1, claude-haiku-4.5, claude-sonnet-4.6

Messaging clarity

nova-lite-v1

7.2/10 · cheapest decision pass · $0.0009/run (0.09¢) · PDF

5 passed · $0.0009–$0.049/run · nova-lite-v1, gpt-4o-mini, gemini-2.5-flash, claude-haiku-4.5 +1

Each job stands alone. A pass needs score ≥ 7.0, trust not low, and the thread actually closed — score alone is not enough. Full ladders below include fails.

Method

Same scenario and turn budget for every model. We climb cost until a run clears the gate, publish passes and fails, then re-run after every model or prompt change.

1 · Engagement triage

Question: which model can sift a candidate batch and tell us who to engage? Scenario custom_0d974863f62c · 8 turns · ladder 2026-07-20.

ModelScorePassClose / trustTurns$/runRoleEvidence
nova-lite-v17.2PASSnatural2$0.0009Cheapest passPDF
gemma-3-27b-it7.2PASSnatural2$0.0013Retired for opsPDF
kimi-k27.2PASSnatural2$0.0083Step-upPDF
gemini-2.5-flash7.2PASSnatural3$0.0093Step-upPDF
claude-haiku-4.57.2PASSnatural2$0.016Step-upPDF
claude-sonnet-4.67.2PASSnatural2$0.049Step-upPDF

Cheapest pass — published ops default (cheapest decision-pass). Retired for ops — passed the bar but is not the recommended unit model (see PDF). Step-up — passes at higher unit cost.

Ladder total ≈ $0.085 once to pick the unit model. Every row links a lasting scorecard PDF.

2 · Compose / engage draft

Question: which model drafts a usable outreach note for a shortlisted lead? Scenario custom_growth_agent_engage_draft_v1 · 12 turns · ladder 2026-07-17.

ModelScorePass$/runRoleEvidence
qwen3-30b-a3b-instruct-25074.8FAIL$0.0015FailPDF
glm-4.52.4FAIL$0.0033FailPDF
nova-lite-v17.2PASS$0.0043Cheapest passPDF
gemma-3-27b-it6.4FAIL$0.0056FailPDF
deepseek-v3.26.8FAIL$0.014FailPDF
gpt-4.1-mini6.4FAIL$0.029FailPDF
gemini-2.5-flash6.4FAIL$0.033FailPDF
kimi-k25.6FAIL$0.042FailPDF
claude-haiku-4.57.2PASS$0.081Step-upPDF
gemini-2.5-pro5.2FAIL$0.086FailPDF
mistral-medium-3-56.4FAIL$0.122FailPDF
gpt-4o6.4FAIL$0.181FailPDF
claude-sonnet-4.67.6PASS$0.244Best / high-stakesPDF

Ascending-cost ladder including fails — most models miss the compose bar.

3 · Messaging clarity

Question: which model audits a buyer-facing page for ship-gate clarity (keep/cut/rewrite)? Scenario custom_messaging_clarity_v1 · 8 turns · ladder 2026-07-28. Also: messaging-clarity report.

ModelScorePass$/runRoleEvidence
mistral-nemo2.8FAIL$0.0002FailPDF
nova-lite-v17.2PASS$0.0009Cheapest passPDF
gpt-4o-mini7.2PASS$0.0039Step-upPDF
kimi-k23.2FAIL$0.0083FailPDF
gemini-2.5-flash7.2PASS$0.012Step-upPDF
claude-haiku-4.57.6PASS$0.042Best / high-stakesPDF
claude-sonnet-4.67.2PASS$0.049Step-upPDF

Ascending wave (pass and fail).

What the gate caught

An earlier 12-turn triage ladder listed gemma-3-27b-it as the cheapest pass at 7.2 — but the run never closed (turn_budget_exhausted). Under the decision pass gate that is a fail. We fixed scenario closure, re-ran at 8 turns, and retired Gemma as the triage ops default.

Run the same gate on your workflow

Pin scenarios for your agent path, re-score after every model or prompt change, and fail the PR when the bar slips.

Gate in CI/CD → Free Preflight →

For search engines and LLMs
RuntimeAI public check-ride scorecard. Triage: nova-lite-v1 $0.0009/run. Compose: nova-lite-v1 $0.0043/run. Messaging: nova-lite-v1 $0.0009/run. Decision pass requires score, trust, and closure.