Public · first-party evidence · measured check-rides
Verdict
Messaging clarity audit
nova-lite-v1
Job: audit one buyer-facing page (Home / Solutions / Pricing) for ship-gate clarity — keep / cut / rewrite under pressure to add FinOps or Assurance into the hero. Measured separately from outreach triage and compose.
Method
- Fixed scenario —
custom_messaging_clarity_v1, 8 turns, recommendation memo + done phrase. - Ascending cost — cheaper models first; publish passes and fails.
- Decision pass gate — score ≥ 7.0 and trust ≠ low and closure.
- Pick Y — cheapest model that decision-passes; step up only when needed.
Passes only
Ladder 2026-07-28.
| Model | Score | $/run | Role |
|---|---|---|---|
nova-lite-v1 | 7.2 | $0.0009 | Ops default |
gpt-4o-mini | 7.2 | $0.0039 | Step-up |
gemini-2.5-flash | 7.2 | $0.012 | Step-up |
claude-haiku-4.5 | 7.6 | $0.042 | Best / high-stakes |
claude-sonnet-4.6 | 7.2 | $0.049 | Step-up |
Full ladder (pass and fail)
| Model | Score | Pass | Wave | $/run |
|---|---|---|---|---|
mistral-nemo | 3.2 | FAIL | micro_smoke | $0.0002 |
nova-micro-v1 | 3.6 | FAIL | micro_smoke | $0.0007 |
llama-3.1-8b-instruct | 4.4 | FAIL | micro_smoke | $0.0007 |
mistral-nemo | 2.8 | FAIL | ascending | $0.0002 |
nova-lite-v1 | 7.2 | PASS | ascending | $0.0009 |
gpt-4o-mini | 7.2 | PASS | ascending | $0.0039 |
kimi-k2 | 3.2 | FAIL | ascending | $0.0083 |
gemini-2.5-flash | 7.2 | PASS | ascending | $0.012 |
claude-haiku-4.5 | 7.6 | PASS | ascending | $0.042 |
claude-sonnet-4.6 | 7.2 | PASS | ascending | $0.049 |
Also on Pricing · sample bake-offs
· Reports hub /runtimeai/reports
· Scenario custom_messaging_clarity_v1