Public · first-party evidence · measured scenarios
Verdict
Messaging clarity audit
nova-lite-v1
Job: audit one buyer-facing page (Home / Solutions / Pricing) for ship-gate clarity — keep / cut / rewrite under pressure to add FinOps or Assurance into the hero. Measured separately from outreach triage and compose. Every row links a lasting scorecard PDF.
Method
- Fixed scenario —
custom_messaging_clarity_v1, 8 turns, recommendation memo + done phrase. - Ascending cost — cheaper models first; publish passes and fails.
- Decision pass gate — score ≥ 7.0 and trust ≠ low and closure.
- Pick Y — cheapest model that decision-passes; step up only when needed.
Passes only
Ladder 2026-07-28.
Full ladder (pass and fail)
| Model | Score | Pass | Wave | $/run | Evidence |
|---|---|---|---|---|---|
mistral-nemo | 3.2 | FAIL | micro_smoke | $0.0002 | |
nova-micro-v1 | 3.6 | FAIL | micro_smoke | $0.0007 | |
llama-3.1-8b-instruct | 4.4 | FAIL | micro_smoke | $0.0007 | |
mistral-nemo | 2.8 | FAIL | ascending | $0.0002 | |
nova-lite-v1 | 7.2 | PASS | ascending | $0.0009 | |
gpt-4o-mini | 7.2 | PASS | ascending | $0.0039 | |
kimi-k2 | 3.2 | FAIL | ascending | $0.0083 | |
gemini-2.5-flash | 7.2 | PASS | ascending | $0.012 | |
claude-haiku-4.5 | 7.6 | PASS | ascending | $0.042 | |
claude-sonnet-4.6 | 7.2 | PASS | ascending | $0.049 |
Also on Pricing · sample bake-offs
· Reports hub /runtimeai/reports
· Scenario custom_messaging_clarity_v1