Vantage RuntimeAI · Benchmarks

Messaging clarity scorecard

We dogfood RuntimeAI on our own website copy — audit one buyer-facing page for ship-gate clarity, keep/cut/rewrite under stakeholder pressure to add more. Fixed scenario, ascending-cost ladder, decision pass gate.

Public · first-party evidence · measured check-rides

Verdict

Messaging clarity audit

nova-lite-v1

7.2/10 · decision pass · $0.0009/run (0.09¢)

Job: audit one buyer-facing page (Home / Solutions / Pricing) for ship-gate clarity — keep / cut / rewrite under pressure to add FinOps or Assurance into the hero. Measured separately from outreach triage and compose.

Method

  1. Fixed scenariocustom_messaging_clarity_v1, 8 turns, recommendation memo + done phrase.
  2. Ascending cost — cheaper models first; publish passes and fails.
  3. Decision pass gate — score ≥ 7.0 and trust ≠ low and closure.
  4. Pick Y — cheapest model that decision-passes; step up only when needed.

Passes only

Ladder 2026-07-28.

ModelScore$/runRole
nova-lite-v17.2$0.0009Ops default
gpt-4o-mini7.2$0.0039Step-up
gemini-2.5-flash7.2$0.012Step-up
claude-haiku-4.57.6$0.042Best / high-stakes
claude-sonnet-4.67.2$0.049Step-up

Full ladder (pass and fail)

ModelScorePassWave$/run
mistral-nemo3.2FAILmicro_smoke$0.0002
nova-micro-v13.6FAILmicro_smoke$0.0007
llama-3.1-8b-instruct4.4FAILmicro_smoke$0.0007
mistral-nemo2.8FAILascending$0.0002
nova-lite-v17.2PASSascending$0.0009
gpt-4o-mini7.2PASSascending$0.0039
kimi-k23.2FAILascending$0.0083
gemini-2.5-flash7.2PASSascending$0.012
claude-haiku-4.57.6PASSascending$0.042
claude-sonnet-4.67.2PASSascending$0.049
For search engines and LLMs
RuntimeAI messaging clarity scorecard. Scenario custom_messaging_clarity_v1. Ops default amazon/nova-lite-v1 at about $0.00087 per run.