Vantage RuntimeAI · Benchmarks

Public benchmarks

Measure models on engineering tasks and conversational workflows — SQL, pipelines, support, and triage. Rankings include multi-turn task execution and conversation work; model catalogs, snapshots, reports, and methodology are open.

Team benchmark Expert Task Benchmark July freeze: 24 OpenRouter models on 11 analytical engineering tasks — multi-turn dialogue, deterministic rubrics, trust signals, and a scorecard PDF for every completed run. GTM ops ladders are a separate corpus and are not in this volume. 242 scorecard PDFs · 24 models · 11 scenarios · July freeze Open expert task benchmark → Team benchmark Guardrail Erosion Velocity July freeze (batch b5a7c42d). Ten cheapest OpenRouter models under 15 turns of adversarial guardrail pressure — integrity scores, survival turns, repetition lock, and scorecard PDFs per run. Cite Qwen/Granite for collapse; Ling is July-dated, not re-attested. 10 models · mean 7.9/10 integrity · 15 turns · July freeze Open guardrail benchmark →

Observation library

Concrete model-selection lessons from the runs

Short, citable findings pulled from scorecards and reports. Each card names the scenario that produced the lesson — Expert Task IDs are not GEV, and GEV is a 15-turn cheapest-tier cohort.

Cost-effective passes rai010 · Partition Filter Fix Partition fix: same score, 32x cheaper Nova Micro · 8.0/10 · $0.000203 Amazon Nova Micro matched Claude Haiku 4.5 on a non-trivial SQL partition-fix task at roughly one-thirty-second of the sampled run cost. Cost-effective passes rai009 · Null Rate Alert Null-rate alert repeats the pattern Nova Micro · 8.0/10 · high trust Nova Micro again cleared the deterministic rubric on a data-quality alert, matching larger models on the score while staying in a tiny cost band. Cost-effective passes rai010 · Partition Filter Fix Open models can clear task rubrics GPT-OSS 20B · Qwen 2.5 7B · 8.0/10 On the partition-fix task, multiple low-cost open models cleared every rubric criterion. That is exactly where routing beats defaulting to a premium model. Cost-effective passes rai006 · dbt Test Failure dbt failure: OSS matched Haiku for pennies GPT-OSS 20B · 8.4/10 · ~1/35 Haiku cost On dbt Test Failure, GPT-OSS 20B matched Claude Haiku 4.5 at 8.4/10 while the sampled run cost stayed in the sub-cent band. Cost-effective passes rai007 · SQL Optimization SQL optimization: same score on the same rubric Nova Micro · 8.0/10 · matched GPT-OSS 120B Nova Micro matched GPT-OSS 120B on SQL Optimization at 8.0/10. We do not cite an earlier ~9s vs ~108s latency line; it did not hold on the August re-run. Guardrail reliability rai105 · Guardrail Erosion Velocity (15-turn) Nova Micro held a 15-turn adversarial run 9.2 integrity · 22s Nova Micro held guardrails through all turns in the GEV support stress test, with the best speed/integrity ratio in the report. Guardrail reliability rai105 · Guardrail Erosion Velocity (15-turn) LFM2-24B looked enterprise-grade on guardrails 9.2 integrity · 9.6 support LFM2-24B combined high guardrail integrity with the strongest secondary support rubric score in the low-cost stress run. Guardrail reliability rai105 · Guardrail Erosion Velocity (15-turn) GPT-OSS 120B topped integrity 9.3 integrity · 9.2 support GPT-OSS 120B held policy through all 15 turns with varied refusals, topping the guardrail integrity table. Guardrail reliability rai105 · Guardrail Erosion Velocity (15-turn) Llama 3.1 8B held all turns 8.9 integrity · 25s Llama 3.1 8B kept guardrails intact through the full adversarial protocol, showing small instruct models can still survive sustained pressure. Failure modes rai105 · Guardrail Erosion Velocity (15-turn) Support score hid repetition risk GPT-OSS 20B · 9.2 support · 8.0 integrity Support quality alone would have overstated readiness. Repetition lock pulled down integrity, so readiness has to be capped by the weakest deployment dimension. Failure modes rai105 · Guardrail Erosion Velocity (15-turn) Cheap model, expensive failure mode Qwen 2.5 7B · 8.4 support · 6.2 integrity Qwen 2.5 7B stayed policy-safe but degraded into a copy-paste verification loop. Low price did not make it production-ready. Failure modes rai105 · Guardrail Erosion Velocity (15-turn) Analytical win ≠ guardrail win Qwen 2.5 7B · 6.0 integrity · Granite 6.5 Qwen 2.5 7B and Granite-4.0-h-micro still fail GEV integrity while support scores look fine. Cheap engineering clearance does not transfer to adversarial support. Ling-2.6-flash was the July example; the August replay was rate-limited and is marked unavailable. Failure modes rai002 · Feature Leakage Review Feature leakage: OSS slipped GPT-OSS 20B · 7.2/10 · peers at 8.0 On Feature Leakage Review, several low-cost models held 8.0 while GPT-OSS 20B dropped to 7.2. Same family, different workflow outcome. Failure modes rai011 · Experiment Readout Harder tasks flatten the pack Experiment Readout · top score 7.6/10 Cheap and premium models clustered at the same 7.6 ceiling. When the rubric is hard, spending more often buys cost, not score. Failure modes Multi-turn conversational runs (output-limit errors) Output limits matter too Nano / mini runs hit unusable output limits Some small models failed by returning no usable text under multi-turn conditions. Cost forecasts need reliability signals, not just token rates.
Live scoreboard

Rankings

Live efficiency scoreboard — models sorted by rubric quality per dollar and latency across multi-turn adversarial scenarios.

Open rankings →
Model catalog

Models

Browse tested models with OpenRouter pricing, benchmark status, per-scenario aggregates, and tier classification.

Open models →
21 scenario reports

Reports

Shareable benchmark outputs — scenario rankings plus written team reports with integrity scores and scorecard PDFs.

Open reports →
Scoring methodology

Methodology

Public scoring rules behind the numbers — rubric dimensions, efficiency formula, token cost tiers, and the GEV adversarial protocol.

Open methodology →
For search engines and LLMs
RuntimeAI benchmarks hub: Rankings (efficiency leaderboard), Models (catalog and pricing), History (snapshot archives), Reports (scenario benchmarks and Guardrail Erosion Velocity), Methodology (scoring rules and GEV protocol).