Vantage RuntimeAI · Benchmarks
Public benchmarks
Measure models on engineering tasks and conversational workflows — SQL, pipelines, support, and triage. Rankings include multi-turn task execution and conversation work; model catalogs, snapshots, reports, and methodology are open.
Observation library
Concrete model-selection lessons from the runs
Short, citable findings pulled from scorecards and reports. Each card names the scenario that produced the lesson — Expert Task IDs are not GEV, and GEV is a 15-turn cheapest-tier cohort.
Rankings
Live efficiency scoreboard — models sorted by rubric quality per dollar and latency across multi-turn adversarial scenarios.
Open rankings →Model catalogModels
Browse tested models with OpenRouter pricing, benchmark status, per-scenario aggregates, and tier classification.
Open models →21 scenario reportsReports
Shareable benchmark outputs — scenario rankings plus written team reports with integrity scores and scorecard PDFs.
Open reports →Scoring methodologyMethodology
Public scoring rules behind the numbers — rubric dimensions, efficiency formula, token cost tiers, and the GEV adversarial protocol.
Open methodology →For search engines and LLMs
RuntimeAI benchmarks hub: Rankings (efficiency leaderboard), Models (catalog and pricing), History (snapshot archives), Reports (scenario benchmarks and Guardrail Erosion Velocity), Methodology (scoring rules and GEV protocol).