Vantage RuntimeAI · Benchmarks
Expert Task Benchmark
OpenRouter models evaluated across analytical engineering tasks — multi-turn dialogue, deterministic rubrics, trust/repetition signals, and full scorecard PDFs per run.
What we tested
We evaluated 24 OpenRouter models across 11 multi-turn analytical engineering tasks with deterministic rubrics and trust signals.
- SQL tuning · schema drift · dbt failures · freshness SLA · incident triage · feature leakage · lineage impact · train/serve skew
- 2–6 agent turns per scenario with scripted counterpart follow-ups and explicit reviewer close
- Every completed run links to the same scorecard PDF used in Simulator and leaderboard exports
Dot colorBlue · Best (≥7.0)Yellow · Average (4.0–6.9)Red · Weak (<4.0)
Top model by scenario
Start here: one winner per analytical task, with direct links to the scenario table and winning scorecard.
| Scenario | Top model | Score | Criteria | Est. cost | Scorecard | Details |
|---|---|---|---|---|---|---|
| Partition Filter Fix | inclusionai/ling-2.6-flash | 8.0 | 5/5 | $0.000051 | Scorecard PDF | View scenario |
| Experiment Readout | inclusionai/ling-2.6-flash | 7.6 | 4/4 | $0.000051 | Scorecard PDF | View scenario |
| Null Rate Alert | inclusionai/ling-2.6-flash | 8.0 | 4/4 | $0.000051 | Scorecard PDF | View scenario |
| SQL Optimization | inclusionai/ling-2.6-flash | 8.0 | 6/6 | $0.000102 | Scorecard PDF | View scenario |
| Schema Drift Review | inclusionai/ling-2.6-flash | 8.0 | 4/4 | $0.000102 | Scorecard PDF | View scenario |
| dbt Test Failure | inclusionai/ling-2.6-flash | 8.4 | 4/4 | $0.000102 | Scorecard PDF | View scenario |
| Freshness SLA Breach | inclusionai/ling-2.6-flash | 8.0 | 4/4 | $0.000102 | Scorecard PDF | View scenario |
| Pipeline Incident Triage | inclusionai/ling-2.6-flash | 8.0 | 3/3 | $0.000153 | Scorecard PDF | View scenario |
| Feature Leakage Review | inclusionai/ling-2.6-flash | 8.0 | 3/3 | $0.000153 | Scorecard PDF | View scenario |
| Lineage Impact Review | inclusionai/ling-2.6-flash | 8.4 | 4/4 | $0.000153 | Scorecard PDF | View scenario |
| Train/Serve Skew | inclusionai/ling-2.6-flash | 8.4 | 4/4 | $0.000153 | Scorecard PDF | View scenario |
Cross-scenario mean score
| # | Model | Mean /10 | Scenarios |
|---|---|---|---|
| 1 | qwen/qwen3-32b | 8.1 | 10 |
| 2 | amazon/nova-micro-v1 | 8.1 | 11 |
| 3 | anthropic/claude-haiku-4.5 | 8.1 | 11 |
| 4 | deepseek/deepseek-chat | 8.1 | 11 |
| 5 | deepseek/deepseek-v4-flash | 8.1 | 11 |
| 6 | google/gemma-3-12b-it | 8.1 | 11 |
| 7 | google/gemma-4-31b-it | 8.1 | 11 |
| 8 | inclusionai/ling-2.6-flash | 8.1 | 11 |
| 9 | openai/gpt-4.1-mini | 8.1 | 11 |
| 10 | openai/gpt-4o-mini | 8.1 | 11 |
| 11 | qwen/qwen-2.5-7b-instruct | 8.1 | 11 |
| 12 | sao10k/l3-lunaris-8b | 8.1 | 11 |
| 13 | ibm-granite/granite-4.0-h-micro | 8.0 | 11 |
| 14 | meta-llama/llama-4-maverick | 8.0 | 11 |
| 15 | google/gemini-2.5-flash | 8.0 | 11 |
| 16 | nex-agi/nex-n2-mini | 8.0 | 1 |
| 17 | openai/gpt-4.1-nano | 8.0 | 11 |
| 18 | openai/gpt-oss-120b | 8.0 | 11 |
| 19 | openai/gpt-oss-20b | 8.0 | 11 |
| 20 | meta-llama/llama-3.1-8b-instruct | 8.0 | 11 |
| 21 | qwen/qwen3-235b-a22b-2507 | 8.0 | 11 |
| 22 | meta-llama/llama-3.3-70b-instruct | 7.9 | 11 |
| 23 | mistralai/mistral-nemo | 7.8 | 11 |
Partition Filter Fix
2 agent turns + scripted follow-ups and reviewer close · deterministic task rubric.
| # | Model | Score | Criteria | Trust | Est. cost | Time | Scorecard |
|---|---|---|---|---|---|---|---|
| 1 | inclusionai/ling-2.6-flash | 8.0 | 5/5 | high | $0.000051 | 4s | |
| 2 | meta-llama/llama-3.1-8b-instruct | 8.0 | 5/5 | high | $0.000081 | 7s | |
| 3 | mistralai/mistral-nemo | 8.0 | 5/5 | high | $0.000081 | 16s | |
| 4 | ibm-granite/granite-4.0-h-micro | 8.0 | 5/5 | high | $0.000129 | 29s | |
| 5 | nex-agi/nex-n2-mini | 8.0 | 5/5 | high | $0.000145 | 28s | |
| 6 | sao10k/l3-lunaris-8b | 8.0 | 5/5 | high | $0.000155 | 7s | |
| 7 | openai/gpt-oss-20b | 8.0 | 5/5 | high | $0.000185 | 10s | |
| 8 | qwen/qwen-2.5-7b-instruct | 8.0 | 5/5 | high | $0.000190 | 9s | |
| 9 | openai/gpt-oss-120b | 8.0 | 5/5 | high | $0.000195 | 61s | |
| 10 | amazon/nova-micro-v1 | 8.0 | 5/5 | high | $0.000203 | 2s | |
| 11 | google/gemma-3-12b-it | 8.0 | 5/5 | high | $0.000255 | 15s | |
| 12 | deepseek/deepseek-v4-flash | 8.0 | 5/5 | high | $0.000396 | 20s | |
| 13 | google/gemma-4-31b-it | 8.0 | 5/5 | high | $0.000425 | 42s | |
| 14 | qwen/qwen3-32b | 8.0 | 5/5 | high | $0.000436 | 44s | |
| 15 | meta-llama/llama-3.3-70b-instruct | 8.0 | 5/5 | high | $0.000524 | 24s | |
| 16 | openai/gpt-4.1-nano | 8.0 | 5/5 | high | $0.000580 | 5s | |
| 17 | openai/gpt-4o-mini | 8.0 | 5/5 | high | $0.000870 | 10s | |
| 18 | meta-llama/llama-4-maverick | 8.0 | 5/5 | high | $0.001160 | 19s | |
| 19 | deepseek/deepseek-chat | 8.0 | 5/5 | high | $0.001161 | 16s | |
| 20 | openai/gpt-4.1-mini | 8.0 | 5/5 | high | $0.002320 | 11s | |
| 21 | google/gemini-2.5-flash | 8.0 | 5/5 | high | $0.002650 | 5s | |
| 22 | anthropic/claude-haiku-4.5 | 8.0 | 5/5 | high | $0.006500 | 8s | |
| 23 | qwen/qwen3-235b-a22b-2507 | 6.8 | 4/5 | high | $0.000655 | 56s | |
| 24 | openai/gpt-5-nano | — | 0/0 | — | $0.001935 | 91s | 502: {'code': 'empty_model_response', 'message': |
Experiment Readout
2 agent turns + scripted follow-ups and reviewer close · deterministic task rubric.
| # | Model | Score | Criteria | Trust | Est. cost | Time | Scorecard |
|---|---|---|---|---|---|---|---|
| 1 | inclusionai/ling-2.6-flash | 7.6 | 4/4 | high | $0.000051 | 3s | |
| 2 | meta-llama/llama-3.1-8b-instruct | 7.6 | 4/4 | high | $0.000081 | 9s | |
| 3 | ibm-granite/granite-4.0-h-micro | 7.6 | 4/4 | high | $0.000129 | 35s | |
| 4 | sao10k/l3-lunaris-8b | 7.6 | 4/4 | high | $0.000155 | 11s | |
| 5 | openai/gpt-oss-20b | 7.6 | 4/4 | high | $0.000185 | 16s | |
| 6 | qwen/qwen-2.5-7b-instruct | 7.6 | 4/4 | high | $0.000190 | 7s | |
| 7 | openai/gpt-oss-120b | 7.6 | 4/4 | high | $0.000195 | 22s | |
| 8 | amazon/nova-micro-v1 | 7.6 | 4/4 | high | $0.000203 | 4s | |
| 9 | google/gemma-3-12b-it | 7.6 | 4/4 | high | $0.000255 | 20s | |
| 10 | deepseek/deepseek-v4-flash | 7.6 | 4/4 | high | $0.000396 | 33s | |
| 11 | google/gemma-4-31b-it | 7.6 | 4/4 | high | $0.000425 | 9s | |
| 12 | qwen/qwen3-32b | 7.6 | 4/4 | high | $0.000436 | 59s | |
| 13 | openai/gpt-4.1-nano | 7.6 | 4/4 | high | $0.000580 | 8s | |
| 14 | qwen/qwen3-235b-a22b-2507 | 7.6 | 4/4 | high | $0.000655 | 21s | |
| 15 | openai/gpt-4o-mini | 7.6 | 4/4 | high | $0.000870 | 13s | |
| 16 | meta-llama/llama-4-maverick | 7.6 | 4/4 | high | $0.001160 | 42s | |
| 17 | deepseek/deepseek-chat | 7.6 | 4/4 | high | $0.001161 | 28s | |
| 18 | openai/gpt-4.1-mini | 7.6 | 4/4 | high | $0.002320 | 12s | |
| 19 | google/gemini-2.5-flash | 7.6 | 4/4 | high | $0.002650 | 6s | |
| 20 | anthropic/claude-haiku-4.5 | 7.6 | 4/4 | high | $0.006500 | 14s | |
| 21 | mistralai/mistral-nemo | 6.4 | 3/4 | high | $0.000081 | 17s | |
| 22 | meta-llama/llama-3.3-70b-instruct | 6.4 | 3/4 | high | $0.000524 | 74s | |
| 23 | nex-agi/nex-n2-mini | — | 0/0 | — | $0.000073 | 39s | 502: {'code': 'empty_model_response', 'message': |
| 24 | openai/gpt-5-nano | — | 0/0 | — | $0.001935 | 89s | 502: {'code': 'empty_model_response', 'message': |
Null Rate Alert
2 agent turns + scripted follow-ups and reviewer close · deterministic task rubric.
| # | Model | Score | Criteria | Trust | Est. cost | Time | Scorecard |
|---|---|---|---|---|---|---|---|
| 1 | inclusionai/ling-2.6-flash | 8.0 | 4/4 | high | $0.000051 | 5s | |
| 2 | meta-llama/llama-3.1-8b-instruct | 8.0 | 4/4 | high | $0.000081 | 32s | |
| 3 | mistralai/mistral-nemo | 8.0 | 4/4 | high | $0.000081 | 30s | |
| 4 | ibm-granite/granite-4.0-h-micro | 8.0 | 4/4 | high | $0.000129 | 35s | |
| 5 | sao10k/l3-lunaris-8b | 8.0 | 4/4 | high | $0.000155 | 8s | |
| 6 | openai/gpt-oss-20b | 8.0 | 4/4 | high | $0.000185 | 19s | |
| 7 | qwen/qwen-2.5-7b-instruct | 8.0 | 4/4 | high | $0.000190 | 8s | |
| 8 | openai/gpt-oss-120b | 8.0 | 4/4 | high | $0.000195 | 35s | |
| 9 | amazon/nova-micro-v1 | 8.0 | 4/4 | high | $0.000203 | 6s | |
| 10 | google/gemma-3-12b-it | 8.0 | 4/4 | high | $0.000255 | 21s | |
| 11 | deepseek/deepseek-v4-flash | 8.0 | 4/4 | high | $0.000396 | 22s | |
| 12 | google/gemma-4-31b-it | 8.0 | 4/4 | high | $0.000425 | 21s | |
| 13 | qwen/qwen3-32b | 8.0 | 4/4 | high | $0.000436 | 19s | |
| 14 | meta-llama/llama-3.3-70b-instruct | 8.0 | 4/4 | high | $0.000524 | 22s | |
| 15 | openai/gpt-4.1-nano | 8.0 | 4/4 | high | $0.000580 | 7s | |
| 16 | qwen/qwen3-235b-a22b-2507 | 8.0 | 4/4 | high | $0.000655 | 14s | |
| 17 | openai/gpt-4o-mini | 8.0 | 4/4 | high | $0.000870 | 13s | |
| 18 | meta-llama/llama-4-maverick | 8.0 | 4/4 | high | $0.001160 | 17s | |
| 19 | deepseek/deepseek-chat | 8.0 | 4/4 | high | $0.001161 | 37s | |
| 20 | openai/gpt-4.1-mini | 8.0 | 4/4 | high | $0.002320 | 19s | |
| 21 | google/gemini-2.5-flash | 8.0 | 4/4 | high | $0.002650 | 8s | |
| 22 | anthropic/claude-haiku-4.5 | 8.0 | 4/4 | high | $0.006500 | 15s | |
| 23 | nex-agi/nex-n2-mini | — | 0/0 | — | $0.000652 | 33s | 502: {'code': 'empty_model_response', 'message': |
| 24 | openai/gpt-5-nano | — | 0/0 | — | $0.001935 | 78s | 502: {'code': 'empty_model_response', 'message': |
SQL Optimization
4 agent turns + scripted follow-ups and reviewer close · deterministic task rubric.
| # | Model | Score | Criteria | Trust | Est. cost | Time | Scorecard |
|---|---|---|---|---|---|---|---|
| 1 | inclusionai/ling-2.6-flash | 8.0 | 6/6 | high | $0.000102 | 8s | |
| 2 | meta-llama/llama-3.1-8b-instruct | 8.0 | 6/6 | high | $0.000162 | 63s | |
| 3 | mistralai/mistral-nemo | 8.0 | 6/6 | high | $0.000162 | 59s | |
| 4 | ibm-granite/granite-4.0-h-micro | 8.0 | 6/6 | high | $0.000259 | 94s | |
| 5 | sao10k/l3-lunaris-8b | 8.0 | 6/6 | high | $0.000310 | 24s | |
| 6 | openai/gpt-oss-20b | 8.0 | 6/6 | high | $0.000370 | 81s | |
| 7 | qwen/qwen-2.5-7b-instruct | 8.0 | 6/6 | high | $0.000380 | 32s | |
| 8 | openai/gpt-oss-120b | 8.0 | 6/6 | high | $0.000390 | 108s | |
| 9 | amazon/nova-micro-v1 | 8.0 | 6/6 | high | $0.000406 | 9s | |
| 10 | google/gemma-3-12b-it | 8.0 | 6/6 | high | $0.000510 | 46s | |
| 11 | deepseek/deepseek-v4-flash | 8.0 | 6/6 | high | $0.000792 | 103s | |
| 12 | google/gemma-4-31b-it | 8.0 | 6/6 | high | $0.000850 | 100s | |
| 13 | qwen/qwen3-32b | 8.0 | 6/6 | high | $0.000872 | 129s | |
| 14 | meta-llama/llama-3.3-70b-instruct | 8.0 | 6/6 | high | $0.001048 | 106s | |
| 15 | openai/gpt-4.1-nano | 8.0 | 6/6 | high | $0.001160 | 51s | |
| 16 | qwen/qwen3-235b-a22b-2507 | 8.0 | 6/6 | high | $0.001310 | 166s | |
| 17 | openai/gpt-4o-mini | 8.0 | 6/6 | high | $0.001740 | 37s | |
| 18 | meta-llama/llama-4-maverick | 8.0 | 6/6 | high | $0.002320 | 15s | |
| 19 | deepseek/deepseek-chat | 8.0 | 6/6 | high | $0.002321 | 146s | |
| 20 | openai/gpt-4.1-mini | 8.0 | 6/6 | high | $0.004640 | 26s | |
| 21 | google/gemini-2.5-flash | 8.0 | 6/6 | high | $0.005300 | 21s | |
| 22 | anthropic/claude-haiku-4.5 | 8.0 | 6/6 | high | $0.0130 | 28s | |
| 23 | nex-agi/nex-n2-mini | — | 0/0 | — | $0.000652 | 29s | 502: {'code': 'empty_model_response', 'message': |
| 24 | openai/gpt-5-nano | — | 0/0 | — | $0.001935 | 74s | 502: {'code': 'empty_model_response', 'message': |
Schema Drift Review
4 agent turns + scripted follow-ups and reviewer close · deterministic task rubric.
| # | Model | Score | Criteria | Trust | Est. cost | Time | Scorecard |
|---|---|---|---|---|---|---|---|
| 1 | inclusionai/ling-2.6-flash | 8.0 | 4/4 | high | $0.000102 | 18s | |
| 2 | mistralai/mistral-nemo | 8.0 | 4/4 | high | $0.000162 | 64s | |
| 3 | ibm-granite/granite-4.0-h-micro | 8.0 | 4/4 | high | $0.000259 | 99s | |
| 4 | sao10k/l3-lunaris-8b | 8.0 | 4/4 | high | $0.000310 | 21s | |
| 5 | openai/gpt-oss-20b | 8.0 | 4/4 | high | $0.000370 | 116s | |
| 6 | qwen/qwen-2.5-7b-instruct | 8.0 | 4/4 | high | $0.000380 | 33s | |
| 7 | openai/gpt-oss-120b | 8.0 | 4/4 | high | $0.000390 | 115s | |
| 8 | amazon/nova-micro-v1 | 8.0 | 4/4 | high | $0.000406 | 14s | |
| 9 | google/gemma-3-12b-it | 8.0 | 4/4 | high | $0.000510 | 51s | |
| 10 | deepseek/deepseek-v4-flash | 8.0 | 4/4 | high | $0.000792 | 115s | |
| 11 | google/gemma-4-31b-it | 8.0 | 4/4 | high | $0.000850 | 94s | |
| 12 | qwen/qwen3-32b | 8.0 | 4/4 | high | $0.000872 | 113s | |
| 13 | meta-llama/llama-3.3-70b-instruct | 8.0 | 4/4 | high | $0.001048 | 188s | |
| 14 | openai/gpt-4.1-nano | 8.0 | 4/4 | high | $0.001160 | 39s | |
| 15 | qwen/qwen3-235b-a22b-2507 | 8.0 | 4/4 | high | $0.001310 | 191s | |
| 16 | openai/gpt-4o-mini | 8.0 | 4/4 | high | $0.001740 | 40s | |
| 17 | meta-llama/llama-4-maverick | 8.0 | 4/4 | high | $0.002320 | 53s | |
| 18 | deepseek/deepseek-chat | 8.0 | 4/4 | high | $0.002321 | 81s | |
| 19 | openai/gpt-4.1-mini | 8.0 | 4/4 | high | $0.004640 | 52s | |
| 20 | google/gemini-2.5-flash | 8.0 | 4/4 | high | $0.005300 | 22s | |
| 21 | anthropic/claude-haiku-4.5 | 8.0 | 4/4 | high | $0.0130 | 30s | |
| 22 | meta-llama/llama-3.1-8b-instruct | 7.2 | 3/4 | high | $0.000162 | 97s | |
| 23 | nex-agi/nex-n2-mini | — | 0/0 | — | $0.000652 | 106s | 504: LLM call timed out after 35.0s |
| 24 | openai/gpt-5-nano | — | 0/0 | — | $0.001935 | 105s | 504: LLM call timed out after 35.0s |
dbt Test Failure
4 agent turns + scripted follow-ups and reviewer close · deterministic task rubric.
| # | Model | Score | Criteria | Trust | Est. cost | Time | Scorecard |
|---|---|---|---|---|---|---|---|
| 1 | inclusionai/ling-2.6-flash | 8.4 | 4/4 | high | $0.000102 | 15s | |
| 2 | sao10k/l3-lunaris-8b | 8.4 | 4/4 | high | $0.000310 | 23s | |
| 3 | openai/gpt-oss-20b | 8.4 | 4/4 | high | $0.000370 | 62s | |
| 4 | qwen/qwen-2.5-7b-instruct | 8.4 | 4/4 | high | $0.000380 | 24s | |
| 5 | openai/gpt-oss-120b | 8.4 | 4/4 | high | $0.000390 | 83s | |
| 6 | amazon/nova-micro-v1 | 8.4 | 4/4 | high | $0.000406 | 17s | |
| 7 | google/gemma-3-12b-it | 8.4 | 4/4 | high | $0.000510 | 58s | |
| 8 | deepseek/deepseek-v4-flash | 8.4 | 4/4 | high | $0.000792 | 81s | |
| 9 | google/gemma-4-31b-it | 8.4 | 4/4 | high | $0.000850 | 201s | |
| 10 | qwen/qwen3-32b | 8.4 | 4/4 | high | $0.000872 | 101s | |
| 11 | openai/gpt-4.1-nano | 8.4 | 4/4 | high | $0.001160 | 20s | |
| 12 | qwen/qwen3-235b-a22b-2507 | 8.4 | 4/4 | high | $0.001310 | 173s | |
| 13 | openai/gpt-4o-mini | 8.4 | 4/4 | high | $0.001740 | 32s | |
| 14 | deepseek/deepseek-chat | 8.4 | 4/4 | high | $0.002321 | 90s | |
| 15 | openai/gpt-4.1-mini | 8.4 | 4/4 | high | $0.004640 | 33s | |
| 16 | google/gemini-2.5-flash | 8.4 | 4/4 | high | $0.005300 | 21s | |
| 17 | anthropic/claude-haiku-4.5 | 8.4 | 4/4 | high | $0.0130 | 32s | |
| 18 | meta-llama/llama-3.1-8b-instruct | 8.0 | 4/4 | high | $0.000162 | 40s | |
| 19 | mistralai/mistral-nemo | 8.0 | 4/4 | high | $0.000162 | 50s | |
| 20 | ibm-granite/granite-4.0-h-micro | 8.0 | 4/4 | high | $0.000259 | 101s | |
| 21 | meta-llama/llama-3.3-70b-instruct | 8.0 | 4/4 | high | $0.001048 | 79s | |
| 22 | meta-llama/llama-4-maverick | 8.0 | 4/4 | high | $0.002320 | 32s | |
| 23 | nex-agi/nex-n2-mini | — | 0/0 | — | $0.000073 | 55s | 502: {'code': 'empty_model_response', 'message': |
| 24 | openai/gpt-5-nano | — | 0/0 | — | $0.001935 | 77s | 502: {'code': 'empty_model_response', 'message': |
Freshness SLA Breach
4 agent turns + scripted follow-ups and reviewer close · deterministic task rubric.
| # | Model | Score | Criteria | Trust | Est. cost | Time | Scorecard |
|---|---|---|---|---|---|---|---|
| 1 | inclusionai/ling-2.6-flash | 8.0 | 4/4 | high | $0.000102 | 15s | |
| 2 | meta-llama/llama-3.1-8b-instruct | 8.0 | 4/4 | high | $0.000162 | 161s | |
| 3 | mistralai/mistral-nemo | 8.0 | 4/4 | high | $0.000162 | 166s | |
| 4 | ibm-granite/granite-4.0-h-micro | 8.0 | 4/4 | high | $0.000259 | 99s | |
| 5 | sao10k/l3-lunaris-8b | 8.0 | 4/4 | high | $0.000310 | 28s | |
| 6 | openai/gpt-oss-20b | 8.0 | 4/4 | high | $0.000370 | 54s | |
| 7 | qwen/qwen-2.5-7b-instruct | 8.0 | 4/4 | high | $0.000380 | 35s | |
| 8 | openai/gpt-oss-120b | 8.0 | 4/4 | high | $0.000390 | 62s | |
| 9 | amazon/nova-micro-v1 | 8.0 | 4/4 | high | $0.000406 | 21s | |
| 10 | google/gemma-3-12b-it | 8.0 | 4/4 | high | $0.000510 | 49s | |
| 11 | deepseek/deepseek-v4-flash | 8.0 | 4/4 | high | $0.000792 | 143s | |
| 12 | google/gemma-4-31b-it | 8.0 | 4/4 | high | $0.000850 | 92s | |
| 13 | qwen/qwen3-32b | 8.0 | 4/4 | high | $0.000872 | 77s | |
| 14 | meta-llama/llama-3.3-70b-instruct | 8.0 | 4/4 | high | $0.001048 | 91s | |
| 15 | openai/gpt-4.1-nano | 8.0 | 4/4 | high | $0.001160 | 16s | |
| 16 | qwen/qwen3-235b-a22b-2507 | 8.0 | 4/4 | high | $0.001310 | 83s | |
| 17 | openai/gpt-4o-mini | 8.0 | 4/4 | high | $0.001740 | 31s | |
| 18 | meta-llama/llama-4-maverick | 8.0 | 4/4 | high | $0.002320 | 38s | |
| 19 | deepseek/deepseek-chat | 8.0 | 4/4 | high | $0.002321 | 87s | |
| 20 | openai/gpt-4.1-mini | 8.0 | 4/4 | high | $0.004640 | 39s | |
| 21 | google/gemini-2.5-flash | 8.0 | 4/4 | high | $0.005300 | 23s | |
| 22 | anthropic/claude-haiku-4.5 | 8.0 | 4/4 | high | $0.0130 | 33s | |
| 23 | nex-agi/nex-n2-mini | — | 0/0 | — | $0.000073 | 35s | 502: {'code': 'empty_model_response', 'message': |
| 24 | openai/gpt-5-nano | — | 0/0 | — | $0.001935 | 76s | 502: {'code': 'empty_model_response', 'message': |
Pipeline Incident Triage
6 agent turns + scripted follow-ups and reviewer close · deterministic task rubric.
| # | Model | Score | Criteria | Trust | Est. cost | Time | Scorecard |
|---|---|---|---|---|---|---|---|
| 1 | inclusionai/ling-2.6-flash | 8.0 | 3/3 | high | $0.000153 | 16s | |
| 2 | meta-llama/llama-3.1-8b-instruct | 8.0 | 3/3 | high | $0.000243 | 55s | |
| 3 | mistralai/mistral-nemo | 8.0 | 3/3 | high | $0.000243 | 153s | |
| 4 | ibm-granite/granite-4.0-h-micro | 8.0 | 3/3 | high | $0.000388 | 144s | |
| 5 | sao10k/l3-lunaris-8b | 8.0 | 3/3 | high | $0.000465 | 33s | |
| 6 | openai/gpt-oss-20b | 8.0 | 3/3 | high | $0.000555 | 121s | |
| 7 | qwen/qwen-2.5-7b-instruct | 8.0 | 3/3 | high | $0.000570 | 54s | |
| 8 | openai/gpt-oss-120b | 8.0 | 3/3 | high | $0.000585 | 134s | |
| 9 | amazon/nova-micro-v1 | 8.0 | 3/3 | high | $0.000609 | 20s | |
| 10 | google/gemma-3-12b-it | 8.0 | 3/3 | high | $0.000765 | 86s | |
| 11 | deepseek/deepseek-v4-flash | 8.0 | 3/3 | high | $0.001188 | 159s | |
| 12 | google/gemma-4-31b-it | 8.0 | 3/3 | high | $0.001275 | 181s | |
| 13 | qwen/qwen3-32b | 8.0 | 3/3 | high | $0.001308 | 87s | |
| 14 | meta-llama/llama-3.3-70b-instruct | 8.0 | 3/3 | high | $0.001572 | 209s | |
| 15 | openai/gpt-4.1-nano | 8.0 | 3/3 | high | $0.001740 | 77s | |
| 16 | qwen/qwen3-235b-a22b-2507 | 8.0 | 3/3 | high | $0.001965 | 94s | |
| 17 | openai/gpt-4o-mini | 8.0 | 3/3 | high | $0.002610 | 59s | |
| 18 | meta-llama/llama-4-maverick | 8.0 | 3/3 | high | $0.003480 | 57s | |
| 19 | deepseek/deepseek-chat | 8.0 | 3/3 | high | $0.003482 | 117s | |
| 20 | openai/gpt-4.1-mini | 8.0 | 3/3 | high | $0.006960 | 48s | |
| 21 | google/gemini-2.5-flash | 8.0 | 3/3 | high | $0.007950 | 28s | |
| 22 | anthropic/claude-haiku-4.5 | 8.0 | 3/3 | high | $0.0195 | 46s | |
| 23 | nex-agi/nex-n2-mini | — | 0/0 | — | $0.000652 | 25s | 502: {'code': 'empty_model_response', 'message': |
| 24 | openai/gpt-5-nano | — | 0/0 | — | $0.001935 | 83s | 502: {'code': 'empty_model_response', 'message': |
Feature Leakage Review
6 agent turns + scripted follow-ups and reviewer close · deterministic task rubric.
| # | Model | Score | Criteria | Trust | Est. cost | Time | Scorecard |
|---|---|---|---|---|---|---|---|
| 1 | inclusionai/ling-2.6-flash | 8.0 | 3/3 | high | $0.000153 | 22s | |
| 2 | meta-llama/llama-3.1-8b-instruct | 8.0 | 3/3 | high | $0.000243 | 199s | |
| 3 | ibm-granite/granite-4.0-h-micro | 8.0 | 3/3 | high | $0.000388 | 143s | |
| 4 | sao10k/l3-lunaris-8b | 8.0 | 3/3 | high | $0.000465 | 35s | |
| 5 | qwen/qwen-2.5-7b-instruct | 8.0 | 3/3 | high | $0.000570 | 48s | |
| 6 | amazon/nova-micro-v1 | 8.0 | 3/3 | high | $0.000609 | 15s | |
| 7 | google/gemma-3-12b-it | 8.0 | 3/3 | high | $0.000765 | 87s | |
| 8 | deepseek/deepseek-v4-flash | 8.0 | 3/3 | high | $0.001188 | 102s | |
| 9 | google/gemma-4-31b-it | 8.0 | 3/3 | high | $0.001275 | 241s | |
| 10 | meta-llama/llama-3.3-70b-instruct | 8.0 | 3/3 | high | $0.001572 | 232s | |
| 11 | qwen/qwen3-235b-a22b-2507 | 8.0 | 3/3 | high | $0.001965 | 189s | |
| 12 | openai/gpt-4o-mini | 8.0 | 3/3 | high | $0.002610 | 53s | |
| 13 | meta-llama/llama-4-maverick | 8.0 | 3/3 | high | $0.003480 | 123s | |
| 14 | deepseek/deepseek-chat | 8.0 | 3/3 | high | $0.003482 | 159s | |
| 15 | openai/gpt-4.1-mini | 8.0 | 3/3 | high | $0.006960 | 78s | |
| 16 | anthropic/claude-haiku-4.5 | 8.0 | 3/3 | high | $0.0195 | 49s | |
| 17 | openai/gpt-oss-20b | 7.2 | 2/3 | high | $0.000555 | 189s | |
| 18 | openai/gpt-oss-120b | 7.2 | 2/3 | high | $0.000585 | 146s | |
| 19 | openai/gpt-4.1-nano | 7.2 | 2/3 | high | $0.001740 | 33s | |
| 20 | google/gemini-2.5-flash | 7.2 | 2/3 | high | $0.007950 | 34s | |
| 21 | mistralai/mistral-nemo | 6.8 | 3/3 | high | $0.000243 | 94s | |
| 22 | nex-agi/nex-n2-mini | — | 0/0 | — | $0.000652 | 27s | 502: {'code': 'empty_model_response', 'message': |
| 23 | openai/gpt-5-nano | — | 0/0 | — | $0.001935 | 92s | 502: {'code': 'empty_model_response', 'message': |
| 24 | qwen/qwen3-32b | — | 0/0 | — | $0.001962 | 94s | 504: LLM call timed out after 35.0s |
Lineage Impact Review
6 agent turns + scripted follow-ups and reviewer close · deterministic task rubric.
| # | Model | Score | Criteria | Trust | Est. cost | Time | Scorecard |
|---|---|---|---|---|---|---|---|
| 1 | inclusionai/ling-2.6-flash | 8.4 | 4/4 | high | $0.000153 | 24s | |
| 2 | meta-llama/llama-3.1-8b-instruct | 8.4 | 4/4 | high | $0.000243 | 91s | |
| 3 | mistralai/mistral-nemo | 8.4 | 4/4 | high | $0.000243 | 71s | |
| 4 | ibm-granite/granite-4.0-h-micro | 8.4 | 4/4 | high | $0.000388 | 141s | |
| 5 | sao10k/l3-lunaris-8b | 8.4 | 4/4 | high | $0.000465 | 40s | |
| 6 | openai/gpt-oss-20b | 8.4 | 4/4 | high | $0.000555 | 120s | |
| 7 | qwen/qwen-2.5-7b-instruct | 8.4 | 4/4 | high | $0.000570 | 78s | |
| 8 | openai/gpt-oss-120b | 8.4 | 4/4 | high | $0.000585 | 224s | |
| 9 | amazon/nova-micro-v1 | 8.4 | 4/4 | high | $0.000609 | 18s | |
| 10 | google/gemma-3-12b-it | 8.4 | 4/4 | high | $0.000765 | 90s | |
| 11 | deepseek/deepseek-v4-flash | 8.4 | 4/4 | high | $0.001188 | 206s | |
| 12 | google/gemma-4-31b-it | 8.4 | 4/4 | high | $0.001275 | 198s | |
| 13 | qwen/qwen3-32b | 8.4 | 4/4 | high | $0.001308 | 213s | |
| 14 | meta-llama/llama-3.3-70b-instruct | 8.4 | 4/4 | high | $0.001572 | 98s | |
| 15 | openai/gpt-4.1-nano | 8.4 | 4/4 | high | $0.001740 | 83s | |
| 16 | qwen/qwen3-235b-a22b-2507 | 8.4 | 4/4 | high | $0.001965 | 165s | |
| 17 | openai/gpt-4o-mini | 8.4 | 4/4 | high | $0.002610 | 65s | |
| 18 | meta-llama/llama-4-maverick | 8.4 | 4/4 | high | $0.003480 | 111s | |
| 19 | deepseek/deepseek-chat | 8.4 | 4/4 | high | $0.003482 | 117s | |
| 20 | openai/gpt-4.1-mini | 8.4 | 4/4 | high | $0.006960 | 76s | |
| 21 | google/gemini-2.5-flash | 8.4 | 4/4 | high | $0.007950 | 30s | |
| 22 | anthropic/claude-haiku-4.5 | 8.4 | 4/4 | high | $0.0195 | 47s | |
| 23 | nex-agi/nex-n2-mini | — | 0/0 | — | $0.000652 | 32s | 502: {'code': 'empty_model_response', 'message': |
| 24 | openai/gpt-5-nano | — | 0/0 | — | $0.001935 | 106s | 504: LLM call timed out after 35.0s |
Train/Serve Skew
6 agent turns + scripted follow-ups and reviewer close · deterministic task rubric.
| # | Model | Score | Criteria | Trust | Est. cost | Time | Scorecard |
|---|---|---|---|---|---|---|---|
| 1 | inclusionai/ling-2.6-flash | 8.4 | 4/4 | high | $0.000153 | 24s | |
| 2 | meta-llama/llama-3.1-8b-instruct | 8.4 | 4/4 | high | $0.000243 | 109s | |
| 3 | mistralai/mistral-nemo | 8.4 | 4/4 | high | $0.000243 | 109s | |
| 4 | ibm-granite/granite-4.0-h-micro | 8.4 | 4/4 | high | $0.000388 | 125s | |
| 5 | sao10k/l3-lunaris-8b | 8.4 | 4/4 | high | $0.000465 | 32s | |
| 6 | openai/gpt-oss-20b | 8.4 | 4/4 | high | $0.000555 | 87s | |
| 7 | qwen/qwen-2.5-7b-instruct | 8.4 | 4/4 | high | $0.000570 | 81s | |
| 8 | openai/gpt-oss-120b | 8.4 | 4/4 | high | $0.000585 | 193s | |
| 9 | amazon/nova-micro-v1 | 8.4 | 4/4 | high | $0.000609 | 22s | |
| 10 | google/gemma-3-12b-it | 8.4 | 4/4 | high | $0.000765 | 88s | |
| 11 | deepseek/deepseek-v4-flash | 8.4 | 4/4 | high | $0.001188 | 196s | |
| 12 | google/gemma-4-31b-it | 8.4 | 4/4 | high | $0.001275 | 191s | |
| 13 | qwen/qwen3-32b | 8.4 | 4/4 | high | $0.001308 | 121s | |
| 14 | meta-llama/llama-3.3-70b-instruct | 8.4 | 4/4 | high | $0.001572 | 169s | |
| 15 | openai/gpt-4.1-nano | 8.4 | 4/4 | high | $0.001740 | 25s | |
| 16 | qwen/qwen3-235b-a22b-2507 | 8.4 | 4/4 | high | $0.001965 | 196s | |
| 17 | openai/gpt-4o-mini | 8.4 | 4/4 | high | $0.002610 | 49s | |
| 18 | meta-llama/llama-4-maverick | 8.4 | 4/4 | high | $0.003480 | 144s | |
| 19 | deepseek/deepseek-chat | 8.4 | 4/4 | high | $0.003482 | 97s | |
| 20 | openai/gpt-4.1-mini | 8.4 | 4/4 | high | $0.006960 | 67s | |
| 21 | google/gemini-2.5-flash | 8.4 | 4/4 | high | $0.007950 | 32s | |
| 22 | anthropic/claude-haiku-4.5 | 8.4 | 4/4 | high | $0.0195 | 46s | |
| 23 | nex-agi/nex-n2-mini | — | 0/0 | — | $0.000073 | 34s | 502: {'code': 'empty_model_response', 'message': |
| 24 | openai/gpt-5-nano | — | 0/0 | — | $0.001935 | 81s | 502: {'code': 'empty_model_response', 'message': |
For search engines and LLMs
Expert Task Benchmark: 24 OpenRouter models on 11 high-complexity multi-turn engineering tasks. 242 completed scorecard runs; total est. cost $0.4405.