Vantage RuntimeAI · Benchmarks

Expert Task Benchmark

OpenRouter models evaluated across analytical engineering tasks — multi-turn dialogue, deterministic rubrics, trust/repetition signals, and full scorecard PDFs per run.

24
Models tested
11
Complex scenarios
242
Scorecard runs
$0.4405
Total est. cost

What we tested

We evaluated 24 OpenRouter models across 11 multi-turn analytical engineering tasks with deterministic rubrics and trust signals.

  • SQL tuning · schema drift · dbt failures · freshness SLA · incident triage · feature leakage · lineage impact · train/serve skew
  • 2–6 agent turns per scenario with scripted counterpart follow-ups and explicit reviewer close
  • Every completed run links to the same scorecard PDF used in Simulator and leaderboard exports

Dot colorBlue · Best (≥7.0)Yellow · Average (4.0–6.9)Red · Weak (<4.0)

Top model by scenario

Start here: one winner per analytical task, with direct links to the scenario table and winning scorecard.

ScenarioTop modelScoreCriteriaEst. costScorecardDetails
Partition Filter Fixinclusionai/ling-2.6-flash8.05/5$0.000051Scorecard PDFView scenario
Experiment Readoutinclusionai/ling-2.6-flash7.64/4$0.000051Scorecard PDFView scenario
Null Rate Alertinclusionai/ling-2.6-flash8.04/4$0.000051Scorecard PDFView scenario
SQL Optimizationinclusionai/ling-2.6-flash8.06/6$0.000102Scorecard PDFView scenario
Schema Drift Reviewinclusionai/ling-2.6-flash8.04/4$0.000102Scorecard PDFView scenario
dbt Test Failureinclusionai/ling-2.6-flash8.44/4$0.000102Scorecard PDFView scenario
Freshness SLA Breachinclusionai/ling-2.6-flash8.04/4$0.000102Scorecard PDFView scenario
Pipeline Incident Triageinclusionai/ling-2.6-flash8.03/3$0.000153Scorecard PDFView scenario
Feature Leakage Reviewinclusionai/ling-2.6-flash8.03/3$0.000153Scorecard PDFView scenario
Lineage Impact Reviewinclusionai/ling-2.6-flash8.44/4$0.000153Scorecard PDFView scenario
Train/Serve Skewinclusionai/ling-2.6-flash8.44/4$0.000153Scorecard PDFView scenario

Cross-scenario mean score

#ModelMean /10Scenarios
1qwen/qwen3-32b8.110
2amazon/nova-micro-v18.111
3anthropic/claude-haiku-4.58.111
4deepseek/deepseek-chat8.111
5deepseek/deepseek-v4-flash8.111
6google/gemma-3-12b-it8.111
7google/gemma-4-31b-it8.111
8inclusionai/ling-2.6-flash8.111
9openai/gpt-4.1-mini8.111
10openai/gpt-4o-mini8.111
11qwen/qwen-2.5-7b-instruct8.111
12sao10k/l3-lunaris-8b8.111
13ibm-granite/granite-4.0-h-micro8.011
14meta-llama/llama-4-maverick8.011
15google/gemini-2.5-flash8.011
16nex-agi/nex-n2-mini8.01
17openai/gpt-4.1-nano8.011
18openai/gpt-oss-120b8.011
19openai/gpt-oss-20b8.011
20meta-llama/llama-3.1-8b-instruct8.011
21qwen/qwen3-235b-a22b-25078.011
22meta-llama/llama-3.3-70b-instruct7.911
23mistralai/mistral-nemo7.811

Partition Filter Fix

2 agent turns + scripted follow-ups and reviewer close · deterministic task rubric.

# Model Score Criteria Trust Est. cost Time Scorecard
1inclusionai/ling-2.6-flash8.05/5high$0.0000514sPDF
2meta-llama/llama-3.1-8b-instruct8.05/5high$0.0000817sPDF
3mistralai/mistral-nemo8.05/5high$0.00008116sPDF
4ibm-granite/granite-4.0-h-micro8.05/5high$0.00012929sPDF
5nex-agi/nex-n2-mini8.05/5high$0.00014528sPDF
6sao10k/l3-lunaris-8b8.05/5high$0.0001557sPDF
7openai/gpt-oss-20b8.05/5high$0.00018510sPDF
8qwen/qwen-2.5-7b-instruct8.05/5high$0.0001909sPDF
9openai/gpt-oss-120b8.05/5high$0.00019561sPDF
10amazon/nova-micro-v18.05/5high$0.0002032sPDF
11google/gemma-3-12b-it8.05/5high$0.00025515sPDF
12deepseek/deepseek-v4-flash8.05/5high$0.00039620sPDF
13google/gemma-4-31b-it8.05/5high$0.00042542sPDF
14qwen/qwen3-32b8.05/5high$0.00043644sPDF
15meta-llama/llama-3.3-70b-instruct8.05/5high$0.00052424sPDF
16openai/gpt-4.1-nano8.05/5high$0.0005805sPDF
17openai/gpt-4o-mini8.05/5high$0.00087010sPDF
18meta-llama/llama-4-maverick8.05/5high$0.00116019sPDF
19deepseek/deepseek-chat8.05/5high$0.00116116sPDF
20openai/gpt-4.1-mini8.05/5high$0.00232011sPDF
21google/gemini-2.5-flash8.05/5high$0.0026505sPDF
22anthropic/claude-haiku-4.58.05/5high$0.0065008sPDF
23qwen/qwen3-235b-a22b-25076.84/5high$0.00065556sPDF
24openai/gpt-5-nano0/0$0.00193591s502: {'code': 'empty_model_response', 'message':

Experiment Readout

2 agent turns + scripted follow-ups and reviewer close · deterministic task rubric.

# Model Score Criteria Trust Est. cost Time Scorecard
1inclusionai/ling-2.6-flash7.64/4high$0.0000513sPDF
2meta-llama/llama-3.1-8b-instruct7.64/4high$0.0000819sPDF
3ibm-granite/granite-4.0-h-micro7.64/4high$0.00012935sPDF
4sao10k/l3-lunaris-8b7.64/4high$0.00015511sPDF
5openai/gpt-oss-20b7.64/4high$0.00018516sPDF
6qwen/qwen-2.5-7b-instruct7.64/4high$0.0001907sPDF
7openai/gpt-oss-120b7.64/4high$0.00019522sPDF
8amazon/nova-micro-v17.64/4high$0.0002034sPDF
9google/gemma-3-12b-it7.64/4high$0.00025520sPDF
10deepseek/deepseek-v4-flash7.64/4high$0.00039633sPDF
11google/gemma-4-31b-it7.64/4high$0.0004259sPDF
12qwen/qwen3-32b7.64/4high$0.00043659sPDF
13openai/gpt-4.1-nano7.64/4high$0.0005808sPDF
14qwen/qwen3-235b-a22b-25077.64/4high$0.00065521sPDF
15openai/gpt-4o-mini7.64/4high$0.00087013sPDF
16meta-llama/llama-4-maverick7.64/4high$0.00116042sPDF
17deepseek/deepseek-chat7.64/4high$0.00116128sPDF
18openai/gpt-4.1-mini7.64/4high$0.00232012sPDF
19google/gemini-2.5-flash7.64/4high$0.0026506sPDF
20anthropic/claude-haiku-4.57.64/4high$0.00650014sPDF
21mistralai/mistral-nemo6.43/4high$0.00008117sPDF
22meta-llama/llama-3.3-70b-instruct6.43/4high$0.00052474sPDF
23nex-agi/nex-n2-mini0/0$0.00007339s502: {'code': 'empty_model_response', 'message':
24openai/gpt-5-nano0/0$0.00193589s502: {'code': 'empty_model_response', 'message':

Null Rate Alert

2 agent turns + scripted follow-ups and reviewer close · deterministic task rubric.

# Model Score Criteria Trust Est. cost Time Scorecard
1inclusionai/ling-2.6-flash8.04/4high$0.0000515sPDF
2meta-llama/llama-3.1-8b-instruct8.04/4high$0.00008132sPDF
3mistralai/mistral-nemo8.04/4high$0.00008130sPDF
4ibm-granite/granite-4.0-h-micro8.04/4high$0.00012935sPDF
5sao10k/l3-lunaris-8b8.04/4high$0.0001558sPDF
6openai/gpt-oss-20b8.04/4high$0.00018519sPDF
7qwen/qwen-2.5-7b-instruct8.04/4high$0.0001908sPDF
8openai/gpt-oss-120b8.04/4high$0.00019535sPDF
9amazon/nova-micro-v18.04/4high$0.0002036sPDF
10google/gemma-3-12b-it8.04/4high$0.00025521sPDF
11deepseek/deepseek-v4-flash8.04/4high$0.00039622sPDF
12google/gemma-4-31b-it8.04/4high$0.00042521sPDF
13qwen/qwen3-32b8.04/4high$0.00043619sPDF
14meta-llama/llama-3.3-70b-instruct8.04/4high$0.00052422sPDF
15openai/gpt-4.1-nano8.04/4high$0.0005807sPDF
16qwen/qwen3-235b-a22b-25078.04/4high$0.00065514sPDF
17openai/gpt-4o-mini8.04/4high$0.00087013sPDF
18meta-llama/llama-4-maverick8.04/4high$0.00116017sPDF
19deepseek/deepseek-chat8.04/4high$0.00116137sPDF
20openai/gpt-4.1-mini8.04/4high$0.00232019sPDF
21google/gemini-2.5-flash8.04/4high$0.0026508sPDF
22anthropic/claude-haiku-4.58.04/4high$0.00650015sPDF
23nex-agi/nex-n2-mini0/0$0.00065233s502: {'code': 'empty_model_response', 'message':
24openai/gpt-5-nano0/0$0.00193578s502: {'code': 'empty_model_response', 'message':

SQL Optimization

4 agent turns + scripted follow-ups and reviewer close · deterministic task rubric.

# Model Score Criteria Trust Est. cost Time Scorecard
1inclusionai/ling-2.6-flash8.06/6high$0.0001028sPDF
2meta-llama/llama-3.1-8b-instruct8.06/6high$0.00016263sPDF
3mistralai/mistral-nemo8.06/6high$0.00016259sPDF
4ibm-granite/granite-4.0-h-micro8.06/6high$0.00025994sPDF
5sao10k/l3-lunaris-8b8.06/6high$0.00031024sPDF
6openai/gpt-oss-20b8.06/6high$0.00037081sPDF
7qwen/qwen-2.5-7b-instruct8.06/6high$0.00038032sPDF
8openai/gpt-oss-120b8.06/6high$0.000390108sPDF
9amazon/nova-micro-v18.06/6high$0.0004069sPDF
10google/gemma-3-12b-it8.06/6high$0.00051046sPDF
11deepseek/deepseek-v4-flash8.06/6high$0.000792103sPDF
12google/gemma-4-31b-it8.06/6high$0.000850100sPDF
13qwen/qwen3-32b8.06/6high$0.000872129sPDF
14meta-llama/llama-3.3-70b-instruct8.06/6high$0.001048106sPDF
15openai/gpt-4.1-nano8.06/6high$0.00116051sPDF
16qwen/qwen3-235b-a22b-25078.06/6high$0.001310166sPDF
17openai/gpt-4o-mini8.06/6high$0.00174037sPDF
18meta-llama/llama-4-maverick8.06/6high$0.00232015sPDF
19deepseek/deepseek-chat8.06/6high$0.002321146sPDF
20openai/gpt-4.1-mini8.06/6high$0.00464026sPDF
21google/gemini-2.5-flash8.06/6high$0.00530021sPDF
22anthropic/claude-haiku-4.58.06/6high$0.013028sPDF
23nex-agi/nex-n2-mini0/0$0.00065229s502: {'code': 'empty_model_response', 'message':
24openai/gpt-5-nano0/0$0.00193574s502: {'code': 'empty_model_response', 'message':

Schema Drift Review

4 agent turns + scripted follow-ups and reviewer close · deterministic task rubric.

# Model Score Criteria Trust Est. cost Time Scorecard
1inclusionai/ling-2.6-flash8.04/4high$0.00010218sPDF
2mistralai/mistral-nemo8.04/4high$0.00016264sPDF
3ibm-granite/granite-4.0-h-micro8.04/4high$0.00025999sPDF
4sao10k/l3-lunaris-8b8.04/4high$0.00031021sPDF
5openai/gpt-oss-20b8.04/4high$0.000370116sPDF
6qwen/qwen-2.5-7b-instruct8.04/4high$0.00038033sPDF
7openai/gpt-oss-120b8.04/4high$0.000390115sPDF
8amazon/nova-micro-v18.04/4high$0.00040614sPDF
9google/gemma-3-12b-it8.04/4high$0.00051051sPDF
10deepseek/deepseek-v4-flash8.04/4high$0.000792115sPDF
11google/gemma-4-31b-it8.04/4high$0.00085094sPDF
12qwen/qwen3-32b8.04/4high$0.000872113sPDF
13meta-llama/llama-3.3-70b-instruct8.04/4high$0.001048188sPDF
14openai/gpt-4.1-nano8.04/4high$0.00116039sPDF
15qwen/qwen3-235b-a22b-25078.04/4high$0.001310191sPDF
16openai/gpt-4o-mini8.04/4high$0.00174040sPDF
17meta-llama/llama-4-maverick8.04/4high$0.00232053sPDF
18deepseek/deepseek-chat8.04/4high$0.00232181sPDF
19openai/gpt-4.1-mini8.04/4high$0.00464052sPDF
20google/gemini-2.5-flash8.04/4high$0.00530022sPDF
21anthropic/claude-haiku-4.58.04/4high$0.013030sPDF
22meta-llama/llama-3.1-8b-instruct7.23/4high$0.00016297sPDF
23nex-agi/nex-n2-mini0/0$0.000652106s504: LLM call timed out after 35.0s
24openai/gpt-5-nano0/0$0.001935105s504: LLM call timed out after 35.0s

dbt Test Failure

4 agent turns + scripted follow-ups and reviewer close · deterministic task rubric.

# Model Score Criteria Trust Est. cost Time Scorecard
1inclusionai/ling-2.6-flash8.44/4high$0.00010215sPDF
2sao10k/l3-lunaris-8b8.44/4high$0.00031023sPDF
3openai/gpt-oss-20b8.44/4high$0.00037062sPDF
4qwen/qwen-2.5-7b-instruct8.44/4high$0.00038024sPDF
5openai/gpt-oss-120b8.44/4high$0.00039083sPDF
6amazon/nova-micro-v18.44/4high$0.00040617sPDF
7google/gemma-3-12b-it8.44/4high$0.00051058sPDF
8deepseek/deepseek-v4-flash8.44/4high$0.00079281sPDF
9google/gemma-4-31b-it8.44/4high$0.000850201sPDF
10qwen/qwen3-32b8.44/4high$0.000872101sPDF
11openai/gpt-4.1-nano8.44/4high$0.00116020sPDF
12qwen/qwen3-235b-a22b-25078.44/4high$0.001310173sPDF
13openai/gpt-4o-mini8.44/4high$0.00174032sPDF
14deepseek/deepseek-chat8.44/4high$0.00232190sPDF
15openai/gpt-4.1-mini8.44/4high$0.00464033sPDF
16google/gemini-2.5-flash8.44/4high$0.00530021sPDF
17anthropic/claude-haiku-4.58.44/4high$0.013032sPDF
18meta-llama/llama-3.1-8b-instruct8.04/4high$0.00016240sPDF
19mistralai/mistral-nemo8.04/4high$0.00016250sPDF
20ibm-granite/granite-4.0-h-micro8.04/4high$0.000259101sPDF
21meta-llama/llama-3.3-70b-instruct8.04/4high$0.00104879sPDF
22meta-llama/llama-4-maverick8.04/4high$0.00232032sPDF
23nex-agi/nex-n2-mini0/0$0.00007355s502: {'code': 'empty_model_response', 'message':
24openai/gpt-5-nano0/0$0.00193577s502: {'code': 'empty_model_response', 'message':

Freshness SLA Breach

4 agent turns + scripted follow-ups and reviewer close · deterministic task rubric.

# Model Score Criteria Trust Est. cost Time Scorecard
1inclusionai/ling-2.6-flash8.04/4high$0.00010215sPDF
2meta-llama/llama-3.1-8b-instruct8.04/4high$0.000162161sPDF
3mistralai/mistral-nemo8.04/4high$0.000162166sPDF
4ibm-granite/granite-4.0-h-micro8.04/4high$0.00025999sPDF
5sao10k/l3-lunaris-8b8.04/4high$0.00031028sPDF
6openai/gpt-oss-20b8.04/4high$0.00037054sPDF
7qwen/qwen-2.5-7b-instruct8.04/4high$0.00038035sPDF
8openai/gpt-oss-120b8.04/4high$0.00039062sPDF
9amazon/nova-micro-v18.04/4high$0.00040621sPDF
10google/gemma-3-12b-it8.04/4high$0.00051049sPDF
11deepseek/deepseek-v4-flash8.04/4high$0.000792143sPDF
12google/gemma-4-31b-it8.04/4high$0.00085092sPDF
13qwen/qwen3-32b8.04/4high$0.00087277sPDF
14meta-llama/llama-3.3-70b-instruct8.04/4high$0.00104891sPDF
15openai/gpt-4.1-nano8.04/4high$0.00116016sPDF
16qwen/qwen3-235b-a22b-25078.04/4high$0.00131083sPDF
17openai/gpt-4o-mini8.04/4high$0.00174031sPDF
18meta-llama/llama-4-maverick8.04/4high$0.00232038sPDF
19deepseek/deepseek-chat8.04/4high$0.00232187sPDF
20openai/gpt-4.1-mini8.04/4high$0.00464039sPDF
21google/gemini-2.5-flash8.04/4high$0.00530023sPDF
22anthropic/claude-haiku-4.58.04/4high$0.013033sPDF
23nex-agi/nex-n2-mini0/0$0.00007335s502: {'code': 'empty_model_response', 'message':
24openai/gpt-5-nano0/0$0.00193576s502: {'code': 'empty_model_response', 'message':

Pipeline Incident Triage

6 agent turns + scripted follow-ups and reviewer close · deterministic task rubric.

# Model Score Criteria Trust Est. cost Time Scorecard
1inclusionai/ling-2.6-flash8.03/3high$0.00015316sPDF
2meta-llama/llama-3.1-8b-instruct8.03/3high$0.00024355sPDF
3mistralai/mistral-nemo8.03/3high$0.000243153sPDF
4ibm-granite/granite-4.0-h-micro8.03/3high$0.000388144sPDF
5sao10k/l3-lunaris-8b8.03/3high$0.00046533sPDF
6openai/gpt-oss-20b8.03/3high$0.000555121sPDF
7qwen/qwen-2.5-7b-instruct8.03/3high$0.00057054sPDF
8openai/gpt-oss-120b8.03/3high$0.000585134sPDF
9amazon/nova-micro-v18.03/3high$0.00060920sPDF
10google/gemma-3-12b-it8.03/3high$0.00076586sPDF
11deepseek/deepseek-v4-flash8.03/3high$0.001188159sPDF
12google/gemma-4-31b-it8.03/3high$0.001275181sPDF
13qwen/qwen3-32b8.03/3high$0.00130887sPDF
14meta-llama/llama-3.3-70b-instruct8.03/3high$0.001572209sPDF
15openai/gpt-4.1-nano8.03/3high$0.00174077sPDF
16qwen/qwen3-235b-a22b-25078.03/3high$0.00196594sPDF
17openai/gpt-4o-mini8.03/3high$0.00261059sPDF
18meta-llama/llama-4-maverick8.03/3high$0.00348057sPDF
19deepseek/deepseek-chat8.03/3high$0.003482117sPDF
20openai/gpt-4.1-mini8.03/3high$0.00696048sPDF
21google/gemini-2.5-flash8.03/3high$0.00795028sPDF
22anthropic/claude-haiku-4.58.03/3high$0.019546sPDF
23nex-agi/nex-n2-mini0/0$0.00065225s502: {'code': 'empty_model_response', 'message':
24openai/gpt-5-nano0/0$0.00193583s502: {'code': 'empty_model_response', 'message':

Feature Leakage Review

6 agent turns + scripted follow-ups and reviewer close · deterministic task rubric.

# Model Score Criteria Trust Est. cost Time Scorecard
1inclusionai/ling-2.6-flash8.03/3high$0.00015322sPDF
2meta-llama/llama-3.1-8b-instruct8.03/3high$0.000243199sPDF
3ibm-granite/granite-4.0-h-micro8.03/3high$0.000388143sPDF
4sao10k/l3-lunaris-8b8.03/3high$0.00046535sPDF
5qwen/qwen-2.5-7b-instruct8.03/3high$0.00057048sPDF
6amazon/nova-micro-v18.03/3high$0.00060915sPDF
7google/gemma-3-12b-it8.03/3high$0.00076587sPDF
8deepseek/deepseek-v4-flash8.03/3high$0.001188102sPDF
9google/gemma-4-31b-it8.03/3high$0.001275241sPDF
10meta-llama/llama-3.3-70b-instruct8.03/3high$0.001572232sPDF
11qwen/qwen3-235b-a22b-25078.03/3high$0.001965189sPDF
12openai/gpt-4o-mini8.03/3high$0.00261053sPDF
13meta-llama/llama-4-maverick8.03/3high$0.003480123sPDF
14deepseek/deepseek-chat8.03/3high$0.003482159sPDF
15openai/gpt-4.1-mini8.03/3high$0.00696078sPDF
16anthropic/claude-haiku-4.58.03/3high$0.019549sPDF
17openai/gpt-oss-20b7.22/3high$0.000555189sPDF
18openai/gpt-oss-120b7.22/3high$0.000585146sPDF
19openai/gpt-4.1-nano7.22/3high$0.00174033sPDF
20google/gemini-2.5-flash7.22/3high$0.00795034sPDF
21mistralai/mistral-nemo6.83/3high$0.00024394sPDF
22nex-agi/nex-n2-mini0/0$0.00065227s502: {'code': 'empty_model_response', 'message':
23openai/gpt-5-nano0/0$0.00193592s502: {'code': 'empty_model_response', 'message':
24qwen/qwen3-32b0/0$0.00196294s504: LLM call timed out after 35.0s

Lineage Impact Review

6 agent turns + scripted follow-ups and reviewer close · deterministic task rubric.

# Model Score Criteria Trust Est. cost Time Scorecard
1inclusionai/ling-2.6-flash8.44/4high$0.00015324sPDF
2meta-llama/llama-3.1-8b-instruct8.44/4high$0.00024391sPDF
3mistralai/mistral-nemo8.44/4high$0.00024371sPDF
4ibm-granite/granite-4.0-h-micro8.44/4high$0.000388141sPDF
5sao10k/l3-lunaris-8b8.44/4high$0.00046540sPDF
6openai/gpt-oss-20b8.44/4high$0.000555120sPDF
7qwen/qwen-2.5-7b-instruct8.44/4high$0.00057078sPDF
8openai/gpt-oss-120b8.44/4high$0.000585224sPDF
9amazon/nova-micro-v18.44/4high$0.00060918sPDF
10google/gemma-3-12b-it8.44/4high$0.00076590sPDF
11deepseek/deepseek-v4-flash8.44/4high$0.001188206sPDF
12google/gemma-4-31b-it8.44/4high$0.001275198sPDF
13qwen/qwen3-32b8.44/4high$0.001308213sPDF
14meta-llama/llama-3.3-70b-instruct8.44/4high$0.00157298sPDF
15openai/gpt-4.1-nano8.44/4high$0.00174083sPDF
16qwen/qwen3-235b-a22b-25078.44/4high$0.001965165sPDF
17openai/gpt-4o-mini8.44/4high$0.00261065sPDF
18meta-llama/llama-4-maverick8.44/4high$0.003480111sPDF
19deepseek/deepseek-chat8.44/4high$0.003482117sPDF
20openai/gpt-4.1-mini8.44/4high$0.00696076sPDF
21google/gemini-2.5-flash8.44/4high$0.00795030sPDF
22anthropic/claude-haiku-4.58.44/4high$0.019547sPDF
23nex-agi/nex-n2-mini0/0$0.00065232s502: {'code': 'empty_model_response', 'message':
24openai/gpt-5-nano0/0$0.001935106s504: LLM call timed out after 35.0s

Train/Serve Skew

6 agent turns + scripted follow-ups and reviewer close · deterministic task rubric.

# Model Score Criteria Trust Est. cost Time Scorecard
1inclusionai/ling-2.6-flash8.44/4high$0.00015324sPDF
2meta-llama/llama-3.1-8b-instruct8.44/4high$0.000243109sPDF
3mistralai/mistral-nemo8.44/4high$0.000243109sPDF
4ibm-granite/granite-4.0-h-micro8.44/4high$0.000388125sPDF
5sao10k/l3-lunaris-8b8.44/4high$0.00046532sPDF
6openai/gpt-oss-20b8.44/4high$0.00055587sPDF
7qwen/qwen-2.5-7b-instruct8.44/4high$0.00057081sPDF
8openai/gpt-oss-120b8.44/4high$0.000585193sPDF
9amazon/nova-micro-v18.44/4high$0.00060922sPDF
10google/gemma-3-12b-it8.44/4high$0.00076588sPDF
11deepseek/deepseek-v4-flash8.44/4high$0.001188196sPDF
12google/gemma-4-31b-it8.44/4high$0.001275191sPDF
13qwen/qwen3-32b8.44/4high$0.001308121sPDF
14meta-llama/llama-3.3-70b-instruct8.44/4high$0.001572169sPDF
15openai/gpt-4.1-nano8.44/4high$0.00174025sPDF
16qwen/qwen3-235b-a22b-25078.44/4high$0.001965196sPDF
17openai/gpt-4o-mini8.44/4high$0.00261049sPDF
18meta-llama/llama-4-maverick8.44/4high$0.003480144sPDF
19deepseek/deepseek-chat8.44/4high$0.00348297sPDF
20openai/gpt-4.1-mini8.44/4high$0.00696067sPDF
21google/gemini-2.5-flash8.44/4high$0.00795032sPDF
22anthropic/claude-haiku-4.58.44/4high$0.019546sPDF
23nex-agi/nex-n2-mini0/0$0.00007334s502: {'code': 'empty_model_response', 'message':
24openai/gpt-5-nano0/0$0.00193581s502: {'code': 'empty_model_response', 'message':
For search engines and LLMs
Expert Task Benchmark: 24 OpenRouter models on 11 high-complexity multi-turn engineering tasks. 242 completed scorecard runs; total est. cost $0.4405.