Leave-behind · investor · executive · engineering · product

Download PDF

Evals are the scanner.
Ship/stop is the layer.

Evals grade the run. The ship decision lives in your CI. For the workflows you can't afford to get wrong when the model layer moves.

Everyone is already testing models and routing for cost-performance — eval adoption is the launchpad, not the finish line. Your stack inspects. What's missing is who answers still ship? on paths you author, every model change, in the pipeline you already merge through.

Why now

Measurement went mainstream (89% obs · 52.4% offline evals). The next model swap still needs before/after on your workflows — and a shared answer when one score isn't enough.

Why this

Portable ship/no-ship beside LangSmith and Braintrust — in your CI store, not their dashboard. Same maturation curve AppSec ran: scanner feed → continuous gate. That layer is still empty.

Same curve as AppSec

Security leaders don't ask "why Snyk when we have grep?" They ask who owns continuous prioritization and fix verification in CI. Same question for AI ship.

AppSec (solved shape)

Scanner feed

Prioritize + verify in CI

Snyk-class gate

AI ship (today)

Eval + trace feed

Ship/stop on your workflows in CI

← missing layer

Two layers. The model layer moves weekly. Measurement vendors won the inspect seat. Nobody owns ship/stop on workflows you author — in the CI you already merge through.
Inspect scaled. Decide didn't. LangChain N=1,340 (Nov–Dec 2025): 89% observability · 52.4% offline evals. Survey does not measure CI ship gates or fleet ship/stop.

Measurement (mainstream)

Evals
89%
Observability
89%
Traces
92%
Budget / FinOps
55%

Ship decision (empty)

Ship bar in CI
Shared owners
Partner-path corpus

No category owns continuous ship/still-trust on partner-authored paths in customer CI.

89%
LangChain obs · N=1,340
52.4%
Offline evals · not ship gates
24×11
Winner changes by workflow
422
OpenRouter routes (proxy)

AppSec parallel

AppSecAI ship
Thousands of CVEs / depsDozens of models + routers + silent swaps
Can't patch everythingCan't re-eval every workflow on every change
Scanning ≠ prioritized fixEval scores ≠ ship/stop on your paths
Not core: triage foreverNot core: pin workflows + synthesize weekly
Snyk in CIShip/still-trust in CI

Why doing this well is hard

Re-running evals after a swap is easy. Ship/stop on owned workflows in CI — when scores conflict — is what teams still do by hand. DIY works at small N; it evaporates at fleet scale.

What breaksSecurityAI ship
Volume10k CVEs/year10+ workflows × 10+ models × weekly motion
PrioritizationCVSS + critical assetsWhich workflows carry client risk
OwnershipSecurity owns the barHero eng owns the gold set
After deployContinuous scanEvals on merge; silent model drift unchecked
EvidenceTicket + scan reportSlack LGTM vs CI ship/stop artifact
Not coreAppSec buys SnykProduct eng shouldn't build ship governance
One line of arithmetic. Each dimension is linear; the burden is multiplicative. Freeze 20260711_215748: 24×11 = 264 model×workflow cells — best winner changes by workflow.
Wworkflows
Mmodels & routes
Cmotions / month
still ship?per month at fleet scale

The math compounds — and your ICP is on the steep part

Inspect tools scale with volume of traces (mostly linear). Ship/stop scales with workflows × models × motion — when every dimension grows with adoption, burden grows multiplicatively, not additively. 89% observability means more firms are crossing from hobbyist to fleet math every quarter.

ProfileWMC/mo W×M cellsW×M×C / moPer yearvs hobbyist
Demo / hobbyist 1 2 1 2 2 24 One path · one API · quarterly swap
Early prod 5 8 4 40 160 1,920 80× First agents live · router on
Real ICP 10 12 8 120 960 11,520 480× 3–10+ agents · weekly motion
Fleet ICP 20 15 16 300 4,800 57,600 2400× Multi-team · silent IDE defaults
Same formula, different universe. Fleet ICP ≈ 4,800 motion decisions/month — 2400× a hobbyist demo. That is why Slack + hero eng stops working: the seat is not a feature, it is capacity.
Demo / hobbyist
2/mo
Early prod
160/mo
Real ICP
960/mo
Fleet ICP
4,800/mo
Typical 18-month ICP trajectory (illustrative — eval launchpad → router → fleet). 30/mo → 3,360/mo ≈ 112× without anyone planning an explosion.
StageW × M × CMotion / mo
Month 0 · eval launchpad 3 × 5 × 2 30/mo
Month 6 · router + second agent 6 × 9 × 5 270/mo
Month 12 · real ICP 10 × 12 × 8 960/mo
Month 18 · fleet pressure 16 × 15 × 14 3,360/mo
Model layer accelerates too. OpenRouter public catalog (proxy for routable churn): 54 → 422 in ~20 months (~8×). Quadratic fit → ~1,041 routes by YE 2027 — more silent swaps, more "which path still clears?"
End 2024
54 routes · 1×
Aug 2026
422 routes · 8×
YE 2026 (forecast)
549 routes · 10×
YE 2027 (forecast)
1,041 routes · 19×

Forecast: same quadratic fit as internal model-churn visual · not the full universe (private weights, on-prem, direct lab APIs excluded).

Eng / platform You have evals like scanners. You don't have a continuous ship bar on your workflows when fifteen models and five agents change every month — that's a control function, not a script.

Product Experiment B won on the benchmark. Checkout, support, and SQL each need a different pass line. Nobody holds that in a spreadsheet when models move weekly.

Investor Inspect vendors won measurement. Routers won cost narrative. Decision layer on partner paths in customer CI is still empty.

Hobbyist vs real player

HobbyistReal ICP
Agents1 demo3–10+ in prod
ModelsOne APILabs + router + IDE defaults
ChangeQuarterlyWeekly
EvalsMaybeScores in CI — not fleet ship bar
DIYWorksGold set evaporates when author leaves

When to walk

One agent and one benchmark — build it yourself. We're for the team that already has Braintrust, a router, and five workflows — and still answers "still ship?" in Slack.

Free gate: pip install vantage-core · author 3 paths · bind to CI.

Sources: LangChain State of Agent Engineering (Nov–Dec 2025, N=1,340) · Vantage freeze 20260711_215748 · LangChain Switchyard Aug 2026 · OpenRouter catalog · September 2026 · vantageai.cc/runtimeai/ship-decision-layer