Leave-behind · investor · executive · engineering · product
Download PDFEvals are the scanner.
Ship/stop is the layer.
Evals grade the run. The ship decision lives in your CI. For the workflows you can't afford to get wrong when the model layer moves.
Everyone is already testing models and routing for cost-performance — eval adoption is the launchpad, not the finish line. Your stack inspects. What's missing is who answers still ship? on paths you author, every model change, in the pipeline you already merge through.
Measurement went mainstream (89% obs · 52.4% offline evals). The next model swap still needs before/after on your workflows — and a shared answer when one score isn't enough.
Portable ship/no-ship beside LangSmith and Braintrust — in your CI store, not their dashboard. Same maturation curve AppSec ran: scanner feed → continuous gate. That layer is still empty.
Same curve as AppSec
Security leaders don't ask "why Snyk when we have grep?" They ask who owns continuous prioritization and fix verification in CI. Same question for AI ship.
AppSec (solved shape)
Scanner feed
Prioritize + verify in CI
Snyk-class gate
AI ship (today)
Eval + trace feed
Ship/stop on your workflows in CI
← missing layer
Measurement (mainstream)
Ship decision (empty)
No category owns continuous ship/still-trust on partner-authored paths in customer CI.
AppSec parallel
| AppSec | AI ship |
|---|---|
| Thousands of CVEs / deps | Dozens of models + routers + silent swaps |
| Can't patch everything | Can't re-eval every workflow on every change |
| Scanning ≠ prioritized fix | Eval scores ≠ ship/stop on your paths |
| Not core: triage forever | Not core: pin workflows + synthesize weekly |
| Snyk in CI | Ship/still-trust in CI |
Why doing this well is hard
Re-running evals after a swap is easy. Ship/stop on owned workflows in CI — when scores conflict — is what teams still do by hand. DIY works at small N; it evaporates at fleet scale.
| What breaks | Security | AI ship |
|---|---|---|
| Volume | 10k CVEs/year | 10+ workflows × 10+ models × weekly motion |
| Prioritization | CVSS + critical assets | Which workflows carry client risk |
| Ownership | Security owns the bar | Hero eng owns the gold set |
| After deploy | Continuous scan | Evals on merge; silent model drift unchecked |
| Evidence | Ticket + scan report | Slack LGTM vs CI ship/stop artifact |
| Not core | AppSec buys Snyk | Product eng shouldn't build ship governance |
motion decisions / month ≈ W × M × C
W = client workflows you own · M = models & routes in play · C = model-layer motions / month (swap, router, prompt, silent default)
Pairwise surface (who beats whom on each path) = W × M cells — Vantage freeze: 264 cells, not one global winner.
The math compounds — and your ICP is on the steep part
Inspect tools scale with volume of traces (mostly linear). Ship/stop scales with workflows × models × motion — when every dimension grows with adoption, burden grows multiplicatively, not additively. 89% observability means more firms are crossing from hobbyist to fleet math every quarter.
| Profile | W | M | C/mo | W×M cells | W×M×C / mo | Per year | vs hobbyist | |
|---|---|---|---|---|---|---|---|---|
| Demo / hobbyist | 1 | 2 | 1 | 2 | 2 | 24 | 1× | One path · one API · quarterly swap |
| Early prod | 5 | 8 | 4 | 40 | 160 | 1,920 | 80× | First agents live · router on |
| Real ICP | 10 | 12 | 8 | 120 | 960 | 11,520 | 480× | 3–10+ agents · weekly motion |
| Fleet ICP | 20 | 15 | 16 | 300 | 4,800 | 57,600 | 2400× | Multi-team · silent IDE defaults |
| Stage | W × M × C | Motion / mo |
|---|---|---|
| Month 0 · eval launchpad | 3 × 5 × 2 | 30/mo |
| Month 6 · router + second agent | 6 × 9 × 5 | 270/mo |
| Month 12 · real ICP | 10 × 12 × 8 | 960/mo |
| Month 18 · fleet pressure | 16 × 15 × 14 | 3,360/mo |
Forecast: same quadratic fit as internal model-churn visual · not the full universe (private weights, on-prem, direct lab APIs excluded).
Eng / platform You have evals like scanners. You don't have a continuous ship bar on your workflows when fifteen models and five agents change every month — that's a control function, not a script.
Product Experiment B won on the benchmark. Checkout, support, and SQL each need a different pass line. Nobody holds that in a spreadsheet when models move weekly.
Investor Inspect vendors won measurement. Routers won cost narrative. Decision layer on partner paths in customer CI is still empty.
Hobbyist vs real player
| Hobbyist | Real ICP | |
|---|---|---|
| Agents | 1 demo | 3–10+ in prod |
| Models | One API | Labs + router + IDE defaults |
| Change | Quarterly | Weekly |
| Evals | Maybe | Scores in CI — not fleet ship bar |
| DIY | Works | Gold set evaporates when author leaves |
When to walk
One agent and one benchmark — build it yourself. We're for the team that already has Braintrust, a router, and five workflows — and still answers "still ship?" in Slack.
Free gate: pip install vantage-core · author 3 paths · bind to CI.
Sources: LangChain State of Agent Engineering (Nov–Dec 2025, N=1,340) · Vantage freeze 20260711_215748 · LangChain Switchyard Aug 2026 · OpenRouter catalog · September 2026 · vantageai.cc/runtimeai/ship-decision-layer