Three domains, measured end to end.

Each result below carries its n count, its standard deviation and its full cost accounting. Where a run lost, the losing number is printed next to the winning one, because a benchmark page that only reports wins is a marketing page.

Sales and CRM agents, on Zapier AutomationBench

Two surfaces from the same benchmark: four reasoning-heavy API tasks, and seven write-heavy CRM tasks. Both scored with partial credit, both against a Sonnet 4.6 reference.

Reasoning-heavy API tasks, 4 tasks, partial credit bars run 0.00 to 0.50

Baseline 0.084 ±0.073 · $0.095
Sonnet 4.6 0.160 ±0.009 · $0.707
Protégé optimized 0.313 ±0.110 · $0.128 · n=10

Write-heavy CRM tasks, 7 tasks, partial credit bars run 0.00 to 1.00

Baseline 0.400 ±0.067 · $0.230
Sonnet 4.6 0.557 ±0.066 · $1.12
Protégé optimized 0.630 ±0.087 · n=10

The honest part: a bigger blind GEPA budget did not win. The 50-call autoresearch prompt landed at 0.252 ±0.144, below the run we shipped. The win came from reading the failed runs and writing small targeted adapters, not from spending more on search.

Zapier AutomationBench, sales surface. Sonnet 4.6 reference. Qwen 3.6 Plus on Fireworks. GEPA with an Opus 4.7 reflection LM. max_steps=10, max_tokens=4096. Cost figures are measured spend for the scored run, not list price.

Operations workflows, one rung at a time

Most of the waste here was output-mode waste, not missing intelligence. Each row adds one rung to the row above it.

Eval score, latency and token burn by rung on a 30-task operations holdout
Rung applied Score p50 latency Eval tokens Reduction
Raw 8B, no prefill 0.956032.60s1.39MBaseline
Plus JSON prefill, rung 3 0.96676.76s74k18.8× fewer
Plus /no_think 0.95985.34s38k36.5× fewer
SFT plus /no_think, rung 5 0.97335.52s39k+3.5pt strict pass

p50 latency, 90 trajectories

Sonnet API at 1.94s against 369 ms on a Fireworks-served 8B. A 5.2× gap on the same slice.

Measured token cost, same slice

$0.0400 on the Sonnet API against $0.0066 on the Fireworks 8B. A 6.0× gap.

Where the misses went

The remaining misses cluster rather than scatter, which is what makes them fixable. A clustered failure comes back as a repair map, not as noise.

A trainingless JSON scaffold scored 0.9667. Adding supervised fine-tuning on top moved the aggregate to 0.9733, a gain of 0.0066 in exchange for a full training run. Training first is the expensive wrong answer, and it is what most teams reach for first.

30-task holdout, 25 samples per task, with a 90-trajectory production-style serving validation on Fireworks. Scores are against a Sonnet reference normalized to 1.0000 on the same bounded slice.

Warehouse-scale labeling, 39,962 rows

Work that was not economically viable at frontier prices becomes a feature you can run monthly. The full table was labeled three ways so the cheap run could be checked against the expensive ones.

Full-table cost, all 39,962 rows linear scale

Protégé 30B $2.82
Sonnet $12 · 4.4× higher
Opus $140 · 50× higher

Agreement

99.30% three-way agreement with Sonnet and Opus on the sparse theater-intent label. Cheap did not mean worse.

Sparse label recall

0.92% positives found, meaning 368 rows, against Sonnet's 0.90% at 360 and Opus's 1.01% at 405.

Privacy

Regulated data never left the warehouse. We needed only the problem shape: the domain, the label set, a few examples and the eval contract.

39,962 non-empty YouTube comments in a Snowflake table, labeled three ways, with full cost accounting and disclosed cold-start caveats. Cost is measured spend for the complete pass, not a per-thousand-row extrapolation.

How we report

Three rules, applied to every run on this page.

Every number carries its spread

Scores are published with standard deviations and the n count behind them. A score without a spread is an anecdote.

Cost is measured, not listed

Dollar figures are what the run actually spent, including the tokens burned by evaluation itself.

Losing runs get printed

The 0.252 ±0.144 GEPA result is on this page because it is the run that explains why the shipped one works.

Private preview

We take on a few design partners at a time.

The fit: a production workflow where cost or latency is starting to hurt, and the people who know what good looks like are not on an ML team.

Book a call