Higher quality on the same benchmark
+13% eval score against Sonnet 4.6
Write-heavy CRM tasks on Zapier AutomationBench, partial credit, n=10. Bars run 0 to 0.70. Protégé scored 1.13× Sonnet's score at a quarter of its cost.
Find the repeated work in your AI spend. Prove the cheapest route that clears your eval. Keep the weights.
90× lower model-token price for the bounded, repeated, verifiable slice of the same workload.
Used by teams at


One agent loop turns a single user action into dozens of model calls. The calls are bounded, repeated and verifiable, and every one of them is billed at the price of a model that could have written the workflow instead of running it.
classify_intent
one bounded call, billed at the frontier rate
The same million calls, routed to an open specialist that clears the same eval: $200
Unit price is a blended frontier serving basis. The specialist figure is that same call at $0.20 per million tokens. The call volumes are a scale, not a customer's bill.
Output tokens bill at three to five times input. Cutting what the model generates is the highest-leverage cost line available, and it needs no training run to collect.
Volume, not novelty, is what you are paying for. The hard part was solved months ago. What remains is the same shape of call, over and over, at a price set for reasoning you no longer need.
Offline work lands first, because it carries no production surface, no latency budget and no rollback risk. It still yields the eval contract that everything after it is measured against.
Point us at two months of inference spend and the traces behind it. We find which workflows repeat, what they cost, and which ones are bounded enough to replace.
An eval contract fixes what a correct answer is, with n counts and standard deviations, before a single optimization runs. Without it, nothing that follows is measurable.
Only when the eval says you have to. We climb the ladder rung by rung and stop at the cheapest one that clears, so most bounded workflows never reach a training run at all.
Downloadable weights or a hosted endpoint, your choice. From there it feeds on its own production traces, which is where the next pass gets its data.
None of them is a prediction. All three already happened, and together they are what makes a specialist cheaper than a rented generalist today rather than eventually.
The best open model more than doubled on the Artificial Analysis Intelligence Index. The gap to frontier halved, from 13 points to 6.
GEPA-style prompt optimization, LoRA, supervised fine-tuning and rejection sampling now run in an afternoon on rented GPUs.
The team that knows the task no longer has to be an ML team.
Every agent loop turns one user action into dozens of model calls. Volume, not novelty, now dominates inference spend.
Volume is exactly what a specialist eats.
Every bar below is a measured run with its n count, its standard deviation and its cost accounting. The runs that lost are published too.
Higher quality on the same benchmark
Write-heavy CRM tasks on Zapier AutomationBench, partial credit, n=10. Bars run 0 to 0.70. Protégé scored 1.13× Sonnet's score at a quarter of its cost.
Fast enough to sit in the request path
30-task operations holdout, 90-trajectory production-style serving run. The same slice also came in 6.0× cheaper on measured token cost.
Cheap enough to run the whole table
39,962 rows labeled three ways, full cost accounting, bars to scale. Cheap did not mean worse: 99.30% three-way agreement, and the regulated data never left the warehouse.
A bigger blind GEPA budget lost, landing at 0.252 ±0.144. We published that too. The failures are why the rest of the numbers are believable.
Read the benchmarksTeams treat optimization as a menu. It is a ladder, and you only climb when the eval proves the rung below did not clear. Most bounded workflows stop by rung three.
Live software, days not weeks
Training, only once the eval demands it
Competitors start at rung five. That is the only rung they sell.
Enterprises get isolation. We get a compounding map of what works, without ever holding their data.
Once our eval defines what good means for a workflow, we are the arbiter of it. Switching vendors means rebuilding the definition of quality first.
Every engagement teaches which rung clears which task shape. That map compounds across customers regardless of what any single contract permits, because it contains no customer content.
A competitor arriving later starts cold against a model that has been improving on production traces for a year.
They optimize the model. We optimize the decision of whether to touch the model at all.
| Capability | Applied Compute | Osmosis | TrainLoop | Protégé |
|---|---|---|---|---|
| Learns from production traces | Yes | Yes | Not claimed | Yes |
| Continuous or automatic retraining | Yes | Hourly | Research only | Yes |
| Eval-gated ladder that can conclude "don't train" | Not claimed | Not claimed | Not claimed | Yes |
| Ships trainingless wins as an outcome | Not claimed | Not claimed | Not claimed | Yes |
| Offline, warehouse-scale batch | Not claimed | Not claimed | Not claimed | Yes |
"Not claimed" means the capability is absent from public materials, not that the vendor is incapable of it. The category is real and well funded: Applied Compute has raised $100M, Osmosis $7M, TrainLoop went through YC W25. All three sell training.
Tushar Jain, founder
Applied Scientist II at Amazon from 2021 to 2025, working on active learning to accelerate LLM training and large-scale optimization. ML Scientist then ML Engineer II at Verisk before that, shipping applied ML into regulated enterprise.
Seven people, already shipping
Founding engineering, ML engineering, software, product design and business operations. The router and the prompt harness are live; supervised fine-tuning and reinforcement learning are in beta.
How we publish
Every result on this site carries its n count, its standard deviation and its cost accounting. When a run loses, it goes on the benchmarks page with the runs that won.
The ladder is productized. Rungs two, three and four are live software, and every engagement returns reusable rungs to the route library, so forward-deployed hours per engagement fall as coverage rises. The self-learning endpoint is the terminal product.
They assembled it to fill their own capacity, because post-training is a feature that sells compute. We stay neutral across providers, we hand back the weights, and we lead with the eval rather than the cluster. Different buyer, different incentive.
It does not have to. The cross-customer asset is the route library, a map of which rung clears which task shape, and it contains zero customer content. Per-tenant learning stays isolated and is sold as the feature.
There is a human approval gate on by default. No checkpoint promotes without sign-off in regulated tenants, and every retrain is versioned, diffable and rollback-able with a full audit trail. Silent self-redeployment is an audit finding, not a feature.
The weights and the eval contract are yours to keep and to run anywhere. Regulated data stays inside your warehouse throughout; on the 39,962-row labeling run we only needed the problem shape, meaning the domain, the labels, a few examples and the eval contract.
We will show you which workflows are overpaying, what the cheapest route that clears your eval looks like, and what it would cost to own it.
Bring the spend breakdown and what currently counts as a correct answer. We do not need your data to start.
Prefer email? tushar@protege.sh