Don't rent their models. Own yours.

Find the repeated work in your AI spend. Prove the cheapest route that clears your eval. Keep the weights.

Frontier generalist $18.00/ M tokens
Open specialist $0.20/ M tokens

90× lower model-token price for the bounded, repeated, verifiable slice of the same workload.

Blended model-token serving basis, before optimization

Used by teams at

  • Sierra
  • Chatbase
  • LimeChat
  • Aidbase
  • TailorTalk
  • Lorikeet

You pay frontier prices for work that stopped being hard months ago.

One agent loop turns a single user action into dozens of model calls. The calls are bounded, repeated and verifiable, and every one of them is billed at the price of a model that could have written the workflow instead of running it.

classify_intent one bounded call, billed at the frontier rate
  1. 1 call $0.018
  2. 1,000 $18
  3. 100,000 $1,800
  4. 1,000,000 $18,000

The same million calls, routed to an open specialist that clears the same eval: $200

Unit price is a blended frontier serving basis. The specialist figure is that same call at $0.20 per million tokens. The call volumes are a scale, not a customer's bill.

Output tokens bill at three to five times input. Cutting what the model generates is the highest-leverage cost line available, and it needs no training run to collect.

Volume, not novelty, is what you are paying for. The hard part was solved months ago. What remains is the same shape of call, over and over, at a price set for reasoning you no longer need.

Four steps, and the fourth one starts the first again.

Offline work lands first, because it carries no production surface, no latency budget and no rollback risk. It still yields the eval contract that everything after it is measured against.

  1. 01

    Capture

    Point us at two months of inference spend and the traces behind it. We find which workflows repeat, what they cost, and which ones are bounded enough to replace.

  2. 02

    Evaluate

    An eval contract fixes what a correct answer is, with n counts and standard deviations, before a single optimization runs. Without it, nothing that follows is measurable.

  3. 03

    Train

    Only when the eval says you have to. We climb the ladder rung by rung and stop at the cheapest one that clears, so most bounded workflows never reach a training run at all.

  4. 04

    Deploy

    Downloadable weights or a hosted endpoint, your choice. From there it feeds on its own production traces, which is where the next pass gets its data.

Three curves crossed in the last twelve months.

None of them is a prediction. All three already happened, and together they are what makes a specialist cheaper than a rented generalist today rather than eventually.

Open weights crossed the line

The best open model more than doubled on the Artificial Analysis Intelligence Index. The gap to frontier halved, from 13 points to 6.

15 30 45 60 APR 2025 APR 2026 gap 13 gap 6 35 60 22 54 Frontier Best open model

Post-training got cheap

GEPA-style prompt optimization, LoRA, supervised fine-tuning and rejection sampling now run in an afternoon on rented GPUs.

The team that knows the task no longer has to be an ML team.

Agents made repetition the bill

Every agent loop turns one user action into dozens of model calls. Volume, not novelty, now dominates inference spend.

Volume is exactly what a specialist eats.

Three domains. Three published benchmarks.

Every bar below is a measured run with its n count, its standard deviation and its cost accounting. The runs that lost are published too.

Higher quality on the same benchmark

+13% eval score against Sonnet 4.6

Baseline open model 0.400
Sonnet 4.6 0.557
Protégé 0.630

Write-heavy CRM tasks on Zapier AutomationBench, partial credit, n=10. Bars run 0 to 0.70. Protégé scored 1.13× Sonnet's score at a quarter of its cost.

See the run

Fast enough to sit in the request path

5.2× lower p50 latency

Sonnet API 1.94s
Protégé, 8B on Fireworks 369ms

30-task operations holdout, 90-trajectory production-style serving run. The same slice also came in 6.0× cheaper on measured token cost.

See the ladder

Cheap enough to run the whole table

50× lower cost than Opus

Protégé, 30B $2.82
Sonnet $12
Opus $140

39,962 rows labeled three ways, full cost accounting, bars to scale. Cheap did not mean worse: 99.30% three-way agreement, and the regulated data never left the warehouse.

See the accounting

A bigger blind GEPA budget lost, landing at 0.252 ±0.144. We published that too. The failures are why the rest of the numbers are believable.

Read the benchmarks

Find the cheapest step that clears the eval.

Teams treat optimization as a menu. It is a ladder, and you only climb when the eval proves the rung below did not clear. Most bounded workflows stop by rung three.

Live software, days not weeks

01 Task contract and eval Hours. No run cost. Defines what correct means before anything is optimized. Live
02 Prompt and scaffold One day, low run cost. GEPA plus autoresearch over the task contract. Live
03 Output control One day, no run cost. Schema, prefill, and /no_think. Where most workflows stop. Live
04 Routing Days, low run cost. Send each task shape to the cheapest model that clears it. Live

Training, only once the eval demands it

05 Supervised fine-tuning One to two weeks, GPU cost. Beta
06 Reinforcement learning Two to six weeks, heavy GPU cost. Beta

Competitors start at rung five. That is the only rung they sell.

.954 .961 .968 .975 30k 100k 400k 1.4M TOKEN BURN PER EVAL, LOG SCALE raw 8B, 32.6s p50 rung 3, no training 18.8× fewer tokens, higher score + /no_think SFT, rung 5, +0.0066 for a training run
On our operations slice a trainingless JSON scaffold scored 0.9667. Adding supervised fine-tuning on top moved the aggregate by 0.0066 and cost a training run. Training first is the expensive wrong answer, and it is what most teams do first.

Anyone can fine-tune once. Three things compound.

Enterprises get isolation. We get a compounding map of what works, without ever holding their data.

QUALITY TIME every other API, frozen Protégé endpoint each step is production traces fed back
The same endpoint and the same integration, against a better model each cycle. Downloadable weights or hosted endpoint is the customer's choice, and the loop runs either way.

Eval contracts and trace corpora

Once our eval defines what good means for a workflow, we are the arbiter of it. Switching vendors means rebuilding the definition of quality first.

The route library, holding zero customer data

Every engagement teaches which rung clears which task shape. That map compounds across customers regardless of what any single contract permits, because it contains no customer content.

The self-learning loop

A competitor arriving later starts cold against a model that has been improving on production traces for a year.

Everyone else sells a training run.

They optimize the model. We optimize the decision of whether to touch the model at all.

Capability comparison between Applied Compute, Osmosis, TrainLoop and Protégé
Capability Applied Compute Osmosis TrainLoop Protégé
Learns from production traces YesYesNot claimedYes
Continuous or automatic retraining YesHourlyResearch onlyYes
Eval-gated ladder that can conclude "don't train" Not claimedNot claimedNot claimedYes
Ships trainingless wins as an outcome Not claimedNot claimedNot claimedYes
Offline, warehouse-scale batch Not claimedNot claimedNot claimedYes

"Not claimed" means the capability is absent from public materials, not that the vendor is incapable of it. The category is real and well funded: Applied Compute has raised $100M, Osmosis $7M, TrainLoop went through YC W25. All three sell training.

The benchmarks were produced by this team.

Tushar Jain, founder

Applied Scientist II at Amazon from 2021 to 2025, working on active learning to accelerate LLM training and large-scale optimization. ML Scientist then ML Engineer II at Verisk before that, shipping applied ML into regulated enterprise.

Seven people, already shipping

Founding engineering, ML engineering, software, product design and business operations. The router and the prompt harness are live; supervised fine-tuning and reinforcement learning are in beta.

How we publish

Every result on this site carries its n count, its standard deviation and its cost accounting. When a run loses, it goes on the benchmarks page with the runs that won.

The questions you should ask.

Isn't this consulting?

The ladder is productized. Rungs two, three and four are live software, and every engagement returns reusable rungs to the route library, so forward-deployed hours per engagement fall as coverage rises. The self-learning endpoint is the terminal product.

CoreWeave already has Weights & Biases, OpenPipe and the GPUs.

They assembled it to fill their own capacity, because post-training is a feature that sells compute. We stay neutral across providers, we hand back the weights, and we lead with the eval rather than the cluster. Different buyer, different incentive.

Our enterprise won't let you train on its data.

It does not have to. The cross-customer asset is the route library, a map of which rung clears which task shape, and it contains zero customer content. Per-tenant learning stays isolated and is sold as the feature.

Self-learning in a regulated environment?

There is a human approval gate on by default. No checkpoint promotes without sign-off in regulated tenants, and every retrain is versioned, diffable and rollback-able with a full audit trail. Silent self-redeployment is an audit finding, not a feature.

What happens to the labeled data and the weights if we stop working together?

The weights and the eval contract are yours to keep and to run anywhere. Regulated data stays inside your warehouse throughout; on the 39,962-row labeling run we only needed the problem shape, meaning the domain, the labels, a few examples and the eval contract.

Send us two months of inference spend.

We will show you which workflows are overpaying, what the cheapest route that clears your eval looks like, and what it would cost to own it.

Book a call

Bring the spend breakdown and what currently counts as a correct answer. We do not need your data to start.

Prefer email? tushar@protege.sh