Don't rent their models.
Own yours.

Custom AI models trained from your work.

One model per job: a sales model, a CX model, an invoice-extraction model, a ticket-triage model, a document-review model.

  • Router · Live
  • Prompt harness · Live
  • SFT + RL · Beta
QUALITY PRODUCTION TRACES → frontier API · frozen Protégé · improving
The only endpoint whose quality goes up between releases. Same integration, better model each cycle, fed by its own production traces.

Used by teams at

  • Sierra
  • Chatbase
  • LimeChat
  • Aidbase
  • TailorTalk
  • Lorikeet

You pay frontier prices for work that stopped being hard.

One agent loop turns a single user action into dozens of model calls, and they are the same few tasks over and over: classify this, route that, pull those fields. Each one is billed at the price of a model that could have written the workflow instead of running it.

classify_intent one task, priced at the frontier rate
Frontier generalist $18.00/ M tokens
Specialist for this task $0.20/ M tokens
Same task, same eval 90×cheaper
  1. 1 call$0.018
  2. 1,000$18
  3. 100,000$1,800
  4. 1,000,000$18,000

The same million calls, served by a model trained on this one task $200

Unit price is a blended frontier serving basis; the specialist figure is the same call at $0.20 per million tokens. The call counts are a scale, not a customer's bill.

Output tokens bill at three to five times input. Cutting what the model generates is the highest-leverage cost line you have, and it needs no training run to collect.

You are paying for volume, not intelligence. The hard part was solved months ago. What is left is the same shape of call, over and over, at a price set for reasoning the task no longer needs.

The model gets better every week you run it.

Four steps, and the fourth feeds the first. Every request your replaced workflow serves becomes training data for the next pass, which is why this is the only endpoint you own whose quality goes up between vendor releases.

production traces
feed the next pass

Step 01

Capture

Send two months of inference spend and the traces behind it. You get back a map of which workflows actually repeat, what each one costs you, and which are bounded enough to replace.

Request the route audit

Three ways a call gets cheaper.

Routing picks the model. Training builds one when no existing model will do. Compression cuts what every call carries. They compound, and none of them touch your integration.

Routing

Per request

Share of one workload

Your model46%
Small open24%
Mid open15%
Sonnet10%
Opus5%

Each request goes to the cheapest model that clears the eval for that task. The mix moves as the task’s own model gets better.

Illustrative shape, not a customer’s bill. Your split depends on your tasks.

Training

τ³-bench

Teachers → your model

Kimi K3 Opus 5 Qwen-3.8-Max 27BYours

Held-out score

27B base0.250
After training0.483
Claude Opus 50.543

Within six points of the frontier at 1/40th the cost per solved task.

29 held-out tasks, 4 trials each. Full accounting on the case study.

Token compression

36.5× fewer

Same task, same eval

Let me think through this step by step. First, I need to consider each of the available fields and determine which ones are relevant. {"intent":"refund","priority":2}

Raw 8B1.39M
Schema, prefill, /no_think38k

Reasoning the task no longer needs is the cheapest thing to remove, and it needs no training run. Score held at 0.9598 against 0.9560.

30-task operations holdout. Measured eval tokens, not an estimate.

The request never changes. The bill does.

Your integration is written once. Behind it, the same call gets routed down, then trained on its own traces, and the price per solved task falls while the score holds.

invoice_extraction Worked example · one workload, three points in its life

First week

“Extract the invoice total and due date.”

Served by
Frontier generalist
Eval score
0.91
Per solved task
$2.40

After the eval contract

“Extract the invoice total and due date.”

Served by
Small open model
Eval score
0.92
Per solved task
$0.31

After training

“Extract the invoice total and due date.”

Served by
Your model
Eval score
0.94
Per solved task
$0.09

The request text is identical in all three columns; that is the point. Figures are a worked example, not a customer’s account. The floor under them is measured: on τ³-bench a trained 27B reached $0.55 per solved task against Opus 5’s $23.22.

Your code, all three times

resp = client.chat.completions.create(
    model=MODEL,
    messages=messages,
    extra_body={"task": "invoice_extraction"},
)

Byte-identical across all three. MODEL is one config value, changed with you and reversible. No redeploy, no SDK change, no migration.

What moved behind it

Modelfrontiersmall openyours
Score0.910.920.94
$ / solve$2.40$0.31$0.09

Each response says which model served it, so a quality change is traceable to a route change instead of guessed at.

invoice_extraction Worked example · version history
Each version with its eval delta, cost delta, traces behind it and outcome
VersionEval$/solve TracesOutcome
v10.910$2.4012,400Baseline
v2+0.006−34%18,900Shipped
v3+0.002−11%24,100Shipped
c4−0.011−29%26,700Rejected, missed the eval
v4+0.009−18%31,500Shipped

Cheaper is necessary and not sufficient. c4 was 29% cheaper and did not ship, because it lost 0.011 on the contract. That row is the whole product: a candidate only reaches your traffic when it clears the eval and you accept it.

Three domains. Three published benchmarks.

Every figure below is a measured run, published with its n count, its standard deviation and its cost accounting. So are the runs that lost.

A bigger blind GEPA budget lost, landing at 0.252 ±0.144. We published that too. The failures are why the rest of the numbers are believable.

Read the benchmarks

Find the cheapest step that clears the eval.

Teams treat optimization as a menu. It is a ladder, and you only climb when the eval proves the rung below did not clear. Most bounded workflows stop by rung three.

Live software, days not weeks

01 Task contract and eval Hours. No run cost. Defines what correct means before anything is optimized. Live
02 Prompt and scaffold One day, low run cost. GEPA plus autoresearch over the task contract. Live
03 Output control One day, no run cost. Schema, prefill, and /no_think. Where most workflows stop. Live
04 Routing Days, low run cost. Send each task shape to the cheapest model that clears it. Live

Training, only once the eval demands it

05 Supervised fine-tuning One to two weeks, GPU cost. Beta
06 Reinforcement learning Two to six weeks, heavy GPU cost. Beta

Competitors start at rung five. That is the only rung they sell.

.954 .961 .968 .975 30k 100k 400k 1.4M TOKEN BURN PER EVAL, LOG SCALE raw 8B, 32.6s p50 rung 3, no training 18.8× fewer tokens, higher score + /no_think SFT, rung 5, +0.0066 for a training run
On our operations slice a trainingless JSON scaffold scored 0.9667. Adding supervised fine-tuning on top moved the aggregate by 0.0066 and cost a training run. Training first is the expensive wrong answer, and it is what most teams do first.

Anyone can fine‑tune once. Three things compound.

Enterprises get isolation. We get a compounding map of what works, without ever holding their data.

QUALITY TIME every other API, frozen Protégé endpoint each step is production traces fed back
The same endpoint and the same integration, against a better model each cycle. Downloadable weights or hosted endpoint is the customer's choice, and the loop runs either way.

Eval contracts and trace corpora

Once our eval defines what good means for a workflow, we are the arbiter of it. Switching vendors means rebuilding the definition of quality first.

The route library, holding zero customer data

Every engagement teaches which rung clears which task shape. That map compounds across customers regardless of what any single contract permits, because it contains no customer content.

The self-learning loop

A competitor arriving later starts cold against a model that has been improving on production traces for a year.

Everyone else sells a training run.

They optimize the model. We optimize the decision of whether to touch the model at all.

Capability comparison between Applied Compute, Osmosis, TrainLoop and Protégé
Capability Applied Compute Osmosis TrainLoop Protégé
Learns from production traces YesYesNot claimedYes
Continuous or automatic retraining YesHourlyResearch onlyYes
Eval-gated ladder that can conclude “don’t train” Not claimedNot claimedNot claimedYes
Ships trainingless wins as an outcome Not claimedNot claimedNot claimedYes
Offline, warehouse-scale batch Not claimedNot claimedNot claimedYes

“Not claimed” means the capability is absent from public materials, not that the vendor is incapable of it. The category is real and well funded: Applied Compute has raised $100M, Osmosis $7M, TrainLoop went through YC W25. All three sell training.

The benchmarks were produced by this team.

Tushar Jain, founder

Applied Scientist II at Amazon from 2021 to 2025, working on active learning to accelerate LLM training and large-scale optimization. ML Scientist then ML Engineer II at Verisk before that, shipping applied ML into regulated enterprise.

Seven people, already shipping

Founding engineering, ML engineering, software, product design and business operations. The router and the prompt harness are live; supervised fine-tuning and reinforcement learning are in beta.

How we publish

Every result on this site carries its n count, its standard deviation and its cost accounting. When a run loses, it goes on the benchmarks page with the runs that won.

The questions you should ask.

Isn't this consulting?

The ladder is productized. Rungs two, three and four are live software, and every engagement returns reusable rungs to the route library, so forward-deployed hours per engagement fall as coverage rises. The self-learning endpoint is the terminal product.

CoreWeave already has Weights & Biases, OpenPipe and the GPUs.

They assembled it to fill their own capacity, because post-training is a feature that sells compute. We stay neutral across providers, we hand back the weights, and we lead with the eval rather than the cluster. Different buyer, different incentive.

Our enterprise won't let you train on its data.

It does not have to. The cross-customer asset is the route library, a map of which rung clears which task shape, and it contains zero customer content. Per-tenant learning stays isolated and is sold as the feature.

Self-learning in a regulated environment?

There is a human approval gate on by default. No checkpoint promotes without sign-off in regulated tenants, and every retrain is versioned, diffable and rollback-able with a full audit trail. Silent self-redeployment is an audit finding, not a feature.

What happens to the labeled data and the weights if we stop working together?

The weights and the eval contract are yours to keep and to run anywhere. Regulated data stays inside your warehouse throughout; on the 39,962-row labeling run we only needed the problem shape, meaning the domain, the labels, a few examples and the eval contract.

Cheaper, or we don’t ship it.

A cheaper model reaches your traffic only after it clears the eval contract you defined, and only after you accept it. If nothing clears, nothing moves and you keep paying what you pay today. That is the outcome we publish alongside the wins.

Interested?

We are in private preview, working with a small group of design partners at a time: capture the traces, write the eval, and ship the first replacement model alongside the people who already know the work.

The fit is a production workflow where cost or latency is starting to hurt, and the people who know what good looks like are not on an ML team.

Book a call

Send two months of inference spend and what currently counts as a correct answer. We do not need your data to start.

Prefer email? contact@protege.sh