Research

Runs we thought were worth writing up: what we measured, what it cost, and what we would still push back on. Every derived figure reproduces from the table printed with it. Where a run has no spread, or the arms were not identically configured, we say so in the piece rather than in a footnote.

Older measured results, across sales and CRM agents, operations workflows and warehouse-scale labeling, are on the benchmarks page.

Interested?

We are in private preview, working with a small group of design partners at a time: capture the traces, write the eval, and ship the first replacement model alongside the people who already know the work.

The fit is a production workflow where cost or latency is starting to hurt, and the people who know what good looks like are not on an ML team.

Book a call

Send two months of inference spend and what currently counts as a correct answer. We do not need your data to start.

Prefer email? contact@protege.sh