Research
Runs we thought were worth writing up: what we measured, what it cost, and what we would still push back on. Every derived figure reproduces from the table printed with it. Where a run has no spread, or the arms were not identically configured, we say so in the piece rather than in a footnote.
Older measured results, across sales and CRM agents, operations workflows and warehouse-scale labeling, are on the benchmarks page.
Interested?
We are in private preview, working with a small group of design partners at a time: capture the traces, write the eval, and ship the first replacement model alongside the people who already know the work.
The fit is a production workflow where cost or latency is starting to hurt, and the people who know what good looks like are not on an ML team.
Book a callSend two months of inference spend and what currently counts as a correct answer. We do not need your data to start.
Prefer email? contact@protege.sh