The benchmarks were produced by this team.

Seven people, already shipping. The router and the prompt harness are live; supervised fine-tuning and reinforcement learning are in beta. Nobody here is between the work and the numbers.

Tushar Jain

Founder & CEO · Chief Scientist

Applied Scientist II at Amazon from 2021 to 2025, working on applied and research machine learning in production: active learning to accelerate LLM training, and large-scale optimization. Before that, ML Scientist then ML Engineer II at Verisk from 2018 to 2021, shipping applied ML into regulated enterprise where the audit trail matters as much as the model.

That sequence is the whole thesis. The expensive part of replacing a frontier model is not the training run, it is knowing what “correct” means for a specific workflow and proving it before you spend a GPU hour. Protégé productizes that judgment as an eval-gated ladder, and publishes the runs that lose alongside the ones that win.

Angel investor in XYMA Analytics.

The rest of the team

Engineering, research, design and operations. Small enough that the people who run the benchmarks are the people who answer your questions about them.

  • Abhishek ChauhanFounding Engineer
  • Amit KumarML Engineer
  • Ansh GajbhiyeSoftware Engineer
  • Anchal SharmaSoftware Engineer
  • Pooja NagarajProduct Design
  • Manya JainBusiness Operations
  • Reinforcement Learning EngineerOpen role

How we work

We publish the failures

Every result carries its n count, its standard deviation and its cost accounting. The 0.252 ±0.144 run that lost is on the benchmarks page next to the one that shipped.

We do not need your data to start

On a 39,962-row labeling engagement the regulated data never left the customer's warehouse. We needed the problem shape: domain, labels, a few examples, an eval contract.

You keep the weights

Downloadable weights or a hosted endpoint, your choice, and the eval contract is yours to run anywhere. Switching away from us should cost you nothing but the switch.

Private preview

We take on a few design partners at a time.

The fit: a production workflow where cost or latency is starting to hurt, and the people who know what good looks like are not on an ML team.

Book a call