Skip to main content
Model Distillation helps product teams turn successful AI interactions into smaller, faster, and less expensive models built for their product. It brings the complete improvement workflow into one place, from collecting examples to deploying a better model.

Get started

Connect your application and follow the complete path from real usage to a deployed model.

What Model Distillation provides

See how the proxy, Studio, training, evaluations, and routing work together.

API reference

Integrate with the management API or call the OpenAI-compatible inference proxy.

The basic idea

Connect an existing AI feature once. Model Distillation collects examples from real usage and lets you turn the best examples into training data. Train several models, compare their quality with the model you use today, and deploy the best one without rewriting your application.
Automation can run the entire improvement loop and deploy a winner only when it meets the quality bar you choose.
Start with llms.txt for a compact map of the customer journey and machine-readable resources. For implementation, give the agent:Ask the agent to use only documented public APIs and to show you any routing or compute-consuming request before it sends it.
Last modified on August 25, 2026