Skip to main content
The inference proxy exposes an OpenAI-compatible Chat Completions endpoint for project routing, provider credentials, parameter overrides, and training-grade traces.
Use a proxy model name as the request’s model value to route production traffic through that project, or use a registered provider model reference for a direct request.
See Connect your application for Python, JavaScript, and cURL examples.
Last modified on September 2, 2026