Skip to main content
The CoreWeave Inference API provides programmatic control over inference gateways, model deployments, and capacity claims. The API is available at api.coreweave.com. This page covers cross-cutting topics: authentication, protocols, status values, error formats, and the OpenAI-compatible inference endpoint. For per-endpoint request and response schemas, see the per-operation pages under each service in the left sidebar:
The Inference API is versioned as v1alpha1. APIs may change before general availability.

Authentication

The Inference API uses bearer token authentication to identify the caller and authorize each request. All API requests must include a CoreWeave API access token in the Authorization header as a Bearer token. The token must belong to a user with the Inference Viewer or Inference Admin role, depending on the operation. Replace [API-TOKEN] with your CoreWeave API access token.
For details on obtaining an API token, see Manage API access tokens.

Protocol support

You can call the Inference API over several transport protocols, depending on your client tooling and performance needs:

Schema and SDKs

The Inference API is defined in Protobuf and published to CoreWeave’s public Buf Schema Registry, so you can install a generated client instead of writing HTTP calls by hand.
  • Public BSR module: buf.build/coreweave/inference (module name buf.build/coreweave/inference).
  • Generated SDKs: Inference SDKs on the Buf Schema Registry. Clients are available for the Connect, gRPC, and Protobuf ecosystems across languages including Go, Python, TypeScript, Java, Kotlin, Rust, and Swift.
  • Services: GatewayService, DeploymentService, and CapacityClaimService (package coreweave.inference.v1alpha1).
Select an ecosystem and language on the Buf Schema Registry to get the install command for your package manager. Generated clients target the same API host and Bearer token authentication described in Authentication.
Generated clients cover the Inference management API, which is what you use to create and manage gateways, deployments, and capacity claims. To send inference requests to a deployed model, use any OpenAI-compatible client against your gateway endpoint. See OpenAI-compatible endpoint.

Query parameters

List endpoints support the following query parameter:

Status values

Use these status values to determine the lifecycle state of a resource when polling or reconciling state in your application. All resources share a common set of status values: Each resource includes a conditions array in its status with detailed information about the current state, including timestamps, reasons, and human-readable messages.

Error responses

When a request fails, the API returns a structured error body that your client can parse to surface details to users or trigger retries. Error responses follow the standard format:

OpenAI-compatible endpoint

Deployed models expose an OpenAI-compatible completions endpoint through their associated gateway. The endpoint URL depends on the gateway’s routing strategy. See Gateways for details on how each routing strategy constructs the request URL. For a complete walkthrough, see the Getting started guide.
Last modified on August 17, 2026