Skip to main content
Use Serverless RL to post-train an LLM with reinforcement learning. You define a reward function that scores your agent’s outputs. Serverless RL collects batches of trajectories, updates a low-rank adapter (LoRA) so higher-scoring behavior becomes more likely, stores the adapter as an artifact in your Weights & Biases account, and hosts every checkpoint on Serverless Inference. Serverless RL is in public preview. You run Serverless RL through OpenPipe’s ART framework or directly through the Serverless Training API.

Prerequisites

Before you start, make sure you have the following:
  • A Forge account. If you don’t have one, sign up.
  • An API key. In Forge, click your profile icon, select Settings, and click Create new API key. Copy the key when it is displayed, because you can’t view it again. The examples on this page read it from the WANDB_API_KEY environment variable; the ART quickstart shows where ART expects it.
  • A project in Weights & Biases to record training metrics and store the trained adapter. See Projects.
  • A reward function that scores your agent’s outputs. Serverless RL trains toward whatever this function rewards, so it defines the task. The ART quickstart shows how to write one.
  • The ART framework, if you use it rather than calling the API directly. Follow the install steps in the ART quickstart.
Also review Usage information and limits to understand costs and restrictions, and choose a base model from the available models.

Train an agent

ART wraps the Serverless Training API and manages rollouts, rewards, and checkpoints for you. It is the recommended way to start. Work through the ART quickstart, or open the example notebook, which trains an agent to play 2048 end to end.

Use your trained models

After you train a model, it is automatically available for inference. This section shows you how to construct the endpoint for a trained model and send requests to it, so you can integrate the model into your application or evaluation workflows. The same steps apply to models trained with Serverless SFT. To send requests to your trained model, you need the following:
  • Your API key. Create one in Forge under Settings.
  • The Serverless Training API base URL, https://forge.coreweave.com/api/training/v1/.
  • Your model’s endpoint.
The model’s endpoint uses the following schema:
The schema consists of:
  • Your entity, which is the name of your team.
  • The name of the project associated with your model.
  • The trained model’s name.
  • The training step of the model you want to deploy. This is usually the step where the model performed best in your evaluations.
For example, if your team is named email-specialists, your project is called mail-search, your trained model is named agent-001, and you want to deploy it on step 25, the endpoint looks like this:
After you have your endpoint, you can integrate it into your normal inference workflows. The following examples show how to make inference requests to your trained model using a cURL request or the Python OpenAI SDK. Choose the example that matches your environment.

cURL

OpenAI SDK

To save storage, delete checkpoints you no longer need. See Usage information and limits.
Last modified on September 28, 2026