Skip to main content
Use Serverless SFT to fine-tune an LLM with supervised learning on curated input and output examples. Serverless SFT trains a low-rank adapter (LoRA) that specializes a base model for your task, stores the adapter as an artifact in your Weights & Biases account, and hosts every checkpoint on Serverless Inference. Serverless SFT is in public preview. You run Serverless SFT through OpenPipe’s ART framework or directly through the Serverless Training API.

Prerequisites

Before you start, make sure you have the following:
  • A Forge account. If you don’t have one, sign up.
  • An API key. In Forge, click your profile icon, select Settings, and click Create new API key. Copy the key when it is displayed, because you can’t view it again. The examples in these docs read it from the WANDB_API_KEY environment variable; the ART quickstart shows where ART expects it.
  • A project in Weights & Biases to record training metrics and store the trained adapter. See Projects.
  • A dataset of input and output examples that show the behavior you want the model to learn. The ART Serverless SFT documentation describes the expected format.
  • The ART framework, if you use it rather than calling the API directly. Follow the install steps in the ART quickstart.
Also review Usage information and limits to understand costs and restrictions, and choose a base model from the available models.

When to use Serverless SFT

Serverless SFT is a good fit when you have, or can produce, examples of the behavior you want:
  • Distillation: Transfer knowledge from a larger, more capable model into a smaller, faster one by training on the larger model’s outputs.
  • Teaching output style and format: Train a model to follow a specific response format, tone, or structure.
  • Warmup before RL: Give a model supervised examples before refining it with Serverless RL.
If you can judge an outcome but cannot write out the ideal answer in advance, use Serverless RL instead.

Train a model

ART wraps the Serverless Training API and manages datasets, training runs, and checkpoints for you. It is the recommended way to start. Follow the ART Serverless SFT documentation to prepare your dataset and start a training job.

After training

Every checkpoint you train is hosted automatically. To send inference requests to it, construct the model endpoint from your team, project, model name, and training step as described in Use your trained models. The same endpoint format applies to models trained with Serverless SFT and Serverless RL. To save storage, delete checkpoints you no longer need. See Usage information and limits
Last modified on September 30, 2026