- Serverless RL: Post-train models with reinforcement learning so they learn new behaviors and improve reliability, speed, and cost on multi-turn agentic tasks. You define a reward function that scores your agent’s outputs. Serverless RL splits the workflow into inference and training phases and multiplexes them across jobs to increase GPU utilization and reduce your training time and costs.
- Serverless SFT: Fine-tune models with supervised learning on curated input and output examples. Use SFT for distillation, for teaching output style and format, or to warm up a model before applying RL.
When to use each method
Serverless RL suits tasks where you can judge an outcome but cannot write out the ideal answer in advance, such as:- Voice agents
- Deep research assistants
- On-prem models
- Content marketing analysis agents
- Distillation: Transfer knowledge from a larger, more capable model into a smaller, faster one.
- Teaching output style and format: Train a model to follow a specific response format, tone, or structure.
- Warmup before RL: Give a model supervised examples before refining it with reinforcement learning.
Why Serverless Training
- Lower training costs: Serverless training multiplexes shared infrastructure across many users, skips the setup process for each job, and scales your GPU costs down to zero when you aren’t training. This reduces training costs significantly.
- Faster training time: Serverless training splits inference requests across many GPUs and provisions training infrastructure the moment you need it, so jobs finish sooner and you iterate faster.
- Automatic deployment: Serverless training deploys every checkpoint you train, so you don’t set up hosting infrastructure. You can access and test trained models immediately in local, staging, or production environments.
How Serverless Training uses Forge services
Serverless training combines the following Forge components:- Serverless Inference: Runs your models, including every trained checkpoint.
- Weights & Biases: Tracks performance metrics while a LoRA adapter trains.
- Artifacts: Stores and versions the LoRA adapters.
- Weave (optional): Shows how the model responds at each step of the training loop.
- Train a Qwen model with OpenPipe RULER and Weave Scorers.
- Track training progress and create custom plots in Weights & Biases.
- Evaluate the final results on a Weave leaderboard.
Serverless RL and Serverless SFT are in public preview. During the preview, you are charged only for inference usage and artifact storage. Adapter training is free during the preview period. See Usage information and limits.
Next steps
- Check the available models.
- Follow Use Serverless SFT or Use Serverless RL. Each starts with its prerequisites: a Forge account, an API key, and a project.
- Look up endpoints in the Serverless Training API reference.