Why use Serverless Inference for LoRAs
Serverless Inference for LoRAs offers the following benefits:- Upload once, deploy without managing servers.
- Track which version is live with artifact versioning.
- Update models by swapping small LoRA files instead of full model weights.
Workflow
At a high level, serving a custom LoRA involves three steps:- Upload your LoRA weights as a Weights & Biases artifact.
- Reference the artifact URI as your model name in the API.
- CoreWeave dynamically loads your weights for inference.
Prerequisites
You need the following:- A CoreWeave Forge API key.
- A Weights & Biases project.
- Python 3.8+ with the
openaiandwandbpackages:pip install wandb openai.
Add and use LoRAs
You can add LoRAs to your CoreWeave Forge account and start using them with the following methods. Choose the tab that matches where your LoRA was trained:- Upload a LoRA you trained elsewhere
- Train a new LoRA with Serverless RL
- Train a new LoRA with supervised fine-tuning
Upload your own custom LoRA directory as a Weights & Biases artifact. Use this method if you trained your LoRA elsewhere (local environment, cloud provider, or partner service).This Python code uploads your locally stored LoRA weights to Weights & Biases as a versioned artifact. It creates a
lora type artifact with the required metadata (base model and storage region), adds your LoRA files from a local directory, and logs it to your Weights & Biases project for use with inference.Key requirements for using an existing LoRA
To use your own LoRAs with Serverless Inference, ensure the following:- The LoRA must have been trained using one of the models listed in the Supported base models section.
- A LoRA saved in PEFT format as a
loratype artifact in your CoreWeave Forge account. - The LoRA must be stored in the
storage_region="coreweave-us"for low latency. - When you upload, include the name of the base model you trained it on (for example,
meta-llama/Llama-3.1-8B-Instruct). This ensures Serverless Inference loads it with the correct model.
Supported base models
Your LoRA must be trained against one of the following base models. Use the exact model ID string when settingwandb.base_model so Serverless Inference can pair your adapter with the correct base model at inference time.
Pricing
You pay only for storage and the inference you run, rather than for always-on servers or dedicated GPU instances. Pricing has two components:- Storage: You’re billed for the storage that holds your LoRA weights.
- Serverless Inference usage: Calls that use LoRA artifacts are billed at the same rates as standard model inference.