- Rollout: In RL post-training, a response, or a sequence of responses, that the current policy generates for scoring and training.
- Policy: The model being optimized.
- Hot-load: Loading full or delta checkpoints from CoreWeave AI Object Storage into a running deployment without restarting the inference processes or recreating their replicas. The Dedicated Inference API, the Cloud Console, and CoreWeave error messages all use this term.
dynamo-vllm runtime engine.
This guide shows you how to create a hot-load-enabled deployment, generate rollouts, publish full and delta checkpoints, and confirm that every replica is serving the new weights. It’s for ML engineers running RL post-training who want CoreWeave to serve rollouts and load new weights.
RL Rollouts is in preview. Hot-loading is enabled per organization and requires the
dynamo-vllm engine. The Dedicated Inference API is versioned as v1alpha1 and might change before general availability.Prerequisites
Before you begin, make sure you have:- A CoreWeave account with Dedicated Inference access enabled.
- Hot-loading enabled for your organization. This is gated separately from Dedicated Inference access. To request it, contact your CoreWeave representative or CoreWeave support with your organization ID. Until it’s enabled, creating a deployment with a
hotLoadsection fails withhot-load is not enabled for this organization. - Inference GPU quota for the instance type you plan to deploy on. To request quota, submit a support ticket.
- A CoreWeave API access token with the Inference Admin role to create deployments and start hot-loads. Polling hot-load and replica status needs only the Inference Viewer role.
- A gateway through which your trainer can submit inference requests. To create one, see Create a gateway. If the deployment uses more than one gateway, all of them must be in the same zone, or creating a hot-load fails.
- An AI Object Storage bucket holding your starting model weights. To make one, see Create a bucket, then grant the Dedicated Inference service account read access to it so the service can read your weights. Your trainer writes each new checkpoint to this same bucket as training proceeds, so the bucket policy must also allow
s3:PutObjectfor the credentials your trainer uploads with. For policy syntax, see Bucket policies. - The AWS CLI configured for AI Object Storage with a profile named
cw, for uploading checkpoints. curlandjqfor the API examples.- A trainer that writes checkpoints in
.safetensorsformat and can call the Dedicated Inference API to start hot-loads and poll their status.
How hot-loading fits your training loop
Your trainer runs rollout generation and weight updates as one loop:- Send inference requests to generate rollouts from the current policy.
- Score the generated responses with your reward function.
- Compute updated model weights.
- Publish a checkpoint to AI Object Storage and start a hot-load through the Dedicated Inference API.
- Generate further rollouts as replicas adopt the updated policy, and repeat.
Checkpoint identities
Each checkpoint you publish needs an identity: a name you choose that the service appends to the deployment’s hot-load path prefix to locate that checkpoint’s files. Identities are single path segments, and a training loop typically derives them from the step number. In the following example,step-000420 is the identity, and the s3:// scheme refers to your S3-compatible AI Object Storage bucket:
previousSnapshotIdentity and snapshotChain refer to checkpoints. A hot-load operation has its own ID, separate from the checkpoint identity. Retrying a checkpoint creates another operation that targets the same identity.
Full and delta checkpoints
The following table compares the two checkpoint types and when to use each:
A delta inherits the model configuration from its base checkpoint and depends on that base to be reconstructed. A full checkpoint followed by a sequence of deltas forms a checkpoint chain. For example, a full checkpoint at
step-000420, followed by deltas at step-000430 and step-000440, forms a chain whose reconstructed result is the policy at step-000440. Keep the baseline and every delta in the chain available in the bucket.
A chain holds at most 128 deltas. When you reach that limit, publish a full checkpoint to start a new chain.
Prepare checkpoints in a format and precision compatible with the serving deployment. When it admits a hot-load request, the service checks that every checkpoint in the resolved chain has its required files.
Policy consistency during an update
A hot-load runs asynchronously, and during the transition, replicas might serve different checkpoints. Existing replicas go through a rolling update: replicas that aren’t updating keep serving, and requests on an updating replica pause while its weights load. Replicas that restart or are added by scaling after the hot-load starts load the target checkpoint directly. The hot-load’spromptCachePolicy controls whether the key-value (KV) cache computed under the earlier weights is kept:
PROMPT_CACHE_POLICY_PRESERVE, the default, keeps the KV cache. Requests that have already begun prefill or decoding resume with their existing cache after the update, so a request that spans an update can include computation under two sets of weights.PROMPT_CACHE_POLICY_RESET_ALLdiscards the KV cache when the update completes, so no cache entry computed under the earlier weights is reused. Use it when the next rollout round needs to start clean.
Create a hot-load-enabled deployment
Create the deployment withruntime.engine set to dynamo-vllm and a hotLoad section that names the bucket and path prefix for your checkpoints. dynamo-vllm is the only engine that supports hot-loading.
To create the deployment in the Cloud Console instead, use the Hot loading section of the create deployment drawer. It has an enable toggle and a checkpoint URI field, which the Console parses into the bucket and path prefix. Hot-loading is available only with the Dynamo vLLM runtime, and selecting another runtime shows an inline error. In the edit drawer, the toggle and URI are read-only. After the Console creates the deployment, complete step 1 of the following procedure, run export CW_DEPLOYMENT_ID="[DEPLOYMENT-ID]" with the deployment’s ID, and continue from step 4.
To create the deployment with the API, follow these steps:
-
Set your API access token and the management API URL. Replace
[API-TOKEN]with your CoreWeave API access token. The rest of this guide reuses these variables: -
Create the deployment. Replace
[GATEWAY-ID]with the ID of your gateway. The deployment parameters endpoint returns valid values for[ENGINE-VERSION]and[INSTANCE-TYPE]. Replace[MODEL-NAME],[MODEL-PATH],[BUCKET-NAME], and[PATH-PREFIX]with the values for your model weights and checkpoint location:modelpoints at the weights the deployment starts from, andhotLoadpoints at the checkpoints your trainer publishes later. This example keeps both in one bucket, separated by path, but they can be different buckets.TRANSITION_MODE_ASYNCis the only supported transition mode. -
Save the deployment ID from the response:
If the ID prints as
null, the request failed, anddeployment-response.jsoncontains the error. -
Wait for the deployment to reach
STATUS_READY. For a polling loop, see Wait for the deployment to start. -
To send rollout requests, export your gateway’s primary endpoint URL, which is the first entry in
gateway.status.endpoints. Replace[GATEWAY-ID]with the ID of your gateway:
Generate rollouts
Send rollout requests to your gateway endpoint using the same OpenAI-compatible API as any other Dedicated Inference deployment. Replace[MODEL-NAME] with your deployment’s model name:
temperature above 0. Greedy decoding bypasses temperature and top-p masking, so the returned log probabilities wouldn’t reflect the distribution the policy sampled from. top_logprobs accepts at most 20 candidates per token.
The response is an OpenAI-compatible chat completion with token IDs and log probabilities for training:
Your trainer computes rewards separately from the inference response.
Load a full checkpoint
A full checkpoint starts a checkpoint chain, so the first checkpoint you hot-load onto a deployment is a full checkpoint. This example loads a full checkpoint with the identitystep-000420.
-
Upload the checkpoint’s
.safetensorsshards,config.json, andmodel.safetensors.index.jsonto a directory named for its identity. Replace[LOCAL-CHECKPOINT-DIR]with the local directory your trainer wrote the checkpoint to. Keep the files at the top level of that directory, because the service doesn’t look for them in subdirectories:If your trainer runs in a CoreWeave cluster, you can upload through the LOTA endpoint instead by adding--endpoint-url http://cwlota.comto the command. For endpoint settings, see Configure endpoints. -
Write the
CreateHotLoadrequest tohot-load.json:identitynames the checkpoint directory beneath the configured path prefix. The bucket and path prefix come from the deployment configuration, so the request doesn’t include a storage URL.PROMPT_CACHE_POLICY_PRESERVEkeeps the KV cache through the transition, which is also the default if you omitpromptCachePolicy. To discard the cache when the update completes, setPROMPT_CACHE_POLICY_RESET_ALL. For how each policy affects requests, see Policy consistency during an update. -
Submit the request and save the hot-load operation ID. The deployment must be in
STATUS_READY, and no other hot-load can be running on it:If the ID prints asnull, the request failed, andhot-load-response.jsoncontains the error. Receiving an operation ID means the service accepted the request, not that the replicas have loaded the weights. - Wait for the hot-load to finish.
Wait for the hot-load to finish
A hot-load is finished when the operation reaches a terminal state and every replica reportsREADY on the target identity. GetHotLoad reports the operation’s aggregate state, and ListHotLoadReplicas reports each replica separately.
-
Poll the operation until it reaches a terminal state. This loop checks every 5 seconds and stops when the state is
HOT_LOAD_STATE_COMPLETED,HOT_LOAD_STATE_FAILED, orHOT_LOAD_STATE_CANCELED, or after 15 minutes:If the final state isHOT_LOAD_STATE_FAILEDorHOT_LOAD_STATE_CANCELED, see Recover a failed hot-load. If the operation is stillHOT_LOAD_STATE_PENDINGorHOT_LOAD_STATE_IN_PROGRESSwhen the loop stops, run the loop again. -
Confirm that every replica serves the target checkpoint. Replica results are paginated, so this loop reads every page, prints each replica, and counts the replicas listed and the replicas that aren’t ready on
CW_TARGET_IDENTITY:When at least one replica is listed and the not-ready count is0, every replica is serving the new checkpoint. A replica inERRORstate prints itserrorReasonin the last column.
Load a delta checkpoint
Compute each delta checkpoint against a checkpoint that you already hot-loaded on the same deployment. This example loadsstep-000430 as a delta against step-000420.
Your trainer writes the delta checkpoint. For an implementation that writes the format RL Rollouts accepts, see the delta checkpoint writer and Delta Weight Sync guide in Slime, an open source RL post-training framework. When you use that writer, set --update-weight-delta-encoding xor and --update-weight-delta-checksum adler32, because Slime also supports encodings and checksums that RL Rollouts doesn’t accept.
CoreWeave offers recipes for RL Rollouts supporting Slime, Miles, NeMo RL, and other trainers on request.
A delta checkpoint directory holds one or more .safetensors shards and a model.safetensors.index.json file:
-
Shards: Each changed tensor is stored as a byte-level XOR delta against the same tensor in the base checkpoint, compressed as its own Zstandard frame. The shard’s safetensors
__metadata__header maps each changed tensor’s name to the Adler-32 checksum of its new full weight, which the service uses to verify the reconstructed tensor. -
Index:
weight_maplists only the tensors that changed, andmetadatadescribes the delta. The service identifies the delta and its base by theidentityandpreviousSnapshotIdentityin your hot-load request, not byversionandbase_version, so those fields can hold other values, such as the zero-padded version numbers that Slime writes:
-
Upload the delta’s
.safetensorsshards andmodel.safetensors.index.json. Replace[LOCAL-DELTA-DIR]with the local directory your trainer wrote the delta to, and keep the files at the top level of that directory: -
Write the request to
hot-load.json:previousSnapshotIdentitymust name the checkpoint the delta was computed against, and that checkpoint must have a completed hot-load on this deployment.COMPRESSION_FORMAT_ZSTDandCHECKSUM_FORMAT_ADLER32are the only formats the service accepts, and your delta artifacts must use them. -
Submit the request and save the new operation ID. If the ID prints as
null, checkhot-load-response.jsonfor the error: - To confirm that the delta loaded, run both loops in Wait for the hot-load to finish again.
Run strictly on-policy rounds
If your training algorithm requires every sample in a round to come from one policy, coordinate the boundary between rounds:- Stop submitting new rollout requests and wait for the current round’s requests and multi-turn trajectories to finish.
- Publish the next checkpoint and start a hot-load, as in Load a full checkpoint or Load a delta checkpoint. To keep the next round from reusing KV cache computed under the previous policy, set
promptCachePolicytoPROMPT_CACHE_POLICY_RESET_ALL. - Wait for the hot-load to finish, including the replica check.
- Resume rollout generation only after every replica is ready on the target checkpoint.
Recover a failed hot-load
If an update stops before every replica loads the target checkpoint, the operation reachesHOT_LOAD_STATE_FAILED, and the replicas might serve different policy versions.
To recover, follow these steps:
- Run the replica loop in Wait for the hot-load to finish to see each replica’s loaded identity and
errorReason. - Address the reported issue, such as an inaccessible checkpoint or a missing base artifact.
- Start a new hot-load for the same checkpoint identity, or for another compatible checkpoint, such as a full checkpoint. A deployment runs one hot-load at a time, so the previous operation must be in a terminal state first, and the deployment must be in
STATUS_READY. The same restriction blocks deployment configuration updates while a hot-load is running. - Repeat both checks in the wait procedure before you resume on-policy generation.
Monitor hot-loads in the Cloud Console
The Console shows hot-load configuration and progress, and polls while an operation is running. It’s read-only: start hot-loads through the API. To track checkpoint adoption across replicas, follow these steps:- Open your deployment in the Cloud Console.
- In the deployment details drawer, select the Hot loads tab.
- Review the hot-load configuration: whether hot-loading is enabled, the checkpoint URI, and the transition behavior.
- Review the operation. Current hot load covers pending and in-progress operations, and Latest hot load shows the newest completed, failed, or canceled one. Each reports the target checkpoint, snapshot type, status, and timestamps in your browser’s local time.
- Review Replica convergence for the count of ready replicas against the total, then the replica table for each replica’s name, state, served checkpoint, zone, transition time, and error reason.
View hot-load history
To list a deployment’s hot-loads, newest first:nextPageToken as pageToken, and keep the same parentDeploymentId filter. The final page omits nextPageToken.
Reference
The following sections list the hot-load API operations and states, with example responses.Hot-load API endpoints
The management API athttps://api.coreweave.com defines these hot-load operations:
Hot-load states
hotLoad.status.state reports the operation’s aggregate state:
Replica states
Each entry in aListHotLoadReplicas response reports one replica’s state:
Example responses
These examples show a delta hot-load forstep-000430 partway through the update.
GetHotLoad returns the operation’s specification, the resolved checkpoint chain (snapshotChain), and the aggregate state. The chain lists the full baseline first, followed by each delta needed to reconstruct the target. It doesn’t include replica status:
ListHotLoadReplicas reports the checkpoint each replica currently serves, so a rolling update shows replicas in different states:
Next steps
- Getting started with Dedicated Inference: Create the gateway and grant the bucket access that this guide depends on.
- CoreWeave AI Object Storage: Configure the bucket that holds your checkpoints.
- Run torchforge on SUNK: Run a GRPO training loop on CoreWeave.
- Scaling: Autoscaling and capacity behavior for the deployment you hot-load into.