Skip to main content
RL Rollouts is a CoreWeave Dedicated Inference capability for serving reinforcement learning (RL) rollouts and hot-loading updated policy weights. This guide uses the following terms:
  • Rollout: In RL post-training, a response, or a sequence of responses, that the current policy generates for scoring and training.
  • Policy: The model being optimized.
  • Hot-load: Loading full or delta checkpoints from CoreWeave AI Object Storage into a running deployment without restarting the inference processes or recreating their replicas. The Dedicated Inference API, the Cloud Console, and CoreWeave error messages all use this term.
Your trainer generates and scores samples, updates the model, and publishes checkpoints. Dedicated Inference serves the rollouts and loads each new checkpoint, so you don’t operate the serving stack yourself. RL Rollouts uses NVIDIA Dynamo with vLLM, configured with the dynamo-vllm runtime engine. This guide shows you how to create a hot-load-enabled deployment, generate rollouts, publish full and delta checkpoints, and confirm that every replica is serving the new weights. It’s for ML engineers running RL post-training who want CoreWeave to serve rollouts and load new weights.
RL Rollouts is in preview. Hot-loading is enabled per organization and requires the dynamo-vllm engine. The Dedicated Inference API is versioned as v1alpha1 and might change before general availability.

Prerequisites

Before you begin, make sure you have:
  • A CoreWeave account with Dedicated Inference access enabled.
  • Hot-loading enabled for your organization. This is gated separately from Dedicated Inference access. To request it, contact your CoreWeave representative or CoreWeave support with your organization ID. Until it’s enabled, creating a deployment with a hotLoad section fails with hot-load is not enabled for this organization.
  • Inference GPU quota for the instance type you plan to deploy on. To request quota, submit a support ticket.
  • A CoreWeave API access token with the Inference Admin role to create deployments and start hot-loads. Polling hot-load and replica status needs only the Inference Viewer role.
  • A gateway through which your trainer can submit inference requests. To create one, see Create a gateway. If the deployment uses more than one gateway, all of them must be in the same zone, or creating a hot-load fails.
  • An AI Object Storage bucket holding your starting model weights. To make one, see Create a bucket, then grant the Dedicated Inference service account read access to it so the service can read your weights. Your trainer writes each new checkpoint to this same bucket as training proceeds, so the bucket policy must also allow s3:PutObject for the credentials your trainer uploads with. For policy syntax, see Bucket policies.
  • The AWS CLI configured for AI Object Storage with a profile named cw, for uploading checkpoints.
  • curl and jq for the API examples.
  • A trainer that writes checkpoints in .safetensors format and can call the Dedicated Inference API to start hot-loads and poll their status.
You create the deployment itself as part of this guide, because hot-loading has to be configured when the deployment is created.

How hot-loading fits your training loop

Your trainer runs rollout generation and weight updates as one loop:
  1. Send inference requests to generate rollouts from the current policy.
  2. Score the generated responses with your reward function.
  3. Compute updated model weights.
  4. Publish a checkpoint to AI Object Storage and start a hot-load through the Dedicated Inference API.
  5. Generate further rollouts as replicas adopt the updated policy, and repeat.
CoreWeave manages rollout serving and weight loading. You manage the trainer, reward computation, checkpoint creation, and the policy-consistency requirements of your training algorithm.

Checkpoint identities

Each checkpoint you publish needs an identity: a name you choose that the service appends to the deployment’s hot-load path prefix to locate that checkpoint’s files. Identities are single path segments, and a training loop typically derives them from the step number. In the following example, step-000420 is the identity, and the s3:// scheme refers to your S3-compatible AI Object Storage bucket:
The Dedicated Inference API calls a checkpoint a snapshot, so fields such as previousSnapshotIdentity and snapshotChain refer to checkpoints. A hot-load operation has its own ID, separate from the checkpoint identity. Retrying a checkpoint creates another operation that targets the same identity.

Full and delta checkpoints

The following table compares the two checkpoint types and when to use each: A delta inherits the model configuration from its base checkpoint and depends on that base to be reconstructed. A full checkpoint followed by a sequence of deltas forms a checkpoint chain. For example, a full checkpoint at step-000420, followed by deltas at step-000430 and step-000440, forms a chain whose reconstructed result is the policy at step-000440. Keep the baseline and every delta in the chain available in the bucket. A chain holds at most 128 deltas. When you reach that limit, publish a full checkpoint to start a new chain. Prepare checkpoints in a format and precision compatible with the serving deployment. When it admits a hot-load request, the service checks that every checkpoint in the resolved chain has its required files.

Policy consistency during an update

A hot-load runs asynchronously, and during the transition, replicas might serve different checkpoints. Existing replicas go through a rolling update: replicas that aren’t updating keep serving, and requests on an updating replica pause while its weights load. Replicas that restart or are added by scaling after the hot-load starts load the target checkpoint directly. The hot-load’s promptCachePolicy controls whether the key-value (KV) cache computed under the earlier weights is kept:
  • PROMPT_CACHE_POLICY_PRESERVE, the default, keeps the KV cache. Requests that have already begun prefill or decoding resume with their existing cache after the update, so a request that spans an update can include computation under two sets of weights.
  • PROMPT_CACHE_POLICY_RESET_ALL discards the KV cache when the update completes, so no cache entry computed under the earlier weights is reused. Use it when the next rollout round needs to start clean.
Replica status reports the checkpoint each replica serves. It doesn’t identify the policy used for each token in a response. If your training algorithm tolerates samples from different policy versions, you can keep submitting rollout requests during an update. If it requires strictly on-policy samples, pause between rollout rounds as described in Run strictly on-policy rounds.

Create a hot-load-enabled deployment

Create the deployment with runtime.engine set to dynamo-vllm and a hotLoad section that names the bucket and path prefix for your checkpoints. dynamo-vllm is the only engine that supports hot-loading.
The hotLoad configuration is immutable after the deployment is created. To add or change it, create a new deployment. When you update the deployment, include the same hotLoad section, because an update that omits or changes it is rejected.
To create the deployment in the Cloud Console instead, use the Hot loading section of the create deployment drawer. It has an enable toggle and a checkpoint URI field, which the Console parses into the bucket and path prefix. Hot-loading is available only with the Dynamo vLLM runtime, and selecting another runtime shows an inline error. In the edit drawer, the toggle and URI are read-only. After the Console creates the deployment, complete step 1 of the following procedure, run export CW_DEPLOYMENT_ID="[DEPLOYMENT-ID]" with the deployment’s ID, and continue from step 4. To create the deployment with the API, follow these steps:
  1. Set your API access token and the management API URL. Replace [API-TOKEN] with your CoreWeave API access token. The rest of this guide reuses these variables:
  2. Create the deployment. Replace [GATEWAY-ID] with the ID of your gateway. The deployment parameters endpoint returns valid values for [ENGINE-VERSION] and [INSTANCE-TYPE]. Replace [MODEL-NAME], [MODEL-PATH], [BUCKET-NAME], and [PATH-PREFIX] with the values for your model weights and checkpoint location:
    model points at the weights the deployment starts from, and hotLoad points at the checkpoints your trainer publishes later. This example keeps both in one bucket, separated by path, but they can be different buckets. TRANSITION_MODE_ASYNC is the only supported transition mode.
  3. Save the deployment ID from the response:
    If the ID prints as null, the request failed, and deployment-response.json contains the error.
  4. Wait for the deployment to reach STATUS_READY. For a polling loop, see Wait for the deployment to start.
  5. To send rollout requests, export your gateway’s primary endpoint URL, which is the first entry in gateway.status.endpoints. Replace [GATEWAY-ID] with the ID of your gateway:

Generate rollouts

Send rollout requests to your gateway endpoint using the same OpenAI-compatible API as any other Dedicated Inference deployment. Replace [MODEL-NAME] with your deployment’s model name:
Set temperature above 0. Greedy decoding bypasses temperature and top-p masking, so the returned log probabilities wouldn’t reflect the distribution the policy sampled from. top_logprobs accepts at most 20 candidates per token. The response is an OpenAI-compatible chat completion with token IDs and log probabilities for training: Your trainer computes rewards separately from the inference response.

Load a full checkpoint

A full checkpoint starts a checkpoint chain, so the first checkpoint you hot-load onto a deployment is a full checkpoint. This example loads a full checkpoint with the identity step-000420.
  1. Upload the checkpoint’s .safetensors shards, config.json, and model.safetensors.index.json to a directory named for its identity. Replace [LOCAL-CHECKPOINT-DIR] with the local directory your trainer wrote the checkpoint to. Keep the files at the top level of that directory, because the service doesn’t look for them in subdirectories:
    If your trainer runs in a CoreWeave cluster, you can upload through the LOTA endpoint instead by adding --endpoint-url http://cwlota.com to the command. For endpoint settings, see Configure endpoints.
  2. Write the CreateHotLoad request to hot-load.json:
    identity names the checkpoint directory beneath the configured path prefix. The bucket and path prefix come from the deployment configuration, so the request doesn’t include a storage URL. PROMPT_CACHE_POLICY_PRESERVE keeps the KV cache through the transition, which is also the default if you omit promptCachePolicy. To discard the cache when the update completes, set PROMPT_CACHE_POLICY_RESET_ALL. For how each policy affects requests, see Policy consistency during an update.
  3. Submit the request and save the hot-load operation ID. The deployment must be in STATUS_READY, and no other hot-load can be running on it:
    If the ID prints as null, the request failed, and hot-load-response.json contains the error. Receiving an operation ID means the service accepted the request, not that the replicas have loaded the weights.
  4. Wait for the hot-load to finish.

Wait for the hot-load to finish

A hot-load is finished when the operation reaches a terminal state and every replica reports READY on the target identity. GetHotLoad reports the operation’s aggregate state, and ListHotLoadReplicas reports each replica separately.
  1. Poll the operation until it reaches a terminal state. This loop checks every 5 seconds and stops when the state is HOT_LOAD_STATE_COMPLETED, HOT_LOAD_STATE_FAILED, or HOT_LOAD_STATE_CANCELED, or after 15 minutes:
    If the final state is HOT_LOAD_STATE_FAILED or HOT_LOAD_STATE_CANCELED, see Recover a failed hot-load. If the operation is still HOT_LOAD_STATE_PENDING or HOT_LOAD_STATE_IN_PROGRESS when the loop stops, run the loop again.
  2. Confirm that every replica serves the target checkpoint. Replica results are paginated, so this loop reads every page, prints each replica, and counts the replicas listed and the replicas that aren’t ready on CW_TARGET_IDENTITY:
    When at least one replica is listed and the not-ready count is 0, every replica is serving the new checkpoint. A replica in ERROR state prints its errorReason in the last column.
For the full response bodies, see Example responses.

Load a delta checkpoint

Compute each delta checkpoint against a checkpoint that you already hot-loaded on the same deployment. This example loads step-000430 as a delta against step-000420. Your trainer writes the delta checkpoint. For an implementation that writes the format RL Rollouts accepts, see the delta checkpoint writer and Delta Weight Sync guide in Slime, an open source RL post-training framework. When you use that writer, set --update-weight-delta-encoding xor and --update-weight-delta-checksum adler32, because Slime also supports encodings and checksums that RL Rollouts doesn’t accept. CoreWeave offers recipes for RL Rollouts supporting Slime, Miles, NeMo RL, and other trainers on request. A delta checkpoint directory holds one or more .safetensors shards and a model.safetensors.index.json file:
  • Shards: Each changed tensor is stored as a byte-level XOR delta against the same tensor in the base checkpoint, compressed as its own Zstandard frame. The shard’s safetensors __metadata__ header maps each changed tensor’s name to the Adler-32 checksum of its new full weight, which the service uses to verify the reconstructed tensor.
  • Index: weight_map lists only the tensors that changed, and metadata describes the delta. The service identifies the delta and its base by the identity and previousSnapshotIdentity in your hot-load request, not by version and base_version, so those fields can hold other values, such as the zero-padded version numbers that Slime writes:
To load the delta, follow these steps:
  1. Upload the delta’s .safetensors shards and model.safetensors.index.json. Replace [LOCAL-DELTA-DIR] with the local directory your trainer wrote the delta to, and keep the files at the top level of that directory:
  2. Write the request to hot-load.json:
    previousSnapshotIdentity must name the checkpoint the delta was computed against, and that checkpoint must have a completed hot-load on this deployment. COMPRESSION_FORMAT_ZSTD and CHECKSUM_FORMAT_ADLER32 are the only formats the service accepts, and your delta artifacts must use them.
  3. Submit the request and save the new operation ID. If the ID prints as null, check hot-load-response.json for the error:
  4. To confirm that the delta loaded, run both loops in Wait for the hot-load to finish again.

Run strictly on-policy rounds

If your training algorithm requires every sample in a round to come from one policy, coordinate the boundary between rounds:
  1. Stop submitting new rollout requests and wait for the current round’s requests and multi-turn trajectories to finish.
  2. Publish the next checkpoint and start a hot-load, as in Load a full checkpoint or Load a delta checkpoint. To keep the next round from reusing KV cache computed under the previous policy, set promptCachePolicy to PROMPT_CACHE_POLICY_RESET_ALL.
  3. Wait for the hot-load to finish, including the replica check.
  4. Resume rollout generation only after every replica is ready on the target checkpoint.

Recover a failed hot-load

If an update stops before every replica loads the target checkpoint, the operation reaches HOT_LOAD_STATE_FAILED, and the replicas might serve different policy versions.
Keep on-policy rollout generation paused until every replica reports the target checkpoint. Replicas serving mixed checkpoints produce samples from more than one policy.
To recover, follow these steps:
  1. Run the replica loop in Wait for the hot-load to finish to see each replica’s loaded identity and errorReason.
  2. Address the reported issue, such as an inaccessible checkpoint or a missing base artifact.
  3. Start a new hot-load for the same checkpoint identity, or for another compatible checkpoint, such as a full checkpoint. A deployment runs one hot-load at a time, so the previous operation must be in a terminal state first, and the deployment must be in STATUS_READY. The same restriction blocks deployment configuration updates while a hot-load is running.
  4. Repeat both checks in the wait procedure before you resume on-policy generation.

Monitor hot-loads in the Cloud Console

The Console shows hot-load configuration and progress, and polls while an operation is running. It’s read-only: start hot-loads through the API. To track checkpoint adoption across replicas, follow these steps:
  1. Open your deployment in the Cloud Console.
  2. In the deployment details drawer, select the Hot loads tab.
  3. Review the hot-load configuration: whether hot-loading is enabled, the checkpoint URI, and the transition behavior.
  4. Review the operation. Current hot load covers pending and in-progress operations, and Latest hot load shows the newest completed, failed, or canceled one. Each reports the target checkpoint, snapshot type, status, and timestamps in your browser’s local time.
  5. Review Replica convergence for the count of ready replicas against the total, then the replica table for each replica’s name, state, served checkpoint, zone, transition time, and error reason.
When no operation is running, the tab reports that no checkpoint update is in progress. If Replica convergence shows replicas on different checkpoints after an operation ends, see Recover a failed hot-load.

View hot-load history

To list a deployment’s hot-loads, newest first:
The history response is paginated. To retrieve the next page, pass the response’s nextPageToken as pageToken, and keep the same parentDeploymentId filter. The final page omits nextPageToken.

Reference

The following sections list the hot-load API operations and states, with example responses.

Hot-load API endpoints

The management API at https://api.coreweave.com defines these hot-load operations:

Hot-load states

hotLoad.status.state reports the operation’s aggregate state:

Replica states

Each entry in a ListHotLoadReplicas response reports one replica’s state:

Example responses

These examples show a delta hot-load for step-000430 partway through the update. GetHotLoad returns the operation’s specification, the resolved checkpoint chain (snapshotChain), and the aggregate state. The chain lists the full baseline first, followed by each delta needed to reconstruct the target. It doesn’t include replica status:
ListHotLoadReplicas reports the checkpoint each replica currently serves, so a rolling update shows replicas in different states:

Next steps

Last modified on September 30, 2026