Prerequisites
Before you begin, make sure you have the following:- A W&B API key for a team that has access to Model Distillation. The examples read it from the
WANDB_API_KEYenvironment variable, pass it in theAuthorizationheader, and name the entity in theWandb-Entityheader. - A project in that team. The project alias appears in every request path. To create one, see the Quick Start.
- The tools for the tab you plan to use in Upload a file. You need only one of the following sets:
curltogether withjq.- Python 3.9 or later with the
requestspackage. - Node.js 18 or later, with no extra packages.
How it works
Every upload goes into an existing dataset. You create an empty dataset once, then add rows to it in either of two ways:- Write entries sends up to 5,000 rows in one request and returns the result immediately. Use it for small datasets and for incremental additions.
- Upload a file sends one JSONL file of up to 1 GiB and 100,000 rows straight to object storage. Model Distillation validates every line and then commits all rows at once.
split on each row, or move a share of the rows afterward. See Choose the validation rows.
Prepare the rows
Each row is an OpenAI Chat Completions request whose last message is the assistant response to train toward. Model Distillation stores the preceding messages, plus anytools, tool_choice, and response_format, as the input, and the final assistant message as the output.
system, developer, user, assistant, or tool, and the last message must be from the assistant. tools accepts up to 128 function tools and 1 MiB of JSON.
Optional fields
Rows can carry the following optional fields:
Rows must not contain other top-level fields.
Append or upsert
Every write uses one of two modes:append, the default, rejects a row whoserow_idalready exists in the dataset.upsertreplaces the existing entry. If the new row omitsgroup_id,split,metadata, orprovenance, the stored value is kept. Relabeled outputs for a replaced entry are discarded.
row_id that appears twice in one request or file is rejected, and a write that would place one group in both splits fails with group_split_conflict.
Create an empty dataset
Create the dataset by sending only a name. Withoutparams, the request creates an empty dataset that is ready immediately, and the response is 201 Created:
id from the response. The following examples refer to it as DATASET_ID.
Write entries
To add up to 5,000 rows in one request, send them to the entries endpoint as a JSON array. You can also send the same rows as NDJSON, one row per line, withContent-Type: application/x-ndjson:
200 OK with created, updated, the new revision, and entry_counts by split. If any row is invalid, the request is rejected and the error names the row by its index. In NDJSON, line N is index N - 1.
To check a batch without changing the dataset, add dry_run=true. Model Distillation validates the rows, reports what would change, and rolls the write back.
Upload a file
For files larger than one request allows, upload the file in four stages: open an upload, send the file with onePUT, complete the upload, and poll until the rows are committed. The following tabs show the whole flow as one script. Numbered comments mark the stages, and How the script works explains each one.
The scripts use a project alias of ticket-classifier and a file named tickets.jsonl, and read the dataset ID from the DATASET_ID environment variable. Before you run a script, replace the entity, alias, and filename with your own, and set WANDB_API_KEY and DATASET_ID in your environment. Use a new idempotency_key for each distinct upload.
- curl
- Python
- JavaScript
How the script works
The numbered comments in each script correspond to the following stages.1
Open the upload
The script sends the filename, its exact byte size, and the write mode. The request also requires an
idempotency_key, so a retried request returns the existing upload instead of opening a second one. Reusing a key with a different request body returns 409 Conflict with type idempotency_conflict.The response is 201 Created with the upload. upload.url is a signed URL for the file, and upload.expires_at is when it stops working. The following excerpt shows those fields:2
Send the file
The script sends the whole file in one
PUT request to upload.url. The URL is bound to the declared size, so the body must be exactly file.size_bytes long. It carries its own authorization, so the script doesn’t add the Authorization or Wandb-Entity headers to this request.The URL must be used within 15 minutes. If it expires, read the upload again: while the upload is in uploading, each read returns a fresh URL.3
Complete the upload
When the file is stored, the script completes the upload. Model Distillation checks that the stored file matches the declared size and queues validation. The response is
202 Accepted.If the file hasn’t arrived, the response is 409 Conflict with type upload_incomplete. If its size doesn’t match, the type is upload_size_mismatch, and the stored file is discarded. In both cases, the upload stays in uploading, so you can send the file again and complete the upload again. Completing an upload that’s already complete returns its current state without error.4
Poll until the rows are committed
The script reads the upload every 10 seconds until it reaches a final state. While validation runs,
validation.validated_lines and validation.validated_bytes report progress.When state is committed, counters reports rows, created, updated, the new revision, and entry_counts by split. If state is failed, see Read validation results.If relabeling, an evaluation, or a fine-tune is running on the dataset, the commit waits until they finish rather than failing.Upload states
An upload moves through the following states:
An upload that is still
uploading 24 hours after creation can no longer be completed. Completing it returns 410 Gone with type dataset_upload_expired, and an hourly job moves it to expired. When an upload finishes in any state, its stored file is deleted.
Limits
Files and uploads must stay within the following limits:
An entity is the team or personal account named in the
Wandb-Entity header. An upload counts toward the entity’s limits while it is uploading, queued, or validating. A request that would exceed either limit returns 429 Too Many Requests. To free capacity, cancel an upload you no longer need, or wait for one to finish.
Read validation results
Afailed upload reports error as a short summary, and validation carries the details:
error_countis the total number of problems found.errorslists up to 100 problems, ordered by line, each withline,code, andmessage.lineisnullfor a problem with the whole upload, such as a commit conflict.errors_truncatedistruewhen more problems exist than the list shows.
invalid_json, invalid_utf8, duplicate_json_key, json_too_deep, row_too_large, empty_line, duplicate_row_id, row_limit_exceeded, and validation_failed, which covers rows that don’t match the row format. Upload-level codes are entry_exists, when an append upload contains a row_id that’s already in the dataset, and group_split_conflict.
A failed upload can’t be reopened. Fix the file and open a new upload.
Cancel an upload
To discard an upload that hasn’t committed, delete it. The response is204 No Content, the state becomes cancelled, and the stored file is deleted. Repeating the request on a cancelled upload returns 204 No Content again.
409 Conflict with type dataset_upload_committed. To remove its rows, delete the dataset instead. See Delete a dataset.
Choose the validation rows
Rows without asplit value go to training. To hold a share of the dataset out for evaluation after you upload it, move entries to the validation split:
filters, which take the same filters as the entry listing, including metadata, and with row_ids. Without either, it selects every entry. fraction moves only that share of the selected groups. Groups are ranked by a digest of the group ID, so a group is never divided, and a later request can take its share from what remains.
To hold out a whole population, such as one customer, filter on metadata instead. Evaluations then measure the model on users it never saw in training:
revision, and entry_counts. Add dry_run=true to see the result without changing the dataset.
Errors
Entry and upload requests can return the following errors.
For request and response schemas, see the following pages in the Management API reference:
- Create a dataset
- Add or replace dataset entries
- Move dataset entries to a split
- Start a JSONL upload into a dataset
- Get a dataset upload
- Finish a dataset upload
- Cancel a dataset upload
Next steps
Relabeling
Ask a stronger model to rewrite the assistant responses in the uploaded dataset without changing the original rows.
Fine-tuning
Train a supported base model on the uploaded dataset, using the original outputs or a relabeled output set.