> ## Documentation Index
> Fetch the complete documentation index at: https://docs.coreweave.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Add or replace dataset entries

> Writes up to 5,000 OpenAI chat rows in one transaction. Send a JSON array, or the same rows as NDJSON (one row per line) with `Content-Type: application/x-ndjson`. Any invalid row rejects the request and is reported by its index; NDJSON line N is index N - 1. `mode=append` rejects a `row_id` that already exists (409 `entry_exists`); `mode=upsert` replaces it, keeping the stored `group_id`, `split`, `metadata`, and `provenance` for any of those the row omits. New entries default to `train` and their own group. A write that would place one group in more than one split returns 409 `group_split_conflict`. Writing moves the dataset to a new revision and marks completed relabels and evaluations stale.



## OpenAPI

````yaml openapi/model-distillation/management.openapi.yaml POST /projects/{alias}/datasets/{datasetId}/entries
openapi: 3.1.0
info:
  title: CoreWeave Model Distillation Management API
  version: 1.0.0
  description: |
    Manage providers, projects, routing versions, datasets, relabeling,
    fine-tunes, evaluations, and analytics. This is the same API used by
    Model Distillation.
servers:
  - url: https://forge.coreweave.com/api/distillation/v1
    description: Production
security:
  - WandbCredential: []
tags:
  - name: Providers
    description: OpenAI-compatible endpoints, credentials, model catalogs, and pricing.
  - name: Projects
    description: Stable project identity and project-level information.
  - name: Routing
    description: Versioned model targets, weights, and request parameters.
  - name: Datasets
    description: >-
      Reproducible data snapshots, entries, relabeling, and reusable model
      outputs.
  - name: Fine-tunes
    description: Supervised fine-tuning jobs and hosted model artifacts.
  - name: Evaluations
    description: Head-to-head, exact-match, and categorization evaluations.
  - name: Analytics
    description: Project traffic, token, error, and estimated-cost read models.
  - name: UI preferences
    description: Per-project display preferences used by the UI.
paths:
  /projects/{alias}/datasets/{datasetId}/entries:
    parameters:
      - $ref: '#/components/parameters/WandbEntity'
      - $ref: '#/components/parameters/ProjectAlias'
      - $ref: '#/components/parameters/DatasetId'
    post:
      tags:
        - Datasets
      summary: Add or replace dataset entries
      description: >-
        Writes up to 5,000 OpenAI chat rows in one transaction. Send a JSON
        array, or the same rows as NDJSON (one row per line) with `Content-Type:
        application/x-ndjson`. Any invalid row rejects the request and is
        reported by its index; NDJSON line N is index N - 1. `mode=append`
        rejects a `row_id` that already exists (409 `entry_exists`);
        `mode=upsert` replaces it, keeping the stored `group_id`, `split`,
        `metadata`, and `provenance` for any of those the row omits. New entries
        default to `train` and their own group. A write that would place one
        group in more than one split returns 409 `group_split_conflict`. Writing
        moves the dataset to a new revision and marks completed relabels and
        evaluations stale.
      operationId: writeDatasetEntries
      parameters:
        - name: mode
          in: query
          description: >-
            append rejects rows whose row_id already exists; upsert replaces
            them.
          required: false
          schema:
            default: append
            type: string
            enum:
              - append
              - upsert
        - name: dry_run
          in: query
          description: Validate and report the change, then roll it back.
          required: false
          schema:
            type: boolean
            default: false
      requestBody:
        required: true
        content:
          application/json:
            schema:
              minItems: 1
              maxItems: 5000
              type: array
              items:
                type: object
                properties:
                  row_id:
                    description: >-
                      Stable identity within the dataset. Upserts match on it.
                      Defaults to a digest of the messages and request fields.
                    type: string
                    minLength: 1
                    maxLength: 512
                  group_id:
                    description: >-
                      Rows that must share a split, such as the turns of one
                      conversation. Defaults to row_id.
                    type: string
                    minLength: 1
                    maxLength: 512
                  split:
                    description: >-
                      train, or val to hold the row out of training for
                      evaluation. New rows default to train.
                    type: string
                    enum:
                      - train
                      - val
                  messages:
                    minItems: 2
                    maxItems: 1000
                    type: array
                    items:
                      type: object
                      propertyNames:
                        type: string
                      additionalProperties: {}
                  tools:
                    maxItems: 128
                    type: array
                    items:
                      type: object
                      properties:
                        type:
                          type: string
                          const: function
                        function:
                          type: object
                          properties:
                            name:
                              type: string
                              minLength: 1
                              maxLength: 256
                            description:
                              type: string
                              maxLength: 16384
                            parameters:
                              type: object
                              propertyNames:
                                type: string
                              additionalProperties: {}
                            strict:
                              type: boolean
                          required:
                            - name
                          additionalProperties: false
                      required:
                        - type
                        - function
                      additionalProperties: false
                  tool_choice: {}
                  response_format: {}
                  metadata:
                    description: >-
                      Filterable attributes, such as the user or customer
                      segment.
                    type: object
                    propertyNames:
                      type: string
                    additionalProperties: {}
                  provenance:
                    description: >-
                      Where the row came from, such as the source system and
                      trace id.
                    type: object
                    propertyNames:
                      type: string
                    additionalProperties: {}
                required:
                  - messages
                additionalProperties: false
          application/x-ndjson:
            schema:
              type: string
              description: >-
                One JSON row per line, in the same shape as the JSON array's
                items.
      responses:
        '200':
          description: Entries were written, or would be for a dry run.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/DatasetEntryWriteResult'
        default:
          $ref: '#/components/responses/Error'
components:
  parameters:
    WandbEntity:
      name: Wandb-Entity
      in: header
      required: false
      description: >-
        Accessible personal or team entity. Omit to use the authenticated user's
        default entity.
      schema:
        type: string
    ProjectAlias:
      name: alias
      in: path
      required: true
      schema:
        type: string
        pattern: ^[a-z0-9-]{1,64}$
    DatasetId:
      name: datasetId
      in: path
      required: true
      schema:
        type: string
        format: uuid
  schemas:
    DatasetEntryWriteResult:
      type: object
      required:
        - dry_run
        - created
        - updated
        - revision
        - entry_counts
        - group_conflicts
      properties:
        dry_run:
          type: boolean
        created:
          type: integer
          description: Rows that became new entries.
        updated:
          type: integer
          description: Existing entries the rows replaced.
        revision:
          type: integer
          description: The dataset revision after the write; unchanged by a dry run.
        entry_counts:
          $ref: '#/components/schemas/DatasetEntryCounts'
        group_conflicts:
          type: array
          items:
            type: string
          description: >-
            Groups spanning more than one split. A real write with conflicts is
            rejected.
    DatasetEntryCounts:
      type: object
      required:
        - total
        - train
        - val
      properties:
        total:
          type: integer
        train:
          type: integer
        val:
          type: integer
    ErrorResponse:
      type: object
      required:
        - error
      properties:
        error:
          type: object
          required:
            - message
            - type
          properties:
            message:
              type: string
            type:
              type: string
      example:
        error:
          message: Project 'missing' not found in entity 'your-team'
          type: not_found
  responses:
    Error:
      description: Request failed.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/ErrorResponse'
  securitySchemes:
    WandbCredential:
      type: http
      scheme: bearer
      bearerFormat: wandb API key or JWT
      description: >-
        A wandb API key or wandb JWT. Browser sessions are exchanged by the
        Model Distillation web app backend; job capabilities are internal only.

````