> ## Documentation Index
> Fetch the complete documentation index at: https://docs.coreweave.com/llms.txt
> Use this file to discover all available pages before exploring further.

# 평가 데이터 내보내기

> Evaluation REST API를 사용하여 평가 결과를 프로그래밍 방식으로 내보냅니다.

W\&B Weave에서 평가를 실행하는 팀은 Weave UI 외부에서 평가 결과가 필요한 경우가 많습니다. 일반적인 사용 사례는 다음과 같습니다:

* 맞춤형 분석 및 시각화를 위해 스프레드시트나 노트북으로 메트릭을 가져옵니다.
* 배포를 제어하기 위해 CI/CD 파이프라인에 평가 결과를 공급합니다.
* W\&B 시트가 없는 이해관계자와 Looker와 같은 BI 도구나 내부 대시보드를 통해 결과를 공유합니다.
* 프로젝트 전반의 점수를 집계하는 자동화된 보고 파이프라인을 구축합니다.

[v2 Evaluation REST API](https://trace.wandb.ai/docs)는 평가 run, 예측, 점수, Scorer와 같은 집중된 평가 개념을 제공합니다. 그 결과는 범용 Calls API에 비해 유형 지정 Scorer statistics 및 resolved dataset inputs가 포함된 더 풍부하고 구조화된 출력을 제공합니다.

<h2 id="api-endpoints-used">
  API endpoints used
</h2>

이 페이지의 스니펫은 [v2 Evaluation REST API](https://trace.wandb.ai/docs)의 다음 엔드포인트를 사용합니다:

* `GET /v2/{entity}/{project}/evaluation_runs`: 프로젝트의 evaluation run을 목록으로 표시하며, evaluation reference, model reference 또는 run ID로 선택 필터링할 수 있습니다.
* `GET /v2/{entity}/{project}/evaluation_runs/{evaluation_run_id}`: 단일 evaluation run을 조회하여 해당 모델, evaluation reference, status, timestamps 및 summary를 가져옵니다.
* `POST /v2/{entity}/{project}/eval_results/query`: 하나 이상의 evaluations에 대한 그룹화된 evaluation result 행을 조회합니다. model output, 점수 및 선택적으로 해결된 dataset row 입력이 포함된 행별 trial을 반환합니다. 요청 시 aggregated scorer statistics도 반환합니다.
* `GET /v2/{entity}/{project}/predictions/{prediction_id}`: 입력, 출력 및 model reference가 포함된 개별 prediction을 조회합니다.

인증은 HTTP Basic을 사용하며, 사용자 이름은 `api`, 비밀번호는 W\&B API 키입니다.

<h2 id="prerequisites">
  사전 요구 사항
</h2>

이 페이지의 예시는 Python을 사용하지만, Evaluation REST API는 언어에 구애받지 않습니다. TypeScript나 다른 HTTP 클라이언트에서 동일한 엔드포인트를 호출할 수 있습니다.

시작하기 전에 다음을 준비하세요:

* Python 3.7 이상.
* `requests` 라이브러리. `pip install requests`로 설치하세요.
* `WANDB_API_KEY` 환경 변수로 설정된 API 키. [CoreWeave Forge UI](https://forge.coreweave.com/settings#apikeys)에서 키를 가져오세요.

<h2 id="set-up-authentication">
  인증 설정
</h2>

다음 스니펫은 이 페이지 전반에 걸쳐 사용되는 라이브러리를 임포트하고 base URL, authentication 튜플, 대상 entity와 프로젝트를 설정합니다. 이후의 모든 예시는 이 변수를 재사용합니다.

```python theme={"system"}
import json
import os

import requests

TRACE_BASE = "https://trace.wandb.ai"
AUTH = ("api", os.environ["WANDB_API_KEY"])

entity = "my-team"
project = "my-project"
```

인증이 설정되면 다음 섹션에 설명된 엔드포인트를 호출할 수 있습니다.

<h2 id="list-evaluation-runs">
  evaluation run 목록
</h2>

evaluation run의 목록은 보통 export 워크플로에서 가장 먼저 필요한 것입니다. 이 목록은 다른 endpoint에서 요구하는 `evaluation_run_id` 값을 제공하기 때문입니다. 프로젝트에서 최근 evaluation run을 가져오고 각 run의 ID와 status 같은 세부 정보를 나열하세요.

```python theme={"system"}
resp = requests.get(
    f"{TRACE_BASE}/v2/{entity}/{project}/evaluation_runs",
    auth=AUTH,
)
runs = [json.loads(line) for line in resp.text.strip().splitlines()]

for run in runs:
    print(run["evaluation_run_id"], run.get("status"))
```

<h2 id="read-a-single-evaluation-run">
  단일 evaluation run 조회
</h2>

`evaluation_run_id`를 가지고 있으면 해당 run의 전체 record를 가져올 수 있습니다. 특정 evaluation run의 상세 정보를 조회하며, 여기에는 모델, evaluation 레퍼런스, status, timestamps가 포함됩니다. 가져오려는 evaluation run의 ID로 `[EVALUATION_RUN_ID]`를 교체하세요.

```python theme={"system"}
eval_run_id = "[EVALUATION_RUN_ID]"

resp = requests.get(
    f"{TRACE_BASE}/v2/{entity}/{project}/evaluation_runs/{eval_run_id}",
    auth=AUTH,
)
eval_run = resp.json()
print(eval_run["evaluation_run_id"], eval_run.get("status"), eval_run.get("model"))
```

<h2 id="get-predictions-and-scores">
  예측 및 점수 조회
</h2>

run의 기본 데이터가 필요한 경우(예: 스프레드시트 내보내기 또는 행 수준 분석), `eval_results/query` 엔드포인트를 사용하여 evaluation run의 행별 결과를 조회하세요. 각 행에는 resolved dataset inputs, 모델 출력 및 개별 Scorer 결과가 포함됩니다. `include_rows`, `include_raw_data_rows`, `resolve_row_refs`를 설정하여 전체 행별 detail을 가져오세요. 조회하려는 evaluation run의 ID로 `[EVALUATION_RUN_ID]`를 교체하세요.

```python theme={"system"}
eval_run_id = "[EVALUATION_RUN_ID]"

resp = requests.post(
    f"{TRACE_BASE}/v2/{entity}/{project}/eval_results/query",
    json={
        "evaluation_run_ids": [eval_run_id],
        "include_rows": True,
        "include_raw_data_rows": True,
        "resolve_row_refs": True,
    },
    auth=AUTH,
)
results = resp.json()

for row in results["rows"]:
    inputs = row.get("raw_data_row")
    for ev in row.get("evaluations", []):
        for trial in ev.get("trials", []):
            output = trial.get("model_output")
            scores = trial.get("scores", {})
            print("Input:", inputs)
            print("Output:", output)
            print("Scores:", scores)
```

<h2 id="get-aggregated-scores">
  집계된 점수 조회
</h2>

고수준 메트릭만 필요한 경우(예: 대시보드 또는 CI/CD 게이팅), 행별 데이터 대신 요약 통계를 요청하세요. 동일한 `eval_results/query` 엔드포인트는 행별 데이터 대신 집계된 Scorer 통계도 반환할 수 있습니다. `include_summary`를 설정하여 이진 Scorer의 통과율이나 연속 Scorer의 평균과 같은 요약 수준 메트릭을 가져오세요.

```python theme={"system"}
resp = requests.post(
    f"{TRACE_BASE}/v2/{entity}/{project}/eval_results/query",
    json={
        "evaluation_run_ids": [eval_run_id],
        "include_summary": True,
        "include_rows": False,
    },
    auth=AUTH,
)
results = resp.json()

for ev in results["summary"]["evaluations"]:
    for stat in ev["scorer_stats"]:
        print(stat["scorer_key"], stat.get("value_type"), stat.get("pass_rate") or stat.get("numeric_mean"))
```

<h2 id="read-a-single-prediction">
  단일 prediction 조회
</h2>

단일 행을 격리하여 검사하려면, 예를 들어 예상치 못한 score를 조사할 때, ID로 직접 prediction을 가져올 수 있습니다. 개별 prediction의 전체 세부 정보를 가져오며, 여기에는 inputs, 출력 및 model 레퍼런스가 포함됩니다. 검색하려는 prediction의 ID로 `[PREDICTION_ID]`를 교체하세요.

```python theme={"system"}
prediction_id = "[PREDICTION_ID]"

resp = requests.get(
    f"{TRACE_BASE}/v2/{entity}/{project}/predictions/{prediction_id}",
    auth=AUTH,
)
prediction = resp.json()
print(prediction)
```

<h2 id="row-digests">
  Row digests
</h2>

각 엔드포인트가 반환하는 원시 데이터 외에도, `eval_results/query`의 응답에는 run 간에 행을 서로 연관 짓는 데 도움이 되는 추가 식별자가 포함됩니다. `eval_results/query` 엔드포인트의 각 결과 행에는 `row_digest`가 포함됩니다. 이는 평가 데이터셋의 특정 입력을 위치가 아닌 내용을 기준으로 고유하게 식별하는 콘텐츠 해시입니다. Row digest는 다음과 같은 경우에 유용합니다.

* **평가 간 비교**: 동일한 데이터셋으로 서로 다른 두 모델을 실행하면 digest가 같은 행은 동일한 입력을 나타냅니다. `row_digest`를 기준으로 조인하면 정확히 같은 작업에서 각 모델의 성능을 비교할 수 있습니다.
* **중복 제거**: 동일한 작업이 여러 평가 스위트에 포함된 경우 digest로 이를 파악할 수 있습니다.
* **재현성**: digest는 콘텐츠 주소 기반이므로 누군가 데이터셋 행을 수정하면(지시문 텍스트, 루브릭 또는 기타 필드 변경) 새 digest가 부여됩니다. 이를 통해 두 evaluation run이 동일한 입력을 사용했는지, 아니면 서로 다른 버전을 사용했는지 확인할 수 있습니다.
