> ## Documentation Index
> Fetch the complete documentation index at: https://docs.coreweave.com/llms.txt
> Use this file to discover all available pages before exploring further.

# 리더보드 퀵스타트

> W&B Weave로 리더보드 퀵스타트를 사용하는 방법을 알아보세요

<Note>
  이것은 대화형 노트북입니다. 로컬에서 실행하거나 아래 링크를 사용할 수 있습니다:

  * [Open in Google Colab](https://colab.research.google.com/github/wandb/docs/blob/main/weave/cookbooks/source/leaderboard_quickstart.ipynb)
  * [View source on GitHub](https://github.com/wandb/docs/blob/main/weave/cookbooks/source/leaderboard_quickstart.ipynb)
</Note>

<h1 id="leaderboard-quickstart">
  리더보드 퀵스타트
</h1>

이 퀵스타트에서는 W\&B Weave 리더보드를 사용하여 여러 데이터셋과 점수화 함수에 걸쳐 모델 성능을 비교하는 방법을 안내합니다. 이 가이드를 마치면 공통 평가 집합을 기준으로 여러 모델의 순위를 매기는 리더보드를 게시하게 됩니다. 그런 다음 각 메트릭에서 가장 뛰어난 성능을 보이는 모델을 파악할 수 있습니다. 이 가이드는 Weave 평가 실행에 익숙하고 결과를 나란히 비교하려는 개발자를 대상으로 합니다.

구체적으로 다음 작업을 수행합니다:

1. 가상의 우편번호 데이터로 데이터셋을 생성합니다.
2. 점수화 함수를 작성하고 기준선 모델을 평가합니다.
3. 이러한 기법을 사용하여 여러 모델과 평가 구성의 조합을 평가합니다.
4. Weights & Biases UI에서 리더보드를 검토합니다.

<h2 id="step-1-generate-a-dataset-of-fake-zip-code-data">
  1단계: 가짜 우편번호 데이터의 데이터셋 생성
</h2>

먼저 `generate_dataset_rows` 함수를 만들어 가짜 우편번호 데이터의 목록을 생성하세요. 이 합성 데이터셋은 리더보드에 일관된 입력 세트와 각 모델을 평가하기 위한 기대값을 제공합니다.

```python lines theme={"system"}
import json

from openai import OpenAI
from pydantic import BaseModel

class Row(BaseModel):
    zip_code: str
    city: str
    state: str
    avg_temp_f: float
    population: int
    median_income: int
    known_for: str

class Rows(BaseModel):
    rows: list[Row]

def generate_dataset_rows(
    location: str = "United States", count: int = 5, year: int = 2022
):
    client = OpenAI()

    completion = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[
            {"role": "system", "content": "You are a helpful assistant."},
            {
                "role": "user",
                "content": f"Please generate {count} rows of data for random zip codes in {location} for the year {year}.",
            },
        ],
        response_format={
            "type": "json_schema",
            "json_schema": {
                "name": "response_format",
                "schema": Rows.model_json_schema(),
            },
        },
    )

    return json.loads(completion.choices[0].message.content)["rows"]
python
import weave

weave.init("leaderboard-demo")
```

<h2 id="step-2-author-scoring-functions">
  2단계: 점수화 함수 작성
</h2>

다음으로 점수화 함수 세 개를 작성하세요. 각 Scorer는 모델 출력의 서로 다른 측면을 평가하므로 리더보드에서 품질의 각 차원별로 모델의 순위를 매길 수 있습니다:

1. `check_concrete_fields`: 모델 출력이 예상 도시 및 주와 일치하는지 확인합니다.
2. `check_value_fields`: 모델 출력이 예상 인구 및 중위 소득의 10% 이내인지 확인합니다.
3. `check_subjective_fields`: LLM을 사용하여 모델 출력이 예상 "known for" 필드와 일치하는지 확인합니다.

```python lines theme={"system"}
@weave.op
def check_concrete_fields(city: str, state: str, output: dict):
    return {
        "city_match": city == output["city"],
        "state_match": state == output["state"],
    }

@weave.op
def check_value_fields(
    avg_temp_f: float, population: int, median_income: int, output: dict
):
    return {
        "avg_temp_f_err": abs(avg_temp_f - output["avg_temp_f"]) / avg_temp_f,
        "population_err": abs(population - output["population"]) / population,
        "median_income_err": abs(median_income - output["median_income"])
        / median_income,
    }

@weave.op
def check_subjective_fields(zip_code: str, known_for: str, output: dict):
    client = OpenAI()

    class Response(BaseModel):
        correct_known_for: bool

    completion = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[
            {"role": "system", "content": "You are a helpful assistant."},
            {
                "role": "user",
                "content": f"My student was asked what the zip code {zip_code} is best known best for. The right answer is '{known_for}', and they said '{output['known_for']}'. Is their answer correct?",
            },
        ],
        response_format={
            "type": "json_schema",
            "json_schema": {
                "name": "response_format",
                "schema": Response.model_json_schema(),
            },
        },
    )

    return json.loads(completion.choices[0].message.content)
```

<h2 id="step-3-create-an-evaluation">
  3단계: 평가 생성
</h2>

다음으로, 가상 데이터와 점수화 함수를 사용하여 평가를 정의하세요. `Evaluation` 객체는 데이터셋과 Scorer를 연결하므로, 동일한 벤치마크로 모든 모델을 평가할 수 있습니다.

```python lines theme={"system"}
rows = generate_dataset_rows()
evaluation = weave.Evaluation(
    name="United States - 2022",
    dataset=rows,
    scorers=[
        check_concrete_fields,
        check_value_fields,
        check_subjective_fields,
    ],
)
```

<h2 id="step-4-evaluate-a-baseline-model">
  4단계: 기준선 모델 평가
</h2>

이제 정적 응답을 반환하는 기준선 모델을 평가하세요. 베이스라인을 설정하면 리더보드에서 비교 기준이 마련되어 이후 각 모델이 정적 구현보다 얼마나 개선되었는지 측정할 수 있습니다.

```python lines theme={"system"}
@weave.op
def baseline_model(zip_code: str):
    return {
        "city": "New York",
        "state": "NY",
        "avg_temp_f": 50.0,
        "population": 1000000,
        "median_income": 100000,
        "known_for": "The Big Apple",
    }

await evaluation.evaluate(baseline_model)
```

<h2 id="step-5-create-more-models">
  5단계: 모델 추가 생성
</h2>

이제 베이스라인과 비교할 모델 두 개를 더 만드세요. 한 모델에는 추가 프롬프트 없이 우편번호만 제공하고, 다른 모델에는 구조화된 프롬프트를 제공합니다. 리더보드에서 두 모델을 비교하면 프롬프트 맥락이 답변 품질에 미치는 영향을 확인할 수 있습니다.

```python lines theme={"system"}
@weave.op
def gpt_4o_mini_no_context(zip_code: str):
    client = OpenAI()

    completion = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": f"""Zip code {zip_code}"""}],
        response_format={
            "type": "json_schema",
            "json_schema": {
                "name": "response_format",
                "schema": Row.model_json_schema(),
            },
        },
    )

    return json.loads(completion.choices[0].message.content)

await evaluation.evaluate(gpt_4o_mini_no_context)
python
@weave.op
def gpt_4o_mini_with_context(zip_code: str):
    client = OpenAI()

    completion = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[
            {
                "role": "user",
                "content": f"""Please answer the following questions about the zip code {zip_code}:
                   1. What is the city?
                   2. What is the state?
                   3. What is the average temperature in Fahrenheit?
                   4. What is the population?
                   5. What is the median income?
                   6. What is the most well known thing about this zip code?
                   """,
            }
        ],
        response_format={
            "type": "json_schema",
            "json_schema": {
                "name": "response_format",
                "schema": Row.model_json_schema(),
            },
        },
    )

    return json.loads(completion.choices[0].message.content)

await evaluation.evaluate(gpt_4o_mini_with_context)
```

<h2 id="step-6-create-more-evaluations">
  6단계: 평가 추가 생성
</h2>

이제 모델과 평가의 조합으로 구성된 행렬을 평가하세요. 모든 모델을 여러 데이터셋(서로 다른 지역과 연도)에 대해 실행하면 리더보드가 여러 조건에서 모델의 순위를 매기는 데 필요한 데이터가 생성됩니다.

```python lines theme={"system"}
scorers = [
    check_concrete_fields,
    check_value_fields,
    check_subjective_fields,
]
evaluations = [
    weave.Evaluation(
        name="United States - 2022",
        dataset=weave.Dataset(
            name="United States - 2022",
            rows=generate_dataset_rows("United States", 5, 2022),
        ),
        scorers=scorers,
    ),
    weave.Evaluation(
        name="California - 2022",
        dataset=weave.Dataset(
            name="California - 2022", rows=generate_dataset_rows("California", 5, 2022)
        ),
        scorers=scorers,
    ),
    weave.Evaluation(
        name="United States - 2000",
        dataset=weave.Dataset(
            name="United States - 2000",
            rows=generate_dataset_rows("United States", 5, 2000),
        ),
        scorers=scorers,
    ),
]
models = [
    baseline_model,
    gpt_4o_mini_no_context,
    gpt_4o_mini_with_context,
]

for evaluation in evaluations:
    for model in models:
        await evaluation.evaluate(
            model, __weave={"display_name": evaluation.name + ":" + model.__name__}
        )
```

<h2 id="step-7-review-the-leaderboard">
  7단계: 리더보드 검토
</h2>

평가 결과를 게시했으므로, 이제 리더보드에 모아 나란히 비교할 수 있습니다.

UI의 리더보드 탭으로 이동하여 **Create Leaderboard**를 클릭하면 새 리더보드를 만들 수 있습니다.

Python에서 직접 리더보드를 생성할 수도 있습니다:

```python lines theme={"system"}
from weave.flow import leaderboard
from weave.trace.ref_util import get_ref

spec = leaderboard.Leaderboard(
    name="Zip Code World Knowledge",
    description="""
This leaderboard compares the performance of models in terms of world knowledge about zip codes.

### Columns

1. **State Match against `United States - 2022`**: The fraction of zip codes that the model correctly identified the state for.
2. **Avg Temp F Error against `California - 2022`**: The mean absolute error of the model's average temperature prediction.
3. **Correct Known For against `United States - 2000`**: The fraction of zip codes that the model correctly identified the most well known thing about the zip code.
""",
    columns=[
        leaderboard.LeaderboardColumn(
            evaluation_object_ref=get_ref(evaluations[0]).uri(),
            scorer_name="check_concrete_fields",
            summary_metric_path="state_match.true_fraction",
        ),
        leaderboard.LeaderboardColumn(
            evaluation_object_ref=get_ref(evaluations[1]).uri(),
            scorer_name="check_value_fields",
            should_minimize=True,
            summary_metric_path="avg_temp_f_err.mean",
        ),
        leaderboard.LeaderboardColumn(
            evaluation_object_ref=get_ref(evaluations[2]).uri(),
            scorer_name="check_subjective_fields",
            summary_metric_path="correct_known_for.true_fraction",
        ),
    ],
)

ref = weave.publish(spec)
```

이제 Weave에 게시된 리더보드가 있어, 정의한 세 가지 평가와 점수화 메트릭에 따라 각 모델을 순위 매깁니다. Weights & Biases UI에서 모델별 점수를 확인하고, 개별 평가 run을 자세히 살펴보며, 향후 모델 iteration을 동일한 베이스라인과 비교할 수 있습니다.
