> ## Documentation Index
> Fetch the complete documentation index at: https://docs.coreweave.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Leaderboard クイックスタート

> W&B Weave で Leaderboard クイックスタートの使い方を学びます

<Note>
  これはインタラクティブなノートブックです。ローカルで実行するか、以下のリンクから利用できます。

  * [Google Colab で開く](https://colab.research.google.com/github/wandb/docs/blob/main/weave/cookbooks/source/leaderboard_quickstart.ipynb)
  * [GitHub でソースを表示](https://github.com/wandb/docs/blob/main/weave/cookbooks/source/leaderboard_quickstart.ipynb)
</Note>

<h1 id="leaderboard-quickstart">
  Leaderboard クイックスタート
</h1>

このクイックスタートでは、W\&B Weave の Leaderboard を使用して、複数のデータセットとスコアリング関数にわたってモデル性能を比較する方法を説明します。最終的には、共通の評価セットで複数のモデルをランク付けする Leaderboard をパブリッシュします。これにより、各メトリクスで最も優れたパフォーマンスを発揮するモデルを特定できます。このガイドは、Weave での評価の実行に慣れていて、結果を並べて比較したい開発者を対象としています。

具体的には、次の作業を行います。

1. 架空の郵便番号データを含むデータセットを生成します。
2. いくつかのスコアリング関数を作成し、ベースラインモデルを評価します。
3. これらの手法を使用して、モデルと評価のマトリクスを評価します。
4. Weights & Biases の UI で Leaderboard を確認します。

<h2 id="step-1-generate-a-dataset-of-fake-zip-code-data">
  ステップ 1: 架空の郵便番号データのデータセットを生成する
</h2>

まず、架空の郵便番号データのリストを生成する関数 `generate_dataset_rows` を作成します。この合成データセットによって、Leaderboard で各モデルをスコアリングする際に、常に同じ入力と期待値を基準として使えるようになります。

```python lines theme={"system"}
import json

from openai import OpenAI
from pydantic import BaseModel

class Row(BaseModel):
    zip_code: str
    city: str
    state: str
    avg_temp_f: float
    population: int
    median_income: int
    known_for: str

class Rows(BaseModel):
    rows: list[Row]

def generate_dataset_rows(
    location: str = "United States", count: int = 5, year: int = 2022
):
    client = OpenAI()

    completion = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[
            {"role": "system", "content": "You are a helpful assistant."},
            {
                "role": "user",
                "content": f"Please generate {count} rows of data for random zip codes in {location} for the year {year}.",
            },
        ],
        response_format={
            "type": "json_schema",
            "json_schema": {
                "name": "response_format",
                "schema": Rows.model_json_schema(),
            },
        },
    )

    return json.loads(completion.choices[0].message.content)["rows"]
python
import weave

weave.init("leaderboard-demo")
```

<h2 id="step-2-author-scoring-functions">
  ステップ 2: スコアリング関数を作成する
</h2>

次に、3 つのスコアリング関数を作成します。各 Scorer はモデル出力のそれぞれ異なる側面を評価するため、Leaderboard では品質の異なる観点ごとにモデルをランク付けできます。

1. `check_concrete_fields`: モデル出力が期待される市および州と一致するかどうかを確認します。
2. `check_value_fields`: モデル出力が、期待される人口および収入の中央値から 10% 以内に収まっているかどうかを確認します。
3. `check_subjective_fields`: LLM を使用して、モデル出力が期待される "known for" フィールドと一致するかどうかを確認します。

```python lines theme={"system"}
@weave.op
def check_concrete_fields(city: str, state: str, output: dict):
    return {
        "city_match": city == output["city"],
        "state_match": state == output["state"],
    }

@weave.op
def check_value_fields(
    avg_temp_f: float, population: int, median_income: int, output: dict
):
    return {
        "avg_temp_f_err": abs(avg_temp_f - output["avg_temp_f"]) / avg_temp_f,
        "population_err": abs(population - output["population"]) / population,
        "median_income_err": abs(median_income - output["median_income"])
        / median_income,
    }

@weave.op
def check_subjective_fields(zip_code: str, known_for: str, output: dict):
    client = OpenAI()

    class Response(BaseModel):
        correct_known_for: bool

    completion = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[
            {"role": "system", "content": "You are a helpful assistant."},
            {
                "role": "user",
                "content": f"My student was asked what the zip code {zip_code} is best known best for. The right answer is '{known_for}', and they said '{output['known_for']}'. Is their answer correct?",
            },
        ],
        response_format={
            "type": "json_schema",
            "json_schema": {
                "name": "response_format",
                "schema": Response.model_json_schema(),
            },
        },
    )

    return json.loads(completion.choices[0].message.content)
```

<h2 id="step-3-create-an-evaluation">
  ステップ 3: 評価を作成する
</h2>

次に、疑似データとスコアリング関数を使って評価を定義します。`Evaluation` オブジェクトはデータセットと Scorer を組み合わせたものなので、どのモデルでも同じベンチマークで実行できます。

```python lines theme={"system"}
rows = generate_dataset_rows()
evaluation = weave.Evaluation(
    name="United States - 2022",
    dataset=rows,
    scorers=[
        check_concrete_fields,
        check_value_fields,
        check_subjective_fields,
    ],
)
```

<h2 id="step-4-evaluate-a-baseline-model">
  ステップ 4: ベースラインモデルを評価する
</h2>

次に、静的な応答を返すベースラインモデルを評価します。ベースラインを設定しておくと、Leaderboard 上に比較の基準点ができるため、以降の各モデルが静的な実装に比べてどの程度改善したかを測定できます。

```python lines theme={"system"}
@weave.op
def baseline_model(zip_code: str):
    return {
        "city": "New York",
        "state": "NY",
        "avg_temp_f": 50.0,
        "population": 1000000,
        "median_income": 100000,
        "known_for": "The Big Apple",
    }

await evaluation.evaluate(baseline_model)
```

<h2 id="step-5-create-more-models">
  ステップ 5: モデルをさらに作成する
</h2>

次に、ベースラインと比較するモデルをさらに 2 つ作成します。一方のモデルには追加のプロンプトを与えずに郵便番号だけを渡し、もう一方には構造化されたプロンプトを渡します。これらを Leaderboard で比較すると、プロンプトのコンテキストが回答の品質にどう影響するかを確認できます。

```python lines theme={"system"}
@weave.op
def gpt_4o_mini_no_context(zip_code: str):
    client = OpenAI()

    completion = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": f"""Zip code {zip_code}"""}],
        response_format={
            "type": "json_schema",
            "json_schema": {
                "name": "response_format",
                "schema": Row.model_json_schema(),
            },
        },
    )

    return json.loads(completion.choices[0].message.content)

await evaluation.evaluate(gpt_4o_mini_no_context)
python
@weave.op
def gpt_4o_mini_with_context(zip_code: str):
    client = OpenAI()

    completion = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[
            {
                "role": "user",
                "content": f"""Please answer the following questions about the zip code {zip_code}:
                   1. What is the city?
                   2. What is the state?
                   3. What is the average temperature in Fahrenheit?
                   4. What is the population?
                   5. What is the median income?
                   6. What is the most well known thing about this zip code?
                   """,
            }
        ],
        response_format={
            "type": "json_schema",
            "json_schema": {
                "name": "response_format",
                "schema": Row.model_json_schema(),
            },
        },
    )

    return json.loads(completion.choices[0].message.content)

await evaluation.evaluate(gpt_4o_mini_with_context)
```

<h2 id="step-6-create-more-evaluations">
  ステップ 6: さらに評価を作成する
</h2>

次に、モデルと評価を組み合わせたマトリクスで評価を実行します。すべてのモデルを複数のデータセット (地域や年が異なるもの) で実行すると、Leaderboard が複数の条件でモデルをランク付けするのに必要なデータが揃います。

```python lines theme={"system"}
scorers = [
    check_concrete_fields,
    check_value_fields,
    check_subjective_fields,
]
evaluations = [
    weave.Evaluation(
        name="United States - 2022",
        dataset=weave.Dataset(
            name="United States - 2022",
            rows=generate_dataset_rows("United States", 5, 2022),
        ),
        scorers=scorers,
    ),
    weave.Evaluation(
        name="California - 2022",
        dataset=weave.Dataset(
            name="California - 2022", rows=generate_dataset_rows("California", 5, 2022)
        ),
        scorers=scorers,
    ),
    weave.Evaluation(
        name="United States - 2000",
        dataset=weave.Dataset(
            name="United States - 2000",
            rows=generate_dataset_rows("United States", 5, 2000),
        ),
        scorers=scorers,
    ),
]
models = [
    baseline_model,
    gpt_4o_mini_no_context,
    gpt_4o_mini_with_context,
]

for evaluation in evaluations:
    for model in models:
        await evaluation.evaluate(
            model, __weave={"display_name": evaluation.name + ":" + model.__name__}
        )
```

<h2 id="step-7-review-the-leaderboard">
  ステップ 7: Leaderboard を確認する
</h2>

評価結果をパブリッシュしたので、これらを Leaderboard にまとめて並べて比較できるようになりました。

新しい Leaderboard を作成するには、UI の Leaderboard タブに移動し、**Create Leaderboard** をクリックします。

Python から直接 Leaderboard を生成することもできます。

```python lines theme={"system"}
from weave.flow import leaderboard
from weave.trace.ref_util import get_ref

spec = leaderboard.Leaderboard(
    name="Zip Code World Knowledge",
    description="""
This leaderboard compares the performance of models in terms of world knowledge about zip codes.

### Columns

1. **State Match against `United States - 2022`**: The fraction of zip codes that the model correctly identified the state for.
2. **Avg Temp F Error against `California - 2022`**: The mean absolute error of the model's average temperature prediction.
3. **Correct Known For against `United States - 2000`**: The fraction of zip codes that the model correctly identified the most well known thing about the zip code.
""",
    columns=[
        leaderboard.LeaderboardColumn(
            evaluation_object_ref=get_ref(evaluations[0]).uri(),
            scorer_name="check_concrete_fields",
            summary_metric_path="state_match.true_fraction",
        ),
        leaderboard.LeaderboardColumn(
            evaluation_object_ref=get_ref(evaluations[1]).uri(),
            scorer_name="check_value_fields",
            should_minimize=True,
            summary_metric_path="avg_temp_f_err.mean",
        ),
        leaderboard.LeaderboardColumn(
            evaluation_object_ref=get_ref(evaluations[2]).uri(),
            scorer_name="check_subjective_fields",
            summary_metric_path="correct_known_for.true_fraction",
        ),
    ],
)

ref = weave.publish(spec)
```

これで、定義した 3 つの評価とスコアリングメトリクスに基づいて各モデルをランク付けする Leaderboard が Weave にパブリッシュされました。Weights & Biases の UI では、モデルごとのスコアの確認や個々の評価 run の詳細な分析に加え、今後のモデルの反復処理を同じベースラインと比較することもできます。
