> ## Documentation Index
> Fetch the complete documentation index at: https://docs.coreweave.com/llms.txt
> Use this file to discover all available pages before exploring further.

# 評価の概要

> W&B Weave を使った評価の基本を学びます

<Note>
  これはインタラクティブなノートブックです。ローカルで実行するか、以下のリンクから利用できます。

  * [Google Colab で開く](https://colab.research.google.com/github/wandb/docs/blob/main/weave/cookbooks/source/Intro_to_Weave_Hello_Eval.ipynb)
  * [GitHub でソースを表示](https://github.com/wandb/docs/blob/main/weave/cookbooks/source/Intro_to_Weave_Hello_Eval.ipynb)
</Note>

このノートブックでは、最小構成のエンドツーエンドの例を通じて、W\&B Weave の評価を紹介します。Weave `Model` を定義して小規模なデータセットで実行し、カスタムのスコアリング関数で出力をスコア付けしたうえで、その結果を Weave で確認します。Weave を初めて使用する開発者を対象としており、より高度な評価ワークフローに取り組む前に、実際に手を動かしながらすばやく始められる内容になっています。

<h2 id="prerequisites">
  前提条件
</h2>

Weave の評価を実行する前に、以下の前提条件を満たしてください。

1. W\&B Weave SDK をインストールし、[APIキー](https://forge.coreweave.com/settings#apikeys)でログインします。
2. OpenAI SDK をインストールし、[APIキー](https://platform.openai.com/api-keys)でログインします。
3. Weights & Biases の project を初期化します。

```python lines theme={"system"}
# 依存関係のインストールとインポート
!pip install wandb weave openai -q

import os
from getpass import getpass

from openai import OpenAI
from pydantic import BaseModel

import weave

# 🔑 APIキーを設定します
# このセルを実行すると、`getpass` で APIキーの入力を求められます。入力内容はターミナルに表示されません。
#####
print("---")
print(
    "Create a W&B API key at: https://forge.coreweave.com/settings#apikeys"
)
os.environ["WANDB_API_KEY"] = getpass("Enter your W&B API key: ")
print("---")
print("You can generate your OpenAI API key here: https://platform.openai.com/api-keys")
os.environ["OPENAI_API_KEY"] = getpass("Enter your OpenAI API key: ")
print("---")
#####

# 🏠 W&B のプロジェクト名を入力します
weave_client = weave.init("MY_PROJECT_NAME")  # 🐝 W&B のプロジェクト名
```

<h2 id="run-your-first-evaluation">
  最初の評価を実行する
</h2>

環境の設定が完了したので、評価を定義して実行してみましょう。

次のコードサンプルは、Weave の `Model` API と `Evaluation` API を使用して LLM を評価する方法を示しています。まず、`weave.Model` をサブクラス化して Weave モデルを定義します。このとき、モデル名とプロンプトの形式を指定し、`predict` メソッドを `@weave.op` でトラッキングします。`predict` メソッドは OpenAI にプロンプトを送信し、Pydantic スキーマ (`FruitExtract`) を使用して応答を構造化された出力にパースします。次に、入力文と期待されるターゲットで構成される小規模な評価データセットを作成します。続いて、モデルの出力をターゲットラベルと比較するカスタムのスコアリング関数を定義します (この関数も `@weave.op` でトラッキングします)。最後に、すべてを `weave.Evaluation` でラップしてデータセットと scorer を指定し、`evaluate()` を呼び出して評価パイプラインを非同期で実行します。

```python lines theme={"system"}
# 1. Weave モデルを構築する
class FruitExtract(BaseModel):
    fruit: str
    color: str
    flavor: str

class ExtractFruitsModel(weave.Model):
    model_name: str
    prompt_template: str

    @weave.op()
    def predict(self, sentence: str) -> dict:
        client = OpenAI()

        response = client.beta.chat.completions.parse(
            model=self.model_name,
            messages=[
                {
                    "role": "user",
                    "content": self.prompt_template.format(sentence=sentence),
                }
            ],
            response_format=FruitExtract,
        )
        result = response.choices[0].message.parsed
        return result

model = ExtractFruitsModel(
    name="gpt4o",
    model_name="gpt-4o",
    prompt_template='Extract fields ("fruit": <str>, "color": <str>, "flavor": <str>) as json, from the following text : {sentence}',
)

# 2. サンプルをいくつか用意する
sentences = [
    "There are many fruits that were found on the recently discovered planet Goocrux. There are neoskizzles that grow there, which are purple and taste like candy.",
    "Pounits are a bright green color and are more savory than sweet.",
    "Finally, there are fruits called glowls, which have a very sour and bitter taste which is acidic and caustic, and a pale orange tinge to them.",
]
labels = [
    {"fruit": "neoskizzles", "color": "purple", "flavor": "candy"},
    {"fruit": "pounits", "color": "green", "flavor": "savory"},
    {"fruit": "glowls", "color": "orange", "flavor": "sour, bitter"},
]
examples = [
    {"id": "0", "sentence": sentences[0], "target": labels[0]},
    {"id": "1", "sentence": sentences[1], "target": labels[1]},
    {"id": "2", "sentence": sentences[2], "target": labels[2]},
]

# 3. 評価用のスコアリング関数を定義する
@weave.op()
def fruit_name_score(target: dict, output: FruitExtract) -> dict:
    target_flavors = [f.strip().lower() for f in target["flavor"].split(",")]
    output_flavors = [f.strip().lower() for f in output.flavor.split(",")]
    # ターゲットの flavor のいずれかが出力の flavor に含まれているかを確認する
    matches = any(tf in of for tf in target_flavors for of in output_flavors)
    return {"correct": matches}

# 4. 評価を実行する
evaluation = weave.Evaluation(
    name="fruit_eval",
    dataset=examples,
    scorers=[fruit_name_score],
)
await evaluation.evaluate(model)
```

評価が完了すると、Weave はモデル、データセット、およびサンプルごとのスコアを project にログします。これにより、Weights & Biases UI で結果を確認できます。

<h2 id="looking-for-more-examples">
  その他のサンプル
</h2>

基本的な評価を実行できたら、次のチュートリアルでより高度なワークフローを確認してみましょう。

* [評価パイプラインをエンドツーエンドで構築する](/ja/products/wandb/weave/tutorial-eval)方法を学ぶ。
* [RAG アプリケーションを評価する](/ja/products/wandb/weave/tutorial-rag)方法を学ぶ。
