> ## Documentation Index
> Fetch the complete documentation index at: https://docs.coreweave.com/llms.txt
> Use this file to discover all available pages before exploring further.

# W&B Models で Weave を使用する

> W&B Models での実験のトラッキングと、Weave での LLM トレースおよび評価を組み合わせた手順を解説するインタラクティブなノートブックです。

このノートブックでは、W\&B Models と Weave を組み合わせて検索拡張生成 (RAG) アプリケーションを構築、評価、パブリッシュする、エンドツーエンドのワークフローを順を追って説明します。Registry からファインチューン済みチャットモデルを取得して、Weave で追跡している既存の `RagModel` に組み込み、`weave.Evaluation` で更新後のアプリケーションを評価したうえで、新しい RAG モデルを Registry にパブリッシュします。このワークフローは、W\&B Models でモデルのトレーニングやファインチューニングを行い、そのモデルを Weave で追跡・評価している LLM アプリケーションに統合したいチームを対象としています。

<Note>
  これはインタラクティブなノートブックです。ローカルで実行するか、次のリンクから利用できます。

  * [Google Colab で開く](https://colab.research.google.com/github/wandb/docs/blob/main/weave/cookbooks/source/Models_and_Weave_Integration_Demo.ipynb)
  * [GitHub でソースを表示](https://github.com/wandb/docs/blob/main/weave/cookbooks/source/Models_and_Weave_Integration_Demo.ipynb)
</Note>

<h2 id="prerequisites">
  前提条件
</h2>

まず、必要なライブラリをインストールし、APIキーを設定したうえで Forge にログインし、新しい Weights & Biases project を作成します。

1. `pip` を使用して `weave`、`pandas`、`unsloth`、`wandb`、`litellm`、`pydantic`、`torch`、`faiss-gpu` をインストールします。

```python lines theme={"system"}
%%capture
!pip install weave wandb pandas pydantic litellm faiss-gpu
python
%%capture
!pip install unsloth
# 最新のナイトリー版 Unsloth もインストールします
!pip uninstall unsloth -y && pip install --upgrade --no-cache-dir "unsloth[colab-new] @ git+https://github.com/unslothai/unsloth.git"
```

2. 環境から必要な APIキーを追加します。

```python lines theme={"system"}
import os

from google.colab import userdata

os.environ["WANDB_API_KEY"] = userdata.get("WANDB_API_KEY")  # W&B Models と Weave 用
os.environ["OPENAI_API_KEY"] = userdata.get(
    "OPENAI_API_KEY"
)  # OpenAI - 検索用の埋め込みに使用
os.environ["GEMINI_API_KEY"] = userdata.get(
    "GEMINI_API_KEY"
)  # Gemini - ベースのチャットモデルに使用
```

3. Forge にログインし、新しい project を作成します。

```python lines theme={"system"}
import pandas as pd
import wandb

import weave

wandb.login()

PROJECT = "weave-cookboook-demo"
ENTITY = "wandb-smle"

weave.init(ENTITY + "/" + PROJECT)
```

<h2 id="download-chatmodel-from-registry-and-implement-unslothlorachatmodel">
  Registry から `ChatModel` をダウンロードして `UnslothLoRAChatModel` を実装する
</h2>

このシナリオでは、Model Team がパフォーマンス最適化のために `unsloth` ライブラリを使用して Llama-3.2 モデルをすでにファインチューニングしており、そのモデルは W\&B Registry で利用できます。このステップでは、ファインチューニングした [`ChatModel`](https://forge.coreweave.com/wandb/wandb-smle/weave-cookboook-demo/weave/object-versions?filter=%7B%22objectName%22%3A%22RagModel%22%7D\&peekPath=%2Fwandb-smle%2Fweave-rag-experiments%2Fobjects%2FChatModelRag%2Fversions%2F2mhdPb667uoFlXStXtZ0MuYoxPaiAXj3KyLS1kYRi84%3F%26) を Registry から取得し、[`RagModel`](https://forge.coreweave.com/wandb/wandb-smle/weave-cookboook-demo/weave/object-versions?filter=%7B%22objectName%22%3A%22RagModel%22%7D\&peekPath=%2Fwandb-smle%2Fweave-cookboook-demo%2Fobjects%2FRagModel%2Fversions%2FcqRaGKcxutBWXyM0fCGTR1Yk2mISLsNari4wlGTwERo%3F%26) で使用できるように `weave.Model` に変換します。

<Note>
  以下のコードで参照している `RagModel` は、完全な RAG アプリケーションと見なせるトップレベルの `weave.Model` です。`RagModel` には `ChatModel`、ベクトルデータベース、プロンプトが含まれます。`ChatModel` も `weave.Model` であり、W\&B Registry からアーティファクトをダウンロードするコードを含んでいます。`ChatModel` はモジュールとして差し替えられるため、`RagModel` の一部として他の任意の LLM チャットモデルに対応させることができます。詳細については、[Weave でモデルを確認してください](https://forge.coreweave.com/wandb/wandb-smle/weave-cookboook-demo/weave/evaluations?peekPath=%2Fwandb-smle%2Fweave-cookboook-demo%2Fobjects%2FRagModel%2Fversions%2Fx7MzcgHDrGXYHHDQ9BA8N89qDwcGkdSdpxH30ubm8ZM%3F%26)。
</Note>

`ChatModel` を読み込むには、アダプターとともに `unsloth.FastLanguageModel` または `peft.AutoPeftModelForCausalLM` を使用します。これにより、アプリへ効率的に組み込めます。Registry からモデルをダウンロードしたら、`model_post_init` メソッドで初期化と予測のロジックを設定します。このステップに必要なコードは Registry の **Use** タブに用意されており、そのまま実装にコピーできます。

以下のコードでは、Registry から取得したファインチューニング済みの Llama-3.2 モデルを管理、初期化、使用するための `UnslothLoRAChatModel` クラスを定義します。`UnslothLoRAChatModel` は、推論を最適化するために `unsloth.FastLanguageModel` を使用します。`model_post_init` メソッドはモデルのダウンロードと設定を行い、`predict` メソッドはユーザーのクエリを処理して応答を生成します。ユースケースに合わせてコードを調整するには、`MODEL_REG_URL` をファインチューニングしたモデルの正しい Registry パスに更新し、ハードウェアや要件に応じて `max_seq_length` や `dtype` などのパラメーターを調整してください。

```python lines theme={"system"}
from typing import Any

from pydantic import PrivateAttr
from unsloth import FastLanguageModel

import weave

class UnslothLoRAChatModel(weave.Model):
    """
    We define an extra ChatModel class to be able store and version more parameters than just the model name.
    Especially, relevant if we consider fine-tuning (locally or aaS) because of specific parameters.
    """

    chat_model: str
    cm_temperature: float
    cm_max_new_tokens: int
    cm_quantize: bool
    inference_batch_size: int
    dtype: Any
    device: str
    _model: Any = PrivateAttr()
    _tokenizer: Any = PrivateAttr()

    def model_post_init(self, __context):
        # Registry の "Use" タブからそのまま貼り付けられます
        run = wandb.init(project=PROJECT, job_type="model_download")
        artifact = run.use_artifact(f"{self.chat_model}")
        model_path = artifact.download()

        # unsloth 版（ネイティブで 2 倍高速な推論を有効化）
        self._model, self._tokenizer = FastLanguageModel.from_pretrained(
            model_name=model_path,
            max_seq_length=self.cm_max_new_tokens,
            dtype=self.dtype,
            load_in_4bit=self.cm_quantize,
        )
        FastLanguageModel.for_inference(self._model)

    @weave.op()
    async def predict(self, query: list[str]) -> dict:
        # add_generation_prompt = true - 生成時には必須です
        input_ids = self._tokenizer.apply_chat_template(
            query,
            tokenize=True,
            add_generation_prompt=True,
            return_tensors="pt",
        ).to("cuda")

        output_ids = self._model.generate(
            input_ids=input_ids,
            max_new_tokens=64,
            use_cache=True,
            temperature=1.5,
            min_p=0.1,
        )

        decoded_outputs = self._tokenizer.batch_decode(
            output_ids[0][input_ids.shape[1] :], skip_special_tokens=True
        )

        return "".join(decoded_outputs).strip()
python
MODEL_REG_URL = "wandb32/wandb-registry-RAG Chat Models/Finetuned Llama-3.2:v3"

max_seq_length = 2048  # 任意の値を指定できます。RoPE Scaling は内部で自動的にサポートされます。
dtype = (
    None  # None で自動検出。Tesla T4、V100 では Float16、Ampere 以降では Bfloat16
)
load_in_4bit = True  # 4bit 量子化を使用してメモリ使用量を削減します。False にすることもできます。

new_chat_model = UnslothLoRAChatModel(
    name="UnslothLoRAChatModelRag",
    chat_model=MODEL_REG_URL,
    cm_temperature=1.0,
    cm_max_new_tokens=max_seq_length,
    cm_quantize=load_in_4bit,
    inference_batch_size=max_seq_length,
    dtype=dtype,
    device="auto",
)
python
await new_chat_model.predict(
    [{"role": "user", "content": "What is the capital of Germany?"}]
)
```

<h2 id="integrate-the-new-chatmodel-version-into-ragmodel">
  新しい `ChatModel` バージョンを `RagModel` に統合する
</h2>

ファインチューン済みチャットモデルを使って RAG アプリケーションを構築すると、パイプライン全体を作り直すことなく、カスタマイズしたコンポーネントを再利用できます。このステップでは、Weave プロジェクトから既存の `RagModel` を取得し、その `ChatModel` を更新して、ファインチューニングしたモデルを使用するようにします。新しいチャットモデルに差し替えても、ベクトルデータベースやプロンプトなどの他のコンポーネントには影響しません。そのため、アプリケーション全体の構造を維持したまま、パフォーマンスを向上させることができます。

次のコードでは、まず Weave プロジェクトの ref を使用して `RagModel` オブジェクトを取得します。続いて、`RagModel` の `chat_model` 属性を更新し、前のステップで作成した新しい `UnslothLoRAChatModel` インスタンスを使用するようにします。その後、更新した `RagModel` をパブリッシュして新しいバージョンを作成します。最後に、更新した `RagModel` でサンプルの予測クエリを実行し、新しいチャットモデルが使用されていることを確認します。

```python lines theme={"system"}
RagModel = weave.ref(
    "weave://wandb-smle/weave-cookboook-demo/object/RagModel:cqRaGKcxutBWXyM0fCGTR1Yk2mISLsNari4wlGTwERo"
).get()
python
RagModel.chat_model.chat_model
python
await RagModel.predict("When was the first conference on climate change?")
python
# MAGIC: chat_model を差し替えて新しいバージョンをパブリッシュします（他の RAG コンポーネントに手を加える必要はありません）
RagModel.chat_model = new_chat_model
python
RagModel.chat_model.chat_model
python
# 予測時に新しいバージョンが参照されるよう、先に新しいバージョンをパブリッシュします
PUB_REFERENCE = weave.publish(RagModel, "RagModel")
python
await RagModel.predict("When was the first conference on climate change?")
```

<h2 id="run-a-weaveevaluation">
  `weave.Evaluation` を実行する
</h2>

更新した `RagModel` をパブリッシュしたら、次は新しいファインチューン済みチャットモデルがアプリケーション内で期待どおりに動作するかを確認します。このステップでは、既存の `weave.Evaluation` を使用して、更新した `RagModel` のパフォーマンスを評価します。これにより、新しいファインチューン済みチャットモデルが RAG アプリケーション内で期待どおりに動作することを確認できます。インテグレーションを効率化し、Models チームと Apps チームが連携しやすくなるよう、評価結果はモデルの W\&B run と Weave の workspace の両方にログします。

Models の場合:

* 評価のサマリーは、ファインチューン済みチャットモデルのダウンロードに使用した W\&B run にログされます。これには、分析用に [Workspace ビュー](https://forge.coreweave.com/wandb/wandb-smle/weave-cookboook-demo/workspace?nw=eglm8z7o9) に表示されるサマリー メトリクスとグラフが含まれます。
* 評価のトレース ID が run の設定に追加され、Weave ページに直接リンクされます。これにより、Model Team によるトレーサビリティが向上します。

Weave の場合:

* `ChatModel` のアーティファクトまたは Registry のリンクが、`RagModel` への入力として保存されます。
* コンテキストを補うため、W\&B の run ID が評価トレースに追加の列として保存されます。

次のコードは、評価オブジェクトを取得し、更新した `RagModel` で評価を実行して、結果を W\&B と Weave の両方にログする方法を示しています。評価 ref (`WEAVE_EVAL`) がお使いの project の設定と一致していることを確認してください。

```python lines theme={"system"}
# MAGIC: 評価用データセットと scorer を含む評価を取得するだけで、そのまま使用できます
WEAVE_EVAL = "weave://wandb-smle/weave-cookboook-demo/object/climate_rag_eval:ntRX6qn3Tx6w3UEVZXdhIh1BWGh7uXcQpOQnIuvnSgo"
climate_rag_eval = weave.ref(WEAVE_EVAL).get()
python
with weave.attributes({"wandb-run-id": wandb.run.id}):
    # 評価トレースを Models に保存するため、.call 属性を使用して結果と Call の両方を取得します
    summary, call = await climate_rag_eval.evaluate.call(climate_rag_eval, RagModel)
python
# Models にログします
wandb.run.log(pd.json_normalize(summary, sep="/").to_dict(orient="records")[0])
wandb.run.config.update(
    {"weave_url": f"https://forge.coreweave.com/wandb/wandb-smle/weave-cookboook-demo/r/call/{call.id}"}
)
wandb.run.finish()
```

<h2 id="save-the-new-rag-model-to-the-registry">
  新しい RAG モデルを Registry に保存する
</h2>

更新した `RagModel` の評価が完了したら、最後のステップとして、他のチームが検索して再利用できるように Registry にパブリッシュします。Models チームと Apps チームの両方が更新した `RagModel` を今後使用できるように、reference artifact として Registry にプッシュします。

次のコードは、更新した `RagModel` の `weave` オブジェクトのバージョンと名前を取得し、それらを使用して参照リンクを作成します。続いて、モデルの Weave URL をメタデータに含む新しいアーティファクトを W\&B に作成します。最後に、このアーティファクトを Registry にログし、指定した Registry パスにリンクします。

コードを実行する前に、`ENTITY` 変数と `PROJECT` 変数が W\&B の設定と一致していることを確認し、正しいリンク先の Registry パスを指定してください。このプロセスで新しい `RagModel` を W\&B エコシステムにパブリッシュすることで、コラボレーションと再利用が可能になり、ワークフローが完了します。

このセクションのコードを実行すると、更新した `RagModel` が reference artifact として Registry で使用できるようになり、W\&B Models と Weave の間のラウンドトリップが完了します。

```python lines theme={"system"}
MODELS_OBJECT_VERSION = PUB_REFERENCE.digest  # weave オブジェクトのバージョン
MODELS_OBJECT_NAME = PUB_REFERENCE.name  # weave オブジェクト名
python
models_url = f"https://wandb.ai/{ENTITY}/{PROJECT}/weave/objects/{MODELS_OBJECT_NAME}/versions/{MODELS_OBJECT_VERSION}"
models_link = (
    f"weave://{ENTITY}/{PROJECT}/object/{MODELS_OBJECT_NAME}:{MODELS_OBJECT_VERSION}"
)

with wandb.init(project=PROJECT, entity=ENTITY) as run:
    # 新しいアーティファクトを作成
    artifact_model = wandb.Artifact(
        name="RagModel",
        type="model",
        description="Models Link from RagModel in Weave",
        metadata={"url": models_url},
    )
    artifact_model.add_reference(models_link, name="model", checksum=False)

    # 新しいアーティファクトをログする
    run.log_artifact(artifact_model, aliases=[MODELS_OBJECT_VERSION])

    # Registry にリンク
    run.link_artifact(
        artifact_model, target_path="wandb32/wandb-registry-RAG Models/RAG Model"
    )
```
