> ## Documentation Index
> Fetch the complete documentation index at: https://docs.coreweave.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Not Diamond のカスタムルーティング

> W&B Weave で Not Diamond のカスタムルーティングを使用する方法を説明します

<Note>
  これはインタラクティブなノートブックです。ローカルで実行するか、以下のリンクから利用できます。

  * [Google Colab で開く](https://colab.research.google.com/github/wandb/docs/blob/main/weave/cookbooks/source/notdiamond_custom_routing.ipynb)
  * [GitHub でソースを表示](https://github.com/wandb/docs/blob/main/weave/cookbooks/source/notdiamond_custom_routing.ipynb)
</Note>

このノートブックでは、Weave と [Not Diamond のカスタムルーティング](https://docs.notdiamond.ai/docs/router-training-quickstart) を使用して、評価結果に基づき LLM プロンプトを最適なモデルにルーティングする方法を紹介します。このノートブックを最後まで進めると、コーディング用プロンプト向けのカスタムルーターをトレーニングし、それを使って新しいプロンプトのターゲットモデルを選択したうえで、単体で最も優れたモデルとパフォーマンスを比較評価できます。

<h2 id="routing-prompts">
  プロンプトのルーティング
</h2>

このセクションでは、複数の LLM を扱う際にカスタムルーティングが役立つ理由を説明します。

複雑な LLM ワークフローを構築する際には、精度、コスト、呼び出しのレイテンシーに応じて、プロンプトを送るモデルを使い分ける必要が生じることがあります。
[Not Diamond](https://www.notdiamond.ai/) を使用すると、こうしたワークフロー内のプロンプトをニーズに最適なモデルへルーティングでき、モデルのコストを抑えつつ精度を最大限に高められます。

どのようなデータ分布であっても、単一のモデルがあらゆるクエリで他のすべてのモデルを上回ることはまれです。複数のモデルを組み合わせ、どの LLM をいつ呼び出すべきかを学習する「メタモデル」を構築することで、個々のどのモデルよりも高いパフォーマンスを実現できるうえ、コストとレイテンシーも削減できます。

<h2 id="custom-routing">
  カスタムルーティング
</h2>

プロンプト用のカスタムルーターをトレーニングするには、次の入力が必要です。

1. **LLM プロンプトのセット**: プロンプトは文字列である必要があります。また、アプリケーションで実際に使用されるプロンプトを代表するものにしてください。
2. **LLM の応答**: 各入力に対する候補 LLM の応答です。候補 LLM には、サポートされる LLM と独自のカスタムモデルの両方を含めることができます。
3. **入力に対する候補 LLM の応答の評価スコア**: スコアは数値で、ニーズに合った任意のメトリクスを使用できます。

これらを Not Diamond API に送信すると、ワークフローごとに調整されたカスタムルーターをトレーニングできます。

<h2 id="set-up-the-training-data">
  トレーニングデータを設定する
</h2>

このセクションでは、カスタムルーターの学習に使用するトレーニングデータとテストデータを準備します。

実際の運用では、独自の評価を使用してカスタムルーターをトレーニングします。ただし、このサンプルノートブックでは、[HumanEval データセット](https://github.com/openai/human-eval)に対する LLM の応答を使用して、コーディングタスク向けのカスタムルーターをトレーニングします。

まず、この例のために用意されたデータセットをダウンロードし、LLM の応答をモデルごとに解析して `EvaluationResults` に変換します。このトレーニングデータとテストデータの分割は、以降のセクションでルーターをトレーニングし、サンプル外データでのパフォーマンスを評価する際に使用します。

```python lines theme={"system"}
!curl -L "https://drive.google.com/uc?export=download&id=1q1zNZHioy9B7M-WRjsJPkfvFosfaHX38" -o humaneval.csv
python
import random

import weave
from weave.flow.dataset import Dataset
from weave.flow.eval import EvaluationResults
from weave.integrations.notdiamond.util import get_model_evals

pct_train = 0.8
pct_test = 1 - pct_train

# 実際には、独自のデータセットで評価を構築し、
# `evaluation.get_eval_results(model)` を呼び出します
model_evals = get_model_evals("./humaneval.csv")
model_train = {}
model_test = {}
for model, evaluation_results in model_evals.items():
    n_results = len(evaluation_results.rows)
    all_idxs = list(range(n_results))
    train_idxs = random.sample(all_idxs, k=int(n_results * pct_train))
    test_idxs = [idx for idx in all_idxs if idx not in train_idxs]

    model_train[model] = EvaluationResults(
        rows=weave.Table([evaluation_results.rows[idx] for idx in train_idxs])
    )
    model_test[model] = Dataset(
        rows=weave.Table([evaluation_results.rows[idx] for idx in test_idxs])
    )
    print(
        f"Found {len(train_idxs)} train rows and {len(test_idxs)} test rows for {model}."
    )
```

<h2 id="train-a-custom-router">
  カスタムルーターをトレーニングする
</h2>

`EvaluationResults` の準備ができたら、それを Not Diamond に送信してカスタムルーターをトレーニングできます。[アカウントを作成](https://app.notdiamond.ai/keys)して
[APIキーを生成](https://app.notdiamond.ai/keys)したうえで、以下のコードに APIキーを入力してください。APIキーはトレーニングリクエストの認証に使用されます。

<Frame>
  <img src="https://mintcdn.com/coreweave-dbfa0e8d/3Dv_sw2eg8feUJlx/products/wandb/weave/_media/api-keys.png?fit=max&auto=format&n=3Dv_sw2eg8feUJlx&q=85&s=1fd162cba9d876866cf649baa1165bf5" alt="APIキーを作成する" width="3454" height="1912" data-path="products/wandb/weave/_media/api-keys.png" />
</Frame>

```python lines theme={"system"}
import os

from weave.integrations.notdiamond.custom_router import train_router

api_key = os.getenv("NOTDIAMOND_API_KEY", "<YOUR_API_KEY>")

preference_id = train_router(
    model_evals=model_train,
    prompt_column="prompt",
    response_column="actual",
    language="en",
    maximize=True,
    api_key=api_key,
    # 初めてカスタムルーターをトレーニングする場合は、この行をコメントアウトしたままにします
    # 既存のカスタムルーターを上書きして再トレーニングする場合は、この行のコメントを解除します
    # preference_id=preference_id,
)
```

その後、Not Diamond アプリでカスタムルーターのトレーニングの進行状況を確認できます。

<Frame>
  <img src="https://mintcdn.com/coreweave-dbfa0e8d/3Dv_sw2eg8feUJlx/products/wandb/weave/_media/router-preferences.png?fit=max&auto=format&n=3Dv_sw2eg8feUJlx&q=85&s=a9bf84f724ddcb3402acc242f4d610d7" alt="ルーターのトレーニングの進行状況を確認する" width="3456" height="1916" data-path="products/wandb/weave/_media/router-preferences.png" />
</Frame>

カスタムルーターのトレーニングが完了したら、そのルーターを使用してプロンプトをルーティングできます。

```python lines theme={"system"}
from notdiamond import NotDiamond

import weave

weave.init("notdiamond-quickstart")

llm_configs = [
    "anthropic/claude-3-5-sonnet-20240620",
    "openai/gpt-4o-2024-05-13",
    "google/gemini-1.5-pro-latest",
    "openai/gpt-4-turbo-2024-04-09",
    "anthropic/claude-3-opus-20240229",
]
client = NotDiamond(api_key=api_key, llm_configs=llm_configs)

new_prompt = (
    """
You are a helpful coding assistant. Using the provided function signature, write the implementation for the function
in Python. Write only the function. Do not include any other text.

from typing import List

def has_close_elements(numbers: List[float], threshold: float) -> bool:
    """
    """ Check if in given list of numbers, are any two numbers closer to each other than
    given threshold.
    >>> has_close_elements([1.0, 2.0, 3.0], 0.5)
    False
    >>> has_close_elements([1.0, 2.8, 3.0, 4.0, 5.0, 2.0], 0.3)
    True
    """
    """
"""
)
session_id, routing_target_model = client.model_select(
    messages=[{"role": "user", "content": new_prompt}],
    preference_id=preference_id,
)

print(f"Session ID: {session_id}")
print(f"Target Model: {routing_target_model}")
```

この例では、Not Diamond が Weave の自動トレースに対応している点も活用しています。結果は Weights & Biases UI で確認できます。

<img src="https://mintcdn.com/coreweave-dbfa0e8d/3Dv_sw2eg8feUJlx/products/wandb/weave/_media/weave-trace.png?fit=max&auto=format&n=3Dv_sw2eg8feUJlx&q=85&s=b667e5a2ed5658032f94d890bd27de27" alt="カスタムルーティングの Weights & Biases UI" width="3340" height="1794" data-path="products/wandb/weave/_media/weave-trace.png" />

<h2 id="evaluate-your-custom-router">
  カスタムルーターを評価する
</h2>

このセクションでは、カスタムルーターが単体で最も優れたモデルを上回るかどうかを測定する方法を説明します。

カスタムルーターのトレーニングが完了したら、次のいずれかの方法でパフォーマンスを評価できます。

* トレーニングに使用したプロンプトを送信して、サンプル内パフォーマンスを評価する。
* 新しいプロンプトまたはホールドアウトしたプロンプトを送信して、サンプル外パフォーマンスを評価する。

次の例では、テストセットをカスタムルーターに送信してパフォーマンスを評価します。

```python lines theme={"system"}
from weave.integrations.notdiamond.custom_router import evaluate_router

eval_prompt_column = "prompt"
eval_response_column = "actual"

best_provider_model, nd_model = evaluate_router(
    model_datasets=model_test,
    prompt_column=eval_prompt_column,
    response_column=eval_response_column,
    api_key=api_key,
    preference_id=preference_id,
)
python
@weave.op()
def is_correct(score: int, output: dict) -> dict:
    # モデルの応答はすでに取得済みなので、score をそのまま流用する簡易的な方法をとっています
    return {"correct": score}

best_provider_eval = weave.Evaluation(
    dataset=best_provider_model.model_results.to_dict(orient="records"),
    scorers=[is_correct],
)
await best_provider_eval.evaluate(best_provider_model)

nd_eval = weave.Evaluation(
    dataset=nd_model.model_results.to_dict(orient="records"), scorers=[is_correct]
)
await nd_eval.evaluate(nd_model)
```

この例では、Not Diamond の"メタモデル"が、プロンプトを複数の異なるモデルにルーティングします。

Weave を使ってカスタムルーターをトレーニングすると、評価も実行され、その結果が Weights & Biases UI にアップロードされます。カスタムルーターのプロセスが完了したら、Weights & Biases UI で結果を確認できます。

UI では、Not Diamond の"メタモデル"が、プロンプトに正確に回答できる可能性がより高い他のモデルへプロンプトをルーティングすることで、単体で最も高いパフォーマンスを示すモデルを上回っていることを確認できます。

<Frame>
  <img src="https://mintcdn.com/coreweave-dbfa0e8d/3Dv_sw2eg8feUJlx/products/wandb/weave/_media/evaluations.png?fit=max&auto=format&n=3Dv_sw2eg8feUJlx&q=85&s=80658faaf13331ef07310a78b9b46db6" alt="Not Diamond の評価" width="3332" height="1792" data-path="products/wandb/weave/_media/evaluations.png" />
</Frame>
