> ## Documentation Index
> Fetch the complete documentation index at: https://docs.coreweave.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Verdict

> Use Verdict evaluation framework with Weave to トレース and monitor your LLM 評価パイプライン

<a target="_blank" href="https://github.com/wandb/examples/blob/master/weave/docs/quickstart_verdict.ipynb">
  <img src="https://colab.research.google.com/assets/colab-badge.svg" alt="Open In Colab" />
</a>

Weave は [Verdict Python ライブラリ](https://verdict.haizelabs.com/docs/) を通じて行われるすべての Call を自動的にトラッキングし、ログするように設計されています。

AI 評価パイプラインを扱う場合、デバッグが重要です。パイプラインのステップが失敗したり、出力が想定外になったり、ネストされた操作が混乱を招いたりする場合、問題の特定は困難です。Verdict アプリケーションは通常、複数のパイプラインステップ、judges、および変換で構成されるため、評価ワークフローの内部動作を理解することが役立ちます。

Weave は [Verdict](https://verdict.readthedocs.io/) アプリケーションのトレースを自動的に取得することで、このプロセスを効率化します。これにより、パイプラインのパフォーマンスを監視および分析して、AI 評価ワークフローをデバッグおよび最適化できます。

<h2 id="getting-started">
  はじめに
</h2>

Verdict パイプラインで Weave トレースを有効にするには、スクリプトの冒頭で `weave.init(project=...)` を呼び出します。`project` 引数を使用して、特定の W\&B チームに `team-name/project-name` の形式でログするか、`project-name` を渡してデフォルトのチームまたは entity にログします。

```python lines {7} theme={"system"}
import weave
from verdict import Pipeline
from verdict.common.judge import JudgeUnit
from verdict.schema import Schema

# Weave をプロジェクト名で初期化します
weave.init("verdict_demo")

# シンプルな評価パイプラインを作成します
pipeline = Pipeline()
pipeline = pipeline >> JudgeUnit().prompt("Rate the quality of this text: {source.text}")

# Create sample data
data = Schema.of(text="This is a sample text for evaluation.")

# パイプラインを実行します - Weave が自動的にトレースします
output = pipeline.run(data)

print(output)
```

<h2 id="tracking-call-metadata">
  Call メタデータのトラッキング
</h2>

Verdict パイプラインの Call にカスタムメタデータを付与するには、[`weave.attributes`](/ja/products/wandb/weave/reference/python-sdk#function-attributes) コンテキストマネージャーを使用します。このコンテキストマネージャーを使うと、パイプラインの run や評価バッチなど、特定のコードブロックにタグを付けられます。これにより、後から Weights & Biases UI で関連するトレースをフィルタリングしたり、グループ化したりできます。

```python lines {7,14} theme={"system"}
import weave
from verdict import Pipeline
from verdict.common.judge import JudgeUnit
from verdict.schema import Schema

# プロジェクト名を指定して Weave を初期化します
weave.init("verdict_demo")

pipeline = Pipeline()
pipeline = pipeline >> JudgeUnit().prompt("Evaluate sentiment: {source.text}")

data = Schema.of(text="I love this product!")

with weave.attributes({"evaluation_type": "sentiment", "batch_id": "batch_001"}):
    output = pipeline.run(data)

print(output)
```

Weave は、Verdict パイプラインの Call のトレースに紐づけてメタデータを自動的にトラッキングします。メタデータは Weave の Web インターフェースで確認できます。

<h2 id="traces">
  トレース
</h2>

AI 評価パイプラインのトレースを一元管理されたデータベースに保存しておくと、開発と本番運用の両方で役立ちます。これらのトレースは評価ワークフローのデバッグや改善に役立つだけでなく、有用なデータセットとしても活用できます。

Weave は Verdict アプリケーションのトレースを自動的に取得し、Verdict ライブラリを通じて行われるすべての Call をトラッキングしてログします。対象には次のものが含まれます。

* `Pipeline` の実行ステップ
* `JudgeUnit` による評価
* `Layer` による変換
* プーリング処理
* カスタムユニットとカスタム変換

トレースは Weave の Web インターフェースで確認できます。Web インターフェースには、パイプライン実行の階層構造が表示されます。

<h2 id="pipeline-tracing-example">
  パイプライン トレースの例
</h2>

以下の例は、Weave がネストされたパイプライン操作をトレースする方法を示しており、マルチステージの Verdict パイプラインの各ステップがどのように取得されるかを確認できます：

```python lines {8} theme={"system"}
import weave
from verdict import Pipeline, Layer
from verdict.common.judge import JudgeUnit
from verdict.transform import MeanPoolUnit
from verdict.schema import Schema

# プロジェクト名を指定して Weave を初期化します
weave.init("verdict_demo")

# 複数のステップで構成されるパイプラインを作成します
pipeline = Pipeline()
pipeline = pipeline >> Layer([
    JudgeUnit().prompt("Rate coherence: {source.text}"),
    JudgeUnit().prompt("Rate relevance: {source.text}"),
    JudgeUnit().prompt("Rate accuracy: {source.text}")
], 3)
pipeline = pipeline >> MeanPoolUnit()

# サンプルデータ
data = Schema.of(text="This is an evaluation of text quality across multiple dimensions.")

# パイプラインを実行します。Weave がすべての操作をトレースします
result = pipeline.run(data)

print(f"Average score: {result}")
```

これにより、詳細なトレースが表示されます:

* メインの `Pipeline` 実行。
* `Layer` 内の各 `JudgeUnit` 評価。
* `MeanPoolUnit` の集約ステップ。
* 各操作のタイミング情報。

<h2 id="configuration">
  設定
</h2>

`weave.init()` を呼び出すと、Weave は Verdict パイプラインのトレースを自動的に有効にします。このインテグレーションは `Pipeline.__init__()` メソッドをパッチして `VerdictTracer` を注入し、すべてのトレースデータを Weave に転送することで動作します。

追加の設定は必要ありません。Weave は自動的に以下の操作を行います。

* すべてのパイプライン操作を取得します。
* 実行タイミングをトラッキングします。
* 入力と出力をログします。
* トレース階層を維持します。
* 同時パイプライン実行を処理します。

<h2 id="custom-tracers-and-weave">
  カスタムトレーサーと Weave
</h2>

アプリケーションですでにカスタムの Verdict トレーサーを使用している場合でも、Weave の `VerdictTracer` はそれらと併用できるため、どちらかのインテグレーションを選ぶ必要はありません。

```python lines {8} theme={"system"}
import weave
from verdict import Pipeline
from verdict.common.judge import JudgeUnit
from verdict.util.tracing import ConsoleTracer
from verdict.schema import Schema

# プロジェクト名を指定して Weave を初期化します
weave.init("verdict_demo")

# Verdict の組み込みトレーサーも引き続き使用できます
console_tracer = ConsoleTracer()

# Weave (自動) と Console の両方でトレースするパイプラインを作成します
pipeline = Pipeline(tracer=[console_tracer])  # Weave トレーサーは自動的に追加されます
pipeline = pipeline >> JudgeUnit().prompt("Evaluate: {source.text}")

data = Schema.of(text="Sample evaluation text")

# Weave とコンソールの両方にトレースが出力されます
result = pipeline.run(data)
```

<h2 id="models-and-evaluations">
  モデルと評価
</h2>

複数のパイプラインコンポーネントで構成される AI システムの整理と評価は、容易ではありません。[`weave.Model`](/ja/products/wandb/weave/guides/core-types/models) を使用すると、プロンプト、パイプラインの設定、評価パラメーターなどの実験の詳細を取得して整理できるため、イテレーションごとの比較が容易になります。

次の例では、Verdict パイプラインを `weave.Model` でラップする方法を示します。

```python lines {8,10,14} theme={"system"}
import asyncio
import weave
from verdict import Pipeline
from verdict.common.judge import JudgeUnit
from verdict.schema import Schema

# プロジェクト名を指定して Weave を初期化します
weave.init("verdict_demo")

class TextQualityEvaluator(weave.Model):
    judge_prompt: str
    pipeline_name: str

    @weave.op()
    async def predict(self, text: str) -> dict:
        pipeline = Pipeline(name=self.pipeline_name)
        pipeline = pipeline >> JudgeUnit().prompt(self.judge_prompt)
        
        data = Schema.of(text=text)
        result = pipeline.run(data)
        
        return {
            "text": text,
            "quality_score": result.score if hasattr(result, 'score') else result,
            "evaluation_prompt": self.judge_prompt
        }

model = TextQualityEvaluator(
    judge_prompt="Rate the quality of this text on a scale of 1-10: {source.text}",
    pipeline_name="text_quality_evaluator"
)

text = "This is a well-written and informative piece of content that provides clear value to readers."

prediction = asyncio.run(model.predict(text))

# Jupyter ノートブックで実行する場合は、次のようにします:
# prediction = await model.predict(text)

print(prediction)
```

このコードで作成したモデルは Weights & Biases UI で可視化でき、パイプラインの構造と評価結果の両方を確認できます。

<h3 id="evaluations">
  評価
</h3>

評価では、評価パイプライン自体のパフォーマンスを測定できます。[`weave.Evaluation`](/ja/products/wandb/weave/guides/core-types/evaluations) クラスを使用すると、特定のタスクやデータセットに対する Verdict パイプラインのパフォーマンスを取得できます。

```python lines {8} theme={"system"}
import asyncio
import weave
from verdict import Pipeline
from verdict.common.judge import JudgeUnit
from verdict.schema import Schema

# Weave を初期化します
weave.init("verdict_demo")

# 評価モデルを作成します
class SentimentEvaluator(weave.Model):
    @weave.op()
    async def predict(self, text: str) -> dict:
        pipeline = Pipeline()
        pipeline = pipeline >> JudgeUnit().prompt(
            "Classify sentiment as positive, negative, or neutral: {source.text}"
        )
        
        data = Schema.of(text=text)
        result = pipeline.run(data)
        
        return {"sentiment": result}

# テストデータ
texts = [
    "I love this product, it's amazing!",
    "This is terrible, worst purchase ever.",
    "The weather is okay today."
]
labels = ["positive", "negative", "neutral"]

examples = [
    {"id": str(i), "text": texts[i], "target": labels[i]}
    for i in range(len(texts))
]

# スコアリング関数
@weave.op()
def sentiment_accuracy(target: str, output: dict) -> dict:
    predicted = output.get("sentiment", "").lower()
    return {"correct": target.lower() in predicted}

model = SentimentEvaluator()

evaluation = weave.Evaluation(
    dataset=examples,
    scorers=[sentiment_accuracy],
)

scores = asyncio.run(evaluation.evaluate(model))
# Jupyter ノートブックの場合は、次を実行します:
# scores = await evaluation.evaluate(model)

print(scores)
```

これにより、さまざまなテストケースにおける Verdict パイプラインのパフォーマンスを示す評価トレースが作成されます。

<h2 id="best-practices">
  ベストプラクティス
</h2>

以下のセクションでは、Verdict パイプラインで Weave を使用する際のパフォーマンスのモニタリングとエラーの処理に関するベストプラクティスを説明します。

<h3 id="performance-monitoring">
  パフォーマンスモニタリング
</h3>

Weave はすべてのパイプライン操作のタイミング情報を自動的に取得するので、run 全体のパフォーマンスボトルネックを特定するために使用できます：

```python lines {6} theme={"system"}
import weave
from verdict import Pipeline, Layer
from verdict.common.judge import JudgeUnit
from verdict.schema import Schema

weave.init("verdict_demo")

# パフォーマンスにばらつきが生じる可能性のあるパイプラインを作成します
pipeline = Pipeline()
pipeline = pipeline >> Layer([
    JudgeUnit().prompt("Quick evaluation: {source.text}"),
    JudgeUnit().prompt("Detailed analysis: {source.text}"),  # こちらは処理に時間がかかる可能性があります
], 2)

data = Schema.of(text="Sample text for performance testing")

# 複数回実行して処理時間の傾向を確認します
for i in range(3):
    with weave.attributes({"run_number": i}):
        result = pipeline.run(data)
```

<h3 id="error-handling">
  エラー処理
</h3>

Weave はパイプラインの実行中に発生した例外を自動的に取得します。そのため、アプリケーション側で例外を処理した場合でも、失敗はトレースに記録されます。

```python lines {6} theme={"system"}
import weave
from verdict import Pipeline
from verdict.common.judge import JudgeUnit
from verdict.schema import Schema

weave.init("verdict_demo")

pipeline = Pipeline()
pipeline = pipeline >> JudgeUnit().prompt("Process: {source.invalid_field}")  # ここでエラーが発生します

data = Schema.of(text="Sample text")

try:
    result = pipeline.run(data)
except Exception as e:
    print(f"Pipeline failed: {e}")
    # エラーの詳細は Weave のトレースに記録されます
```

Weave を Verdict と統合すると、AI 評価パイプラインを可視化できるため、評価ワークフローのデバッグや最適化、理解が容易になります。


## Related topics

- [RAG アプリケーションを評価する](/ja/products/wandb/weave/tutorial-rag.md)
