> ## Documentation Index
> Fetch the complete documentation index at: https://docs.coreweave.com/llms.txt
> Use this file to discover all available pages before exploring further.

# ガードレールを設定する

> 本番アプリケーションで LLM の安全性を確保し、出力品質を測定します

ガードレールは、LLM ジャッジモデルのスコアに基づいて、LLM アプリケーションの動作に介入します。出力がユーザーに届く前にリアルタイムで動作し、スコアがしきい値を超えると応答をブロックまたは変更できます。ガードレールを使用すると、有害な内容のブロック、個人を特定できる情報 (PII) を含む応答のフィルタリング、ユーザーからの攻撃的な入力のブロックができます。

このガイドでは、Weave のガードレールの仕組みとパフォーマンスの調整方法を説明し、組み込みの Scorer、カスタム Scorer、AWS Bedrock Guardrails を使用して本番環境の LLM アプリケーションを保護する例を紹介します。

<h2 id="how-weave-guardrails-work">
  Weave ガードレールの仕組み
</h2>

Weave ガードレールは、インラインの [Weave Scorer](/ja/products/wandb/weave/guides/evaluation/scorers) を使用して、ユーザーからの入力または LLM からの出力を評価し、LLM の応答をリアルタイムで調整します。カスタム Scorer を設定するか、[組み込み Scorer](/ja/products/wandb/weave/guides/evaluation/builtin_scorers) を使用して、さまざまな目的でコンテンツを評価できます。このガイドでは、両方のタイプの Scorer をガードレールとして使用する方法を説明します。

本番トラフィックをアプリケーションの制御フローを変更せずにパッシブにスコアリングしたい場合は、代わりに [モニター](/ja/products/wandb/weave/guides/evaluation/monitors) を使用してください。

モニターとは異なり、ガードレールはアプリケーションの制御フローに影響を与えるため、コードの変更が必要です。ただし、ガードレールからのすべての Scorer 結果は Weave のデータベースに自動的に保存されるため、ガードレールは追加の設定なしでモニターとしても機能します。元々どのように使用されたかに関係なく、過去の Scorer 結果を分析できます。

<Note>
  Weave TypeScript SDK は、ガードレールを設定するために必要なツールをサポートしていません。
</Note>

<h3 id="optimize-your-weave-guardrail-performance">
  Weave ガードレールのパフォーマンスを最適化する
</h3>

ガードレールはアプリケーションの制御フローを中断し、応答の流れを変える可能性があるため、複雑すぎるとパフォーマンスに影響を与えることがあります。最適なパフォーマンスを得るには、以下の推奨事項に従ってください：

* ガードレールのロジックを最小限かつ高速に保ちます。
* 一般的な結果をキャッシュします。
* 重い外部 API 呼び出しを避けます。
* 繰り返しの初期化コストを避けるため、ガードレールをメイン関数外で初期化します。

ガードレールをメイン関数外で初期化することは、特に以下の場合に重要です：

* Scorer が ML モデルを読み込む場合。
* ローカル LLM を使用していてレイテンシーが重要な場合。
* Scorer がネットワーク接続を維持する場合。
* 高トラフィックのアプリケーションがある場合。

<h2 id="example-create-a-guardrail-using-a-built-in-moderation-scorer">
  例: 組み込みのモデレーション Scorer を使用してガードレールを作成する
</h2>

以下の例では、ユーザー プロンプトを OpenAI の GPT-4o mini モデルに送信します。モデルの応答は、LLM の応答に有害または毒性のあるコンテンツが含まれているかどうかを評価するため、[OpenAI の moderation API](https://platform.openai.com/docs/guides/moderation) に渡されます。モデルの応答は guardrail 関数 (`generate_safe_response()`) に渡され、この関数は `OpenAIModerationScorer` を使用して LLM の元の応答をチェックします。関数のロジックは、OpenAI の評価応答で `passed` フィールドの boolean をチェックし、アプリケーションがどのように応答するかを決定します。

```python lines {28-45} theme={"system"}
import weave
import openai
from weave.scorers import OpenAIModerationScorer
import asyncio

# Weave を初期化する
weave.init("your-team-name/your-project-name")

# OpenAI クライアントを初期化する
client = openai.OpenAI()  # Uses OPENAI_API_KEY env var

# モデレーション Scorer を初期化する
moderation_scorer = OpenAIModerationScorer()

# OpenAI にプロンプトを送信する
@weave.op
def generate_response(prompt: str) -> str:
    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[
            {"role": "system", "content": "You are a helpful assistant."},
            {"role": "user", "content": prompt}
        ],
        max_tokens=200
    )
    return response.choices[0].message.content

# guardrail 関数は応答の毒性をチェックする
async def generate_safe_response(prompt: str) -> str:
    """Generate a response with content moderation guardrail."""
    # 結果と Call オブジェクトの両方を取得する
    result, call = generate_response.call(prompt)
    
    # ユーザーに返す前にモデレーション Scorer を適用する
    score = await call.apply_scorer(moderation_scorer)
    print("This is the score object:", score)
    
    # コンテンツがフラグ付けされたかどうかを確認する
    if not score.result.get("passed", True): 
        categories = score.result.get("categories", {})
        flagged_categories = list(categories.keys()) if categories else []
        print(f"Content blocked. Flagged categories: {flagged_categories}")
        return "I'm sorry, I can't provide that response due to content policy restrictions."
    
    return result

# サンプルを実行する
if __name__ == "__main__":
    
    prompts = [
        "What's the capital of France?",
        "Tell me a funny fact about dogs.",
    ]
    
    for prompt in prompts:
        print(f"\nPrompt: {prompt}")
        response = asyncio.run(generate_safe_response(prompt))
        print(f"Response: {response}")
```

LLM-as-a-judge Scorer を使用する場合、Op の変数をスコアリング プロンプトで参照できます。たとえば、「`{output}` が `{ground_truth}` に基づいて正確かどうかを評価します。」 詳細については、[prompt variables](/ja/products/wandb/weave/guides/evaluation/scorers#access-variables-from-your-ops-in-scoring-prompts) を参照してください。

<h2 id="example-create-a-guardrail-using-a-custom-scorer">
  Example: Create a guardrail using a カスタム Scorer
</h2>

以下の例では、LLM の応答 (メールアドレス、電話番号、社会保障番号など) に含まれる 個人を特定できる情報 (PII) を検出する custom ガードレールを作成します。これにより、生成されたコンテンツに機密情報が露出するのを防ぎます。`generate_safe_response` 関数は、custom `PIIDetectionScorer` を適用します。

```python lines {14-39, 57-69} theme={"system"}
import weave
import openai
import re
import asyncio
from weave import Scorer

weave.init("your-team-name/your-project-name")

client = openai.OpenAI()

class PIIDetectionScorer(Scorer):
    """Detects PII in LLM outputs to prevent data leaks."""
    
    @weave.op
    def score(self, output: str) -> dict:
        """
        Check for common PII patterns in the output.
        
        Returns:
            dict: Contains 'passed' (bool) and 'detected_types' (list)
        """
        detected_types = []
        
        # メールパターン
        if re.search(r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b', output):
            detected_types.append("email")
        
        # 電話番号パターン（米国形式）
        if re.search(r'\b\d{3}[-.]?\d{3}[-.]?\d{4}\b', output):
            detected_types.append("phone")
        
        # SSNパターン
        if re.search(r'\b\d{3}-\d{2}-\d{4}\b', output):
            detected_types.append("ssn")
        
        return {
            "passed": len(detected_types) == 0,
            "detected_types": detected_types
        }

# パフォーマンスを最大化するために関数の外でscorerを初期化する
pii_scorer = PIIDetectionScorer()

@weave.op
def generate_response(prompt: str) -> str:
    """Generate a response using an LLM."""
    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[
            {"role": "system", "content": "You are a helpful assistant."},
            {"role": "user", "content": prompt}
        ],
        max_tokens=200
    )
    return response.choices[0].message.content

async def generate_safe_response(prompt: str) -> str:
    """Generate a response with PII detection guardrail."""
    result, call = generate_response.call(prompt)
    
    # PII検出scorerを適用する
    score = await call.apply_scorer(pii_scorer)
    
    # PIIが検出された場合は応答をブロックする
    if not score.result.get("passed", True):
        detected_types = score.result.get("detected_types", [])
        return f"I cannot provide a response that may contain sensitive information (detected: {', '.join(detected_types)})."
    
    return result

# 使用例
if __name__ == "__main__":
    prompts = [
        "What's the weather like today?",
        "Can you help me contact someone at john.doe@example.com?",
        "Tell me about machine learning.",
    ]
    
    for prompt in prompts:
        print(f"\nPrompt: {prompt}")
        response = asyncio.run(generate_safe_response(prompt))
        print(f"Response: {response}")
```

<h2 id="integrate-weave-with-aws-bedrock-guardrails">
  Weave を AWS Bedrock Guardrails と統合する
</h2>

AWS でコンテンツポリシーをすでに管理している場合は、Weave で `BedrockGuardrailScorer` を使用して適用できます。この Scorer は AWS Bedrock Guardrails を使用して、設定済みのポリシーに基づいてコンテンツを検出およびフィルターします。

Bedrock Guardrails のインテグレーションを設定する前に、以下が必要です。

* Bedrock へのアクセス権を持つ AWS アカウント。
* [AWS Bedrock コンソールで設定済みのガードレール](https://docs.aws.amazon.com/bedrock/latest/userguide/guardrails-components.html)。
* `boto3` [Python パッケージ](https://boto3.amazonaws.com/v1/documentation/api/latest/reference/services/s3.html)。

独自の Bedrock クライアントを作成する必要はありません。Weave が自動的に作成します。リージョンを指定するには、Scorer の `bedrock_runtime_kwargs` パラメーターにリージョン値を渡してください。

AWS Bedrock でガードレールを作成する方法の例については、[Bedrock guardrails notebook](https://github.com/aws-samples/amazon-bedrock-samples/blob/main/responsible_ai/bedrock-guardrails/guardrails-api.ipynb) を参照してください。

次の例では、テキスト生成を AWS Bedrock Guardrails のポリシーに対してチェックしてから、ユーザーに結果を返します。

```python theme={"system"}
import weave
from weave.scorers.bedrock_guardrails import BedrockGuardrailScorer

weave.init("your-team-name/your-project-name")

guardrail_scorer = BedrockGuardrailScorer(
    guardrail_id="your-guardrail-id",
    guardrail_version="DRAFT",
    source="INPUT",
    bedrock_runtime_kwargs={"region_name": "us-east-1"}
)

@weave.op
def generate_text(prompt: str) -> str:
    # ここにテキスト生成ロジックを記述します
    return "Generated text..."

async def generate_safe_text(prompt: str) -> str:
    result, call = generate_text.call(prompt)

    score = await call.apply_scorer(guardrail_scorer)

    if not score.result.passed:
        if score.result.metadata.get("modified_output"):
            return score.result.metadata["modified_output"]
        return "I cannot generate that content due to content policy restrictions."

    return result
```
