> ## Documentation Index
> Fetch the complete documentation index at: https://docs.coreweave.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Verdict

> Weave로 Verdict 평가 프레임워크를 사용하여 LLM 평가 파이프라인을 트레이스하고 모니터링하세요

<a target="_blank" href="https://github.com/wandb/examples/blob/master/weave/docs/quickstart_verdict.ipynb">
  <img src="https://colab.research.google.com/assets/colab-badge.svg" alt="Open In Colab" />
</a>

Weave는 [Verdict Python 라이브러리](https://verdict.haizelabs.com/docs/)를 통해 이루어지는 모든 Call을 자동으로 추적하고 로깅합니다.

AI 평가 파이프라인을 다룰 때는 디버깅이 중요합니다. 파이프라인 단계가 실패하거나, 출력이 예상과 다르거나, 중첩된 오퍼레이션 때문에 흐름을 파악하기 어려울 때 문제의 원인을 정확히 찾아내기란 쉽지 않습니다. Verdict 애플리케이션은 대개 여러 파이프라인 단계, 평가자, 변환으로 구성되므로 평가 워크플로의 내부 동작을 파악해 두면 도움이 됩니다.

Weave는 [Verdict](https://verdict.readthedocs.io/) 애플리케이션의 트레이스를 자동으로 캡처하여 이 과정을 간소화합니다. 이를 통해 파이프라인의 성능을 모니터링하고 분석하면서 AI 평가 워크플로를 디버깅하고 최적화할 수 있습니다.

<h2 id="getting-started">
  시작하기
</h2>

Verdict 파이프라인에서 Weave 트레이싱을 활성화하려면 스크립트 시작 부분에서 `weave.init(project=...)`를 호출하세요. 특정 W\&B 팀에 로깅하려면 `project` 인수에 `team-name/project-name` 형식으로 지정하고, 기본 팀 또는 entity에 로깅하려면 `project-name`만 전달하세요.

```python lines {7} theme={"system"}
import weave
from verdict import Pipeline
from verdict.common.judge import JudgeUnit
from verdict.schema import Schema

# 프로젝트 이름으로 Weave 초기화
weave.init("verdict_demo")

# 단순한 평가 파이프라인 생성
pipeline = Pipeline()
pipeline = pipeline >> JudgeUnit().prompt("Rate the quality of this text: {source.text}")

# 샘플 데이터 생성
data = Schema.of(text="This is a sample text for evaluation.")

# 파이프라인 실행 - Weave가 자동으로 트레이스를 기록합니다
output = pipeline.run(data)

print(output)
```

<h2 id="tracking-call-metadata">
  Call 메타데이터 추적
</h2>

Verdict 파이프라인 Call에 맞춤형 메타데이터를 추가하려면 [`weave.attributes`](/ko/products/wandb/weave/reference/python-sdk#function-attributes) 컨텍스트 관리자를 사용하세요. 이 컨텍스트 관리자로 파이프라인 run이나 평가 배치와 같은 특정 코드 블록에 태그를 지정해 두면, 나중에 Weights & Biases UI에서 관련 트레이스를 필터링하고 그룹화할 수 있습니다.

```python lines {7,14} theme={"system"}
import weave
from verdict import Pipeline
from verdict.common.judge import JudgeUnit
from verdict.schema import Schema

# 프로젝트 이름을 지정하여 Weave를 초기화합니다
weave.init("verdict_demo")

pipeline = pipeline >> JudgeUnit().prompt("Evaluate sentiment: {source.text}")

data = Schema.of(text="I love this product!")

with weave.attributes({"evaluation_type": "sentiment", "batch_id": "batch_001"}):
    output = pipeline.run(data)

print(output)
```

Weave는 Verdict 파이프라인 Call의 트레이스에 메타데이터를 자동으로 기록합니다. 기록된 메타데이터는 Weave 웹 인터페이스에서 확인할 수 있습니다.

<h2 id="traces">
  트레이스
</h2>

AI 평가 파이프라인의 트레이스를 중앙 데이터베이스에 저장해 두면 개발 단계와 프로덕션 환경 모두에서 유용합니다. 저장된 트레이스는 평가 워크플로를 디버깅하고 개선하는 데 도움이 되며, 유용한 데이터셋으로도 활용할 수 있습니다.

Weave는 Verdict 애플리케이션의 트레이스를 자동으로 캡처합니다. Verdict 라이브러리를 통해 이루어지는 모든 Call을 추적하고 로깅하며, 여기에는 다음이 포함됩니다.

* `Pipeline` 실행 단계
* `JudgeUnit` 평가
* `Layer` 변환
* 풀링 오퍼레이션
* 맞춤형 유닛 및 변환

Weave 웹 인터페이스에서 트레이스를 확인할 수 있으며, 이 인터페이스에는 파이프라인 실행의 계층 구조가 표시됩니다.

<h2 id="pipeline-tracing-example">
  파이프라인 트레이싱 예시
</h2>

다음 예시에서는 Weave가 중첩된 파이프라인 오퍼레이션을 트레이스하는 방식을 보여 줍니다. 이 예시를 통해 다단계 Verdict 파이프라인의 각 단계가 어떻게 캡처되는지 확인할 수 있습니다.

```python lines {8} theme={"system"}
import weave
from verdict import Pipeline, Layer
from verdict.common.judge import JudgeUnit
from verdict.transform import MeanPoolUnit
from verdict.schema import Schema

# 프로젝트 이름으로 Weave 초기화
weave.init("verdict_demo")

# 여러 단계로 구성된 파이프라인 생성
pipeline = Pipeline()
pipeline = pipeline >> Layer([
    JudgeUnit().prompt("Rate coherence: {source.text}"),
    JudgeUnit().prompt("Rate relevance: {source.text}"),
    JudgeUnit().prompt("Rate accuracy: {source.text}")
], 3)
pipeline = pipeline >> MeanPoolUnit()

# 샘플 데이터
data = Schema.of(text="This is an evaluation of text quality across multiple dimensions.")

# 파이프라인 실행 - Weave가 모든 오퍼레이션을 자동으로 트레이스함
result = pipeline.run(data)

print(f"Average score: {result}")
```

이렇게 하면 다음 내용을 보여 주는 상세한 트레이스가 생성됩니다.

* 메인 `Pipeline` 실행
* `Layer` 내 각 `JudgeUnit`의 평가
* `MeanPoolUnit` 집계 단계
* 각 오퍼레이션의 소요 시간 정보

<h2 id="configuration">
  설정
</h2>

`weave.init()`을 호출하면 Weave가 Verdict 파이프라인의 트레이싱을 자동으로 활성화합니다. 이 인테그레이션은 `Pipeline.__init__()` 메서드를 패치하고, 모든 트레이스 데이터를 Weave로 전달하는 `VerdictTracer`를 주입하는 방식으로 작동합니다.

별도의 설정은 필요하지 않습니다. Weave는 다음 작업을 자동으로 수행합니다.

* 모든 파이프라인 오퍼레이션을 캡처합니다.
* 실행 시간을 추적합니다.
* 입력과 출력을 로깅합니다.
* 트레이스 계층 구조를 유지합니다.
* 파이프라인의 동시 실행을 처리합니다.

<h2 id="custom-tracers-and-weave">
  맞춤형 트레이서와 Weave
</h2>

애플리케이션에서 이미 맞춤형 Verdict 트레이서를 사용하고 있다면 Weave의 `VerdictTracer`를 기존 트레이서와 함께 실행할 수 있습니다. 따라서 여러 인테그레이션 중 하나만 골라야 할 필요가 없습니다.

```python lines {8} theme={"system"}
import weave
from verdict import Pipeline
from verdict.common.judge import JudgeUnit
from verdict.util.tracing import ConsoleTracer
from verdict.schema import Schema

# 프로젝트 이름으로 Weave를 초기화합니다
weave.init("verdict_demo")

# Verdict의 기본 제공 트레이서도 그대로 사용할 수 있습니다
console_tracer = ConsoleTracer()

# Weave(자동) 트레이싱과 Console 트레이싱을 함께 사용하는 파이프라인을 생성합니다
pipeline = Pipeline(tracer=[console_tracer])  # Weave 트레이서는 자동으로 추가됩니다
pipeline = pipeline >> JudgeUnit().prompt("Evaluate: {source.text}")

data = Schema.of(text="Sample evaluation text")

# Weave와 콘솔 양쪽에 트레이스가 기록됩니다
result = pipeline.run(data)
```

<h2 id="models-and-evaluations">
  모델과 평가
</h2>

여러 파이프라인 컴포넌트로 구성된 AI 시스템은 체계적으로 관리하고 평가하기가 까다로울 수 있습니다. [`weave.Model`](/ko/products/wandb/weave/guides/core-types/models)을 사용하면 프롬프트, 파이프라인 설정, 평가 매개변수 같은 실험 세부 정보를 캡처하고 정리할 수 있어 여러 반복 버전을 더 쉽게 비교할 수 있습니다.

다음 예시는 Verdict 파이프라인을 `weave.Model`로 래핑하는 방법을 보여 줍니다.

```python lines {8,10,14} theme={"system"}
import asyncio
import weave
from verdict import Pipeline
from verdict.common.judge import JudgeUnit
from verdict.schema import Schema

# 프로젝트 이름으로 Weave를 초기화합니다
weave.init("verdict_demo")

class TextQualityEvaluator(weave.Model):
    judge_prompt: str
    pipeline_name: str

    @weave.op()
    async def predict(self, text: str) -> dict:
        pipeline = Pipeline(name=self.pipeline_name)
        pipeline = pipeline >> JudgeUnit().prompt(self.judge_prompt)
        
        data = Schema.of(text=text)
        result = pipeline.run(data)
        
        return {
            "text": text,
            "quality_score": result.score if hasattr(result, 'score') else result,
            "evaluation_prompt": self.judge_prompt
        }

model = TextQualityEvaluator(
    judge_prompt="Rate the quality of this text on a scale of 1-10: {source.text}",
    pipeline_name="text_quality_evaluator"
)

text = "This is a well-written and informative piece of content that provides clear value to readers."

prediction = asyncio.run(model.predict(text))

# Jupyter Notebook에서는 다음을 실행하세요:
# prediction = await model.predict(text)

print(prediction)
```

이 코드는 Weights & Biases UI에서 시각화할 수 있는 모델을 생성합니다. UI에서는 파이프라인 구조와 평가 결과를 함께 확인할 수 있습니다.

<h3 id="evaluations">
  평가
</h3>

평가를 사용하면 평가 파이프라인 자체의 성능을 측정할 수 있습니다. [`weave.Evaluation`](/ko/products/wandb/weave/guides/core-types/evaluations) 클래스를 사용하면 Verdict 파이프라인이 특정 작업이나 데이터셋에서 얼마나 잘 작동하는지 캡처할 수 있습니다.

```python lines {8} theme={"system"}
import asyncio
import weave
from verdict import Pipeline
from verdict.common.judge import JudgeUnit
from verdict.schema import Schema

# Weave 초기화
weave.init("verdict_demo")

# 평가 모델 생성
class SentimentEvaluator(weave.Model):
    @weave.op()
    async def predict(self, text: str) -> dict:
        pipeline = Pipeline()
        pipeline = pipeline >> JudgeUnit().prompt(
            "Classify sentiment as positive, negative, or neutral: {source.text}"
        )
        
        data = Schema.of(text=text)
        result = pipeline.run(data)
        
        return {"sentiment": result}

# 테스트 데이터
texts = [
    "I love this product, it's amazing!",
    "This is terrible, worst purchase ever.",
    "The weather is okay today."
]
labels = ["positive", "negative", "neutral"]

examples = [
    {"id": str(i), "text": texts[i], "target": labels[i]}
    for i in range(len(texts))
]

# 점수화 함수
@weave.op()
def sentiment_accuracy(target: str, output: dict) -> dict:
    predicted = output.get("sentiment", "").lower()
    return {"correct": target.lower() in predicted}

model = SentimentEvaluator()

evaluation = weave.Evaluation(
    dataset=examples,
    scorers=[sentiment_accuracy],
)

scores = asyncio.run(evaluation.evaluate(model))
# Jupyter Notebook에서는 대신 다음을 실행하세요:
# scores = await evaluation.evaluate(model)

print(scores)
```

이렇게 하면 다양한 테스트 케이스에서 Verdict 파이프라인이 어떤 성능을 내는지 보여 주는 평가 트레이스가 생성됩니다.

<h2 id="best-practices">
  모범 사례
</h2>

다음 섹션에서는 Verdict 파이프라인에서 Weave를 사용할 때 성능을 모니터링하고 오류를 처리하는 모범 사례를 설명합니다.

<h3 id="performance-monitoring">
  성능 모니터링
</h3>

Weave는 모든 파이프라인 오퍼레이션의 실행 시간 정보를 자동으로 캡처합니다. 이 정보를 활용하면 여러 run에 걸쳐 성능 병목 지점을 파악할 수 있습니다.

```python lines {6} theme={"system"}
import weave
from verdict import Pipeline, Layer
from verdict.common.judge import JudgeUnit
from verdict.schema import Schema

weave.init("verdict_demo")

# 성능 편차가 발생할 수 있는 파이프라인 생성
pipeline = Pipeline()
pipeline = pipeline >> Layer([
    JudgeUnit().prompt("Quick evaluation: {source.text}"),
    JudgeUnit().prompt("Detailed analysis: {source.text}"),  # 더 느릴 수 있음
], 2)

data = Schema.of(text="Sample text for performance testing")

# 여러 번 실행하여 실행 시간 패턴 확인
for i in range(3):
    with weave.attributes({"run_number": i}):
        result = pipeline.run(data)
```

<h3 id="error-handling">
  오류 처리
</h3>

Weave는 파이프라인 실행 중에 발생하는 예외를 자동으로 캡처하므로, 애플리케이션에서 예외를 처리하더라도 해당 실패가 트레이스에 기록됩니다.

```python lines {6} theme={"system"}
import weave
from verdict import Pipeline
from verdict.common.judge import JudgeUnit
from verdict.schema import Schema

weave.init("verdict_demo")

pipeline = Pipeline()
pipeline = pipeline >> JudgeUnit().prompt("Process: {source.invalid_field}")  # 이 부분에서 오류가 발생합니다

data = Schema.of(text="Sample text")

try:
    result = pipeline.run(data)
except Exception as e:
    print(f"Pipeline failed: {e}")
    # 오류 세부 정보는 Weave 트레이스에 캡처됩니다
```

Weave를 Verdict와 통합하면 AI 평가 파이프라인을 한눈에 파악할 수 있어 평가 워크플로를 더 쉽게 디버그하고, 최적화하고, 이해할 수 있습니다.


## Related topics

- [RAG 애플리케이션 평가](/ko/products/wandb/weave/tutorial-rag.md)
