> ## Documentation Index
> Fetch the complete documentation index at: https://docs.coreweave.com/llms.txt
> Use this file to discover all available pages before exploring further.

# 퀵스타트: 맞춤형 에이전트 관측성 설정

> Weave SDK로 멀티턴 에이전트 트레이싱을 하세요. 대화, 턴, LLM Call, 도구 Call이 프로젝트의 Agents 뷰에 표시됩니다.

export const AgentLensBanner = ({href}) => <Tip>
    <strong>This workflow is also available in CoreWeave Agent Lens.</strong> Agent Lens is the Forge experience built for tracing, monitoring, and analyzing AI agents, with automated insights into agent failures and user intents. It uses the same trace data as Weights & Biases Weave, so the traces you already send appear there with nothing to migrate.{' '}
    <a href={href || '/products/agent-lens'}>{href ? 'See how to do this in Agent Lens' : 'Learn about Agent Lens'}</a>.
  </Tip>;

<AgentLensBanner href="/ko/products/agent-lens/get-started/custom-agents" />

[Colab에서 사용해 보기](https://colab.research.google.com/github/wandb/docs/blob/main/weave/cookbooks/source/custom-agents-quickstart.ipynb) · [GitHub 소스](https://github.com/wandb/docs/blob/main/weave/cookbooks/source/custom-agents-quickstart.ipynb)

Weave SDK를 사용하면 널리 쓰이는 SDK나 맞춤형 하니스로 구축한 에이전트 트레이싱을 할 수 있습니다. 이 퀵스타트에서는 직접 구축한 맞춤형 멀티턴 에이전트에 Weave를 수동으로 통합하고, OpenTelemetry span을 내보내고 캡처하는 방법을 설명합니다. 에이전트에서 Weave를 활용하는 방식에 대한 개념 설명은 [에이전트 트레이싱](/ko/products/wandb/weave/guides/tracking/trace-agents)을 참조하세요.

Claude Agent SDK나 Codex 같은 SDK 또는 하니스와 Weave를 통합하려면 [에이전트 인테그레이션 선택](/ko/products/wandb/weave/agent-integration-quickstart)을 참조하세요. Weave는 여러 에이전트 구축용 SDK와 에이전트 하니스를 자동 패치하므로 빠르게 통합할 수 있습니다.

<h2 id="what-youll-learn">
  학습 내용
</h2>

이 퀵스타트를 마치면 Weave와 호환되는 OTel span을 내보내는 멀티턴 에이전트를 직접 실행해 볼 수 있습니다. 또한 Weave가 대화, 턴, LLM Call, 도구 Call을 에이전트 코드에 어떻게 매핑하는지 이해할 수 있으므로, 같은 패턴을 직접 만든 맞춤형 에이전트에도 적용할 수 있습니다.

이 가이드의 코드는 Wikipedia에서 정보를 찾아볼 수 있는 간단한 리서치 에이전트를 구성합니다. 이 에이전트는 세 가지 질문(세 번의 턴)을 던지며, Wikipedia에서 답을 검색할 시점은 LLM이 판단합니다. Weave는 모든 단계(대화, 각 질문, 각 AI 응답, 각 Wikipedia 조회)를 기록하므로 Weave Agents 뷰에서 어떤 일이 일어났는지 확인할 수 있습니다.

이 가이드에서는 다음 내용을 다룹니다.

* `weave.init()`으로 에이전트 트레이싱용 Weave를 초기화합니다.
* `start_conversation` / `startConversation` 및 `start_turn` / `startTurn`으로 대화와 턴을 시작합니다.
* `start_llm` / `startLLM`으로 LLM Call을 래핑하고 사용량을 기록합니다.
* `start_tool` / `startTool`로 도구 실행을 래핑하고 결과를 기록합니다.
* 토큰 수와 비용이 표시되도록 전체 토큰 사용량과 가격 산정이 가능한 모델을 기록합니다.
* 생성된 대화, 턴, 도구 Call을 Agents 뷰에서 확인합니다.

<h2 id="how-the-weave-sdk-works-with-agents">
  Weave SDK가 에이전트와 함께 작동하는 방식
</h2>

Weave SDK에는 에이전트용 범용 OTel 수집 시스템이 포함되어 있어, Weave는 에이전트 코드의 모든 OTel span에서 정보를 캡처할 수 있습니다. 다만 Weights & Biases UI의 Agents 뷰에 에이전트의 트레이스를 렌더링하려면 Weave가 다음 span을 별도로 처리해야 합니다.

| 개념 | Python | TypeScript | OTel span |
| - | - | - | - |
| 하나의 대화 | `weave.start_conversation(...)` | `weave.startConversation(...)` | (span 없음, 턴을 그룹화) |
| 사용자 또는 에이전트 간의 한 차례 주고받기 | `weave.start_turn(...)` | `weave.startTurn(...)` | `invoke_agent` |
| 한 번의 LLM API 호출 | `weave.start_llm(...)` | `weave.startLLM(...)` | `chat` |
| 한 번의 도구 실행 | `weave.start_tool(...)` | `weave.startTool(...)` | `execute_tool` |

Python에서는 네 가지 함수 모두 컨텍스트 관리자(`with weave.start_*(...) as obj:`)로 작동합니다. 블록을 벗어나면 예외가 발생한 경우에도 span을 종료하고 속성을 플러시합니다. TypeScript에서는 반환된 각 객체에서 `.end()`를 호출하세요. 예외가 발생해도 정리 작업이 확실히 수행되도록 `try { ... } finally { obj.end(); }`를 사용하세요.

`gen_ai.usage.*`, `gen_ai.agent.name` 등 그 밖의 [GenAI 시맨틱 규칙 속성](https://opentelemetry.io/docs/specs/semconv/gen-ai/gen-ai-agent-spans/)을 사용하면 추가 렌더링이 가능하지만, 필수는 아닙니다.

<h2 id="prerequisites">
  사전 요구 사항
</h2>

* CoreWeave Forge 계정 및 [API 키](https://forge.coreweave.com/settings#apikeys)
* OpenAI API 키
* Python 3.10 이상(Python 예시 실행 시)
* Node.js 18 이상(TypeScript 예시는 내장 `fetch`가 필요함)

<h2 id="install-packages">
  패키지 설치
</h2>

개발 환경에 다음 패키지를 설치하세요.

<CodeGroup>
  ```bash Python theme={"system"}
  pip install weave openai requests
  ```

  ```bash TypeScript theme={"system"}
  npm install weave openai
  ```
</CodeGroup>

<h2 id="initialize-weave">
  Weave 초기화
</h2>

`weave.init()`는 W\&B 인증을 수행하고, 에이전트 span을 **Agents** 뷰로 전송하는 OTel 익스포터를 설정합니다. 팀에 해당 프로젝트가 없으면 처음 데이터를 기록할 때 Weave가 프로젝트를 자동으로 생성합니다.

<CodeGroup>
  ```python lines Python theme={"system"}
  import getpass
  import os

  os.environ["WANDB_API_KEY"] = getpass.getpass("Enter your CoreWeave Forge API key: ")
  os.environ["OPENAI_API_KEY"] = getpass.getpass("Enter your OpenAI API key: ")

  TEAM = input("Enter your CoreWeave Forge team name: ")
  PROJECT = input("Enter your Weights & Biases project name: ")

  import weave
  weave.init(f"{TEAM}/{PROJECT}")
  ```

  ```typescript lines highlight="4" TypeScript twoslash theme={"system"}
  // @noErrors
  // 이 프로젝트를 실행하기 전에 환경 변수 WANDB_API_KEY, OPENAI_API_KEY를 설정하세요
  import * as weave from 'weave';

  await weave.init(`[YOUR-TEAM]/[YOUR-PROJECT]`);
  ```
</CodeGroup>

<h2 id="define-a-tool">
  도구 정의하기
</h2>

다음 코드는 에이전트가 사용할 Wikipedia 검색 도구를 정의하고, 이 도구를 언제 어떻게 사용할지 지정하는 OpenAI 도구 스키마도 함께 정의합니다.

<CodeGroup>
  ```python lines Python theme={"system"}
  import json
  import requests

  def wikipedia_search(query: str) -> str:
      r = requests.get(
          "https://en.wikipedia.org/w/api.php",
          params={
              "action": "query", "generator": "search", "gsrsearch": query, "gsrlimit": 1,
              "prop": "extracts", "exintro": True, "explaintext": True, "format": "json",
          },
          headers={"User-Agent": "weave-demo"},
      ).json()
      return next(iter(r["query"]["pages"].values()))["extract"]

  wikipedia_tool_schema = {
      "type": "function",
      "function": {
          "name": "wikipedia_search",
          "description": "Search Wikipedia for a topic and return its intro paragraph.",
          "parameters": {
              "type": "object",
              "properties": {"query": {"type": "string"}},
              "required": ["query"],
          },
      },
  }
  ```

  ```typescript lines TypeScript twoslash theme={"system"}
  // @noErrors
  async function wikipediaSearch(query: string): Promise<string> {
    const url = new URL('https://en.wikipedia.org/w/api.php');
    url.search = new URLSearchParams({
      action: 'query',
      generator: 'search',
      gsrsearch: query,
      gsrlimit: '1',
      prop: 'extracts',
      exintro: 'true',
      explaintext: 'true',
      format: 'json',
    }).toString();
    const res = await fetch(url, { headers: { 'User-Agent': 'weave-demo' } });
    const data = (await res.json()) as {
      query: { pages: Record<string, { extract: string }> };
    };
    return Object.values(data.query.pages)[0].extract;
  }

  const wikipediaToolSchema = {
    type: 'function' as const,
    function: {
      name: 'wikipedia_search',
      description: 'Search Wikipedia for a topic and return its intro paragraph.',
      parameters: {
        type: 'object',
        properties: { query: { type: 'string' } },
        required: ['query'],
      },
    },
  };
  ```
</CodeGroup>

<h2 id="run-a-traced-multi-turn-agent">
  트레이스되는 멀티턴 에이전트 실행하기
</h2>

도구와 Weave 초기화가 준비되었으니, 다음 단계에서는 이 둘을 결합해 완전한 에이전트 루프를 구성합니다. 이 루프를 통해 대화, 턴, LLM Call, 도구 Call이 어떻게 중첩되는지 확인할 수 있습니다.

다음 예시에서는 하나의 대화에서 세 개의 턴을 실행합니다. 각 턴은 다음과 같이 동작합니다.

1. `chat` span을 열고 LLM이 도구 호출 여부를 결정하도록 합니다.
2. LLM이 도구를 요청하면 해당 호출을 감싸는 `execute_tool` span을 열고, 그 결과를 LLM에 다시 전달합니다.
3. 두 번째 `chat` span을 열어 최종 답변을 생성합니다.

<CodeGroup>
  ```python lines highlight="10,12,19,38,51,55,69" Python theme={"system"}
  import weave
  from openai import OpenAI

  openai_client = OpenAI()
  MODEL = "gpt-4o-mini"

  def run_turn(history, user_message):
      history.append({"role": "user", "content": user_message})

      with weave.start_turn(user_message=user_message, model=MODEL):
          # LLM Call 1: 모델이 도구를 사용하기로 결정할 수 있습니다.
          with weave.start_llm(model=MODEL, provider_name="openai") as llm:
              resp = openai_client.chat.completions.create(
                  model=MODEL, messages=history, tools=[wikipedia_tool_schema],
              )
              msg = resp.choices[0].message
              llm.output(msg.content or "")
              # record()는 사용량, 가격 산정 기준 모델, 응답 ID를 한 번의 호출로 설정합니다.
              llm.record(
                  usage=weave.Usage(
                      input_tokens=resp.usage.prompt_tokens,
                      output_tokens=resp.usage.completion_tokens,
                      cache_read_input_tokens=getattr(
                          resp.usage.prompt_tokens_details, "cached_tokens", 0
                      ),
                  ),
                  response_id=resp.id,
                  response_model=resp.model,
              )
              history.append(msg.model_dump(exclude_none=True))

          # 도구 요청이 없으면 첫 번째 LLM 응답이 곧 답변입니다.
          if not msg.tool_calls:
              return msg.content

          # 요청된 각 도구 Call을 실행합니다.
          for tc in msg.tool_calls:
              with weave.start_tool(
                  name=tc.function.name,
                  arguments=tc.function.arguments,
                  tool_call_id=tc.id,
              ) as tool:
                  tool.result = wikipedia_search(**json.loads(tc.function.arguments))
                  history.append({
                      "role": "tool",
                      "tool_call_id": tc.id,
                      "content": tool.result,
                  })

          # LLM Call 2: 최종 답변을 생성합니다.
          with weave.start_llm(model=MODEL, provider_name="openai") as llm:
              resp = openai_client.chat.completions.create(model=MODEL, messages=history)
              msg = resp.choices[0].message
              llm.output(msg.content)
              llm.record(
                  usage=weave.Usage(
                      input_tokens=resp.usage.prompt_tokens,
                      output_tokens=resp.usage.completion_tokens,
                      cache_read_input_tokens=getattr(
                          resp.usage.prompt_tokens_details, "cached_tokens", 0
                      ),
                  ),
                  response_id=resp.id,
                  response_model=resp.model,
              )
              history.append({"role": "assistant", "content": msg.content})
              return msg.content

  with weave.start_conversation(agent_name="research-bot") as conversation:
      history = []
      for question in [
          "Who founded Anthropic?",
          "What is Claude (the AI assistant)?",
          "Summarize what we discussed in one sentence.",
      ]:
          print(f"USER: {question}")
          print(f"AGENT: {run_turn(history, question)}\n")
  ```

  ```typescript lines highlight="11,14,25,47,62,70,89" theme={"system"}
  // @noErrors
  import * as weave from 'weave';
  import OpenAI from 'openai';

  const openaiClient = new OpenAI();
  const MODEL = 'gpt-4o-mini';

  // history는 OpenAI 채팅 메시지 목록입니다. 간결하게 작성하기 위해 유형을 느슨하게 지정했습니다.
  async function runTurn(history: any[], userMessage: string): Promise<string | null> {
    history.push({ role: 'user', content: userMessage });

    const turn = weave.startTurn({ model: MODEL });
    try {
      // LLM Call 1: 모델이 도구를 사용하기로 결정할 수 있습니다.
      const llm1 = weave.startLLM({ model: MODEL, providerName: 'openai' });
      let msg;
      try {
        const resp = await openaiClient.chat.completions.create({
          model: MODEL,
          messages: history,
          tools: [wikipediaToolSchema],
        });
        msg = resp.choices[0].message;
        llm1.output(msg.content ?? '');
        // record()는 사용량, 과금 기준 모델, 응답 ID를 한 번에 설정합니다.
        llm1.record({
          usage: {
            inputTokens: resp.usage?.prompt_tokens,
            outputTokens: resp.usage?.completion_tokens,
            cacheReadInputTokens: resp.usage?.prompt_tokens_details?.cached_tokens,
          },
          responseId: resp.id,
          responseModel: resp.model,
        });
        history.push(msg);
      } finally {
        llm1.end();
      }

      // 도구 요청이 없으면 첫 번째 LLM 응답이 곧 답변입니다.
      if (!msg.tool_calls?.length) {
        return msg.content ?? null;
      }

      // 요청된 각 도구 Call을 실행합니다.
      for (const tc of msg.tool_calls) {
        if (tc.type !== 'function') continue;
        const tool = weave.startTool({
          name: tc.function.name,
          args: tc.function.arguments,
          toolCallId: tc.id,
        });
        try {
          const { query } = JSON.parse(tc.function.arguments);
          tool.result = await wikipediaSearch(query);
          history.push({ role: 'tool', tool_call_id: tc.id, content: tool.result });
        } finally {
          tool.end();
        }
      }

      // LLM Call 2: 최종 답변을 종합합니다.
      const llm2 = weave.startLLM({ model: MODEL, providerName: 'openai' });
      try {
        const resp = await openaiClient.chat.completions.create({
          model: MODEL,
          messages: history,
        });
        const msg2 = resp.choices[0].message;
        llm2.output(msg2.content ?? '');
        llm2.record({
          usage: {
            inputTokens: resp.usage?.prompt_tokens,
            outputTokens: resp.usage?.completion_tokens,
            cacheReadInputTokens: resp.usage?.prompt_tokens_details?.cached_tokens,
          },
          responseId: resp.id,
          responseModel: resp.model,
        });
        history.push({ role: 'assistant', content: msg2.content });
        return msg2.content ?? null;
      } finally {
        llm2.end();
      }
    } finally {
      turn.end();
    }
  }

  const conversation = weave.startConversation({ agentName: 'research-bot' });
  try {
    const history: any[] = [];
    for (const question of [
      'Who founded Anthropic?',
      'What is Claude (the AI assistant)?',
      'Summarize what we discussed in one sentence.',
    ]) {
      console.log(`USER: ${question}`);
      console.log(`AGENT: ${await runTurn(history, question)}\n`);
    }
  } finally {
    conversation.end();
  }
  ```
</CodeGroup>

<h2 id="record-token-usage-and-cost">
  토큰 사용량 및 비용 기록
</h2>

각 `chat` span에는 토큰 사용량과 모델 ID가 포함됩니다. Weave는 사용량을 바탕으로 토큰 수를 표시하고, 사용량과 모델 ID를 함께 사용해 비용을 산출합니다. 따라서 값이 불완전하거나 가격을 책정할 수 없으면 트레이스의 나머지 부분이 정상으로 보이더라도 토큰이 `0 in / 0 out`으로 표시되거나 비용이 `Cost -`로 표시됩니다. `record(...)`를 사용하면 이러한 필드(`output_messages`, `response_id`, `reasoning` 등 포함)를 한 번의 호출로 설정할 수 있습니다. 이때 전달한 필드만 적용됩니다.

비용이 표시되려면 다음 두 가지 조건을 충족해야 합니다.

* **완전한 사용량.** `input_tokens`는 캐시된 토큰을 포함한 *전체* 입력 토큰 수입니다. Weave는 캐시 읽기와 캐시 쓰기에 각각의 요율을 적용하고 이를 입력 합계에서 차감합니다. 따라서 `cache_read_input_tokens`와 `cache_creation_input_tokens`는 이들을 포함한 전체 `input_tokens`와 *함께* 보고해야 합니다. 프롬프트 캐싱을 지원하는 공급자(예: Anthropic)에서는 캐시된 토큰이 입력의 대부분을 차지하는 경우가 많으므로, 이를 누락하면 사용량과 비용이 거의 0으로 표시됩니다.
* **가격 책정이 가능한 모델 ID.** 비용은 모델을 기준으로 조회됩니다. Weave는 `response_model`(공급자가 실제로 서빙한 정확한 모델)을 우선 사용하고, 이 값이 없으면 `start_llm`에 전달한 `model`을 사용합니다. `opus`나 `sonnet` 같은 별칭으로는 가격을 책정할 수 없어 `Cost -`로 표시되므로, 응답에서 반환된 구체적인 ID(`resp.model`)를 `response_model`로 전달하세요.

OpenAI는 캐시된 토큰을 `prompt_tokens`에 포함해 집계하므로 위 예시를 그대로 적용할 수 있습니다. 반면 Anthropic은 캐시된 토큰을 `input_tokens`와 *별도로* 보고하므로, Weave가 가격 책정에 사용하는 합계에 이를 다시 더해야 합니다.

<CodeGroup>
  ```python lines Python theme={"system"}
  with weave.start_llm(model=MODEL, provider_name="anthropic") as llm:
      resp = anthropic_client.messages.create(
          model=MODEL, max_tokens=1024, messages=history,
      )
      u = resp.usage
      llm.output(resp.content[0].text)
      llm.record(
          usage=weave.Usage(
              input_tokens=u.input_tokens
              + u.cache_read_input_tokens
              + u.cache_creation_input_tokens,
              output_tokens=u.output_tokens,
              cache_read_input_tokens=u.cache_read_input_tokens,
              cache_creation_input_tokens=u.cache_creation_input_tokens,
          ),
          response_id=resp.id,
          response_model=resp.model,
      )
  ```

  ```typescript lines TypeScript twoslash theme={"system"}
  // @noErrors
  const u = resp.usage;
  llm.record({
    usage: {
      inputTokens:
        u.input_tokens + u.cache_read_input_tokens + u.cache_creation_input_tokens,
      outputTokens: u.output_tokens,
      cacheReadInputTokens: u.cache_read_input_tokens,
      cacheCreationInputTokens: u.cache_creation_input_tokens,
    },
    responseId: resp.id,
    responseModel: resp.model,
  });
  ```
</CodeGroup>

<h2 id="see-your-agent-traces-in-the-agents-view">
  Agents 뷰에서 에이전트 트레이스 확인하기
</h2>

`weave.init()`이 실행되면 프로젝트 링크가 출력됩니다. 이 링크에서 다음 항목을 확인할 수 있습니다.

* **Agents** 탭의 `research-bot` 행
* 턴 세 개로 구성된 대화 하나
* 두 개의 `chat` span과 하나의 `execute_tool` span이 중첩된 각 턴(`invoke_agent`)
* 각 `chat`의 토큰 수, 지연 시간, 모델, 전체 메시지 교환 내용

턴을 클릭하면 입력, 출력, 도구 인수, 도구 결과를 자세히 확인할 수 있습니다.

<h2 id="link-to-a-conversation-from-your-app">
  앱에서 대화로 연결하기
</h2>

자체 UI에서 Weave Agents 뷰의 대화로 바로 이동하는 딥 링크를 만들려면 entity, 프로젝트, 대화 ID를 조합해 URL을 구성하세요. `weave.init()`은 `entity`와 `project`를 담은 클라이언트를 반환하고, `start_conversation`은 `conversation_id`를 제공합니다.

<CodeGroup>
  ```python lines Python theme={"system"}
  from weave.trace.urls import agent_conversation_path

  client = weave.init(f"{TEAM}/{PROJECT}")

  with weave.start_conversation(agent_name="research-bot") as conversation:
      # ... 턴 실행 ...
      url = agent_conversation_path(
          client.entity, client.project, conversation.conversation_id
      )
      print(f"View this conversation at {url}")
  ```

  ```plaintext TypeScript lines theme={"system"}
  이 기능은 TypeScript에서 사용할 수 없습니다.
  ```
</CodeGroup>

<h2 id="next-steps">
  다음 단계
</h2>

* [Weave로 에이전트 트레이싱](/ko/products/wandb/weave/guides/tracking/trace-agents)하는 방법과 Weave SDK에서 사용 가능한 기능 및 옵션을 알아보세요.
* Weave를 에이전트와 통합하는 다른 방법은 [에이전트 인테그레이션 선택](/ko/products/wandb/weave/agent-integration-quickstart)을 참조하세요.
