> ## Documentation Index
> Fetch the complete documentation index at: https://docs.coreweave.com/llms.txt
> Use this file to discover all available pages before exploring further.

# LlamaIndex

> Weave を使用して LlamaIndex アプリケーションをトレースおよびデバッグし、LLM Call、RAG パイプライン、エージェントのステップ、評価を自動的に取得します。

このガイドでは、Weave を使用して [LlamaIndex](https://docs.llamaindex.ai/en/stable/) アプリケーションをトレース、デバッグ、評価する方法を説明します。このガイドを通じて、[LlamaIndex Python ライブラリ](https://github.com/run-llama/llama_index) 経由で行われた Call を Weave が自動的に取得する仕組みを学びます。これにより、独自のログ用コードを書かなくても、RAG パイプライン、エージェントのステップ、LLM Call を監視できます。このガイドは、LlamaIndex で LLM アプリケーションを構築している開発者のうち、デバッグ、パフォーマンス分析、評価のためにワークフローの可視性を高めたい方を対象としています。

LLM を扱う以上、デバッグは避けて通れません。モデルの Call の失敗、出力形式の崩れ、ネストされたモデルの Call による混乱など、問題の原因を特定するのは容易ではありません。LlamaIndex アプリケーションは複数のステップと LLM Call の invocation で構成されることが多いため、チェーンやエージェントの内部動作を把握することが重要です。

Weave は LlamaIndex アプリケーションのトレースを自動的に取得し、このプロセスを効率化します。アプリケーションのパフォーマンスを監視・分析できるため、LLM ワークフローのデバッグや最適化に役立ちます。さらに、Weave は評価ワークフローもサポートします。

<h2 id="get-started">
  はじめに
</h2>

まず、スクリプトの冒頭で `weave.init()` を呼び出します。これにより Weave が初期化され、以降に実行されるすべての LlamaIndex の Call についてトレースの取得が開始されます。`weave.init()` の引数にはプロジェクト名を指定します。プロジェクト名はトレースを整理するために使用されます。

```python lines {5} theme={"system"}
import weave
from llama_index.core.chat_engine import SimpleChatEngine

# プロジェクト名を指定して Weave を初期化します
weave.init("llamaindex_demo")

chat_engine = SimpleChatEngine.from_defaults()
response = chat_engine.chat(
    "Say something profound and romantic about fourth of July"
)
print(response)
```

前の例では、内部で OpenAI の Call を行う基本的な LlamaIndex チャットエンジンを作成しています。このコードを実行すると、Weave がチャットエンジンの実行トレースを取得するため、Weave の Web インターフェースで確認できます。以下のトレースを参照してください。

[<img src="https://mintcdn.com/coreweave-dbfa0e8d/3Dv_sw2eg8feUJlx/products/wandb/weave/_media/simple_llamaindex.png?fit=max&auto=format&n=3Dv_sw2eg8feUJlx&q=85&s=07e1b9836baeb21cbf8fc94e08be35b8" alt="simple_llamaindex.png" width="3340" height="1866" data-path="products/wandb/weave/_media/simple_llamaindex.png" />](https://forge.coreweave.com/wandb/wandbot/test-llamaindex-weave/weave/calls/b6b5d898-2df8-4e14-b553-66ce84661e74)

<h2 id="traces">
  トレース
</h2>

このセクションでは、RAG パイプラインのような複数ステップで構成される LlamaIndex のワークフローを、Weave がどのように取得するかを説明します。

LlamaIndex は、データと LLM を簡単に接続できることで知られています。基本的な RAG アプリケーションには、埋め込みステップ、取得ステップ、応答合成ステップが必要です。複雑になるにつれて、開発時にも本番環境でも、個々のステップのトレースを一元化されたデータベースに保存することが重要になります。

これらのトレースは、アプリケーションのデバッグや改善に欠かせません。Weave は、プロンプトテンプレート、LLM Call、ツール、エージェントのステップなど、LlamaIndex ライブラリを介して行われるすべての Call を自動的にトラッキングします。トレースは Weave の Web インターフェースで確認できます。

次の例は、LlamaIndex の [Starter Tutorial (OpenAI)](https://docs.llamaindex.ai/en/stable/getting_started/starter_example/) に掲載されている基本的な RAG パイプラインです。

```python lines {5} theme={"system"}
import weave
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader

# プロジェクト名を指定して Weave を初期化します
weave.init("llamaindex_demo")

# `data` ディレクトリに `.txt` ファイルがあることを前提とします
documents = SimpleDirectoryReader("data").load_data()
index = VectorStoreIndex.from_documents(documents)

query_engine = index.as_query_engine()
response = query_engine.query("What did the author do growing up?")
print(response)
```

トレースのタイムラインには、「イベント」だけでなく、実行時間、コスト、トークン数 (該当する場合) も取得されます。トレースをドリルダウンすると、各ステップの入力と出力を確認できます。

[<img src="https://mintcdn.com/coreweave-dbfa0e8d/3Dv_sw2eg8feUJlx/products/wandb/weave/_media/llamaindex_rag.png?fit=max&auto=format&n=3Dv_sw2eg8feUJlx&q=85&s=926c43d10205a3ebbc0907357949b57b" alt="llamaindex_rag.png" width="3340" height="1866" data-path="products/wandb/weave/_media/llamaindex_rag.png" />](https://forge.coreweave.com/wandb/wandbot/test-llamaindex-weave/weave/calls?filter=%7B%22traceRootsOnly%22%3Atrue%7D\&peekPath=%2Fwandbot%2Ftest-llamaindex-weave%2Fcalls%2F6ac53407-1bb7-4c38-b5a3-c302bd877a11%3Ftracetree%3D1)

<h2 id="one-click-observability">
  ワンクリック可観測性
</h2>

このセクションでは、Weave インテグレーションが LlamaIndex に組み込まれた可観測性システムとどのように連携するかを説明します。この仕組みにより、ハンドラを手動で設定する必要はありません。

LlamaIndex は、本番環境で堅牢な LLM アプリケーションを構築するための[ワンクリック可観測性](https://docs.llamaindex.ai/en/stable/module_guides/observability/)を提供しています。

Weave インテグレーションは LlamaIndex のこの機能を使用して、[`WeaveCallbackHandler()`](https://github.com/wandb/weave/blob/master/weave/integrations/llamaindex/llamaindex.py) を `llama_index.core.global_handler` に自動的に設定します。LlamaIndex と Weave を使用する場合、`weave.init([NAME_OF_PROJECT])` で Weave run を初期化するだけで済みます。

<h2 id="create-a-model-for-easier-experimentation">
  実験を効率化するために `Model` を作成する
</h2>

プロンプト、モデルの設定、推論パラメーターなど複数のコンポーネントがある場合、さまざまなユースケースに合わせてアプリケーション内の LLM を整理し、評価するのは容易ではありません。[`weave.Model`](/ja/products/wandb/weave/guides/core-types/models) を使用すると、システムプロンプトや使用するモデルといった実験の詳細を取得して整理できるため、異なる反復処理を比較しやすくなります。

次の例では、[weave/data](https://github.com/wandb/weave/tree/master/data) フォルダーにあるデータを使用して、`WeaveModel` 内に LlamaIndex のクエリエンジンを構築する方法を示します。

```python lines {16,52,61} theme={"system"}
import weave

from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
from llama_index.core.node_parser import SentenceSplitter
from llama_index.llms.openai import OpenAI
from llama_index.core import PromptTemplate


PROMPT_TEMPLATE = """
You are given with relevant information about Paul Graham. Answer the user query only based on the information provided. Don't make up stuff.

User Query: {query_str}
Context: {context_str}
Answer:
"""

class SimpleRAGPipeline(weave.Model):
    chat_llm: str = "gpt-4"
    temperature: float = 0.1
    similarity_top_k: int = 2
    chunk_size: int = 256
    chunk_overlap: int = 20
    prompt_template: str = PROMPT_TEMPLATE

    def get_llm(self):
        return OpenAI(temperature=self.temperature, model=self.chat_llm)

    def get_template(self):
        return PromptTemplate(self.prompt_template)

    def load_documents_and_chunk(self, data):
        documents = SimpleDirectoryReader(data).load_data()
        splitter = SentenceSplitter(
            chunk_size=self.chunk_size,
            chunk_overlap=self.chunk_overlap,
        )
        nodes = splitter.get_nodes_from_documents(documents)
        return nodes

    def get_query_engine(self, data):
        nodes = self.load_documents_and_chunk(data)
        index = VectorStoreIndex(nodes)

        llm = self.get_llm()
        prompt_template = self.get_template()

        return index.as_query_engine(
            similarity_top_k=self.similarity_top_k,
            llm=llm,
            text_qa_template=prompt_template,
        )

    @weave.op()
    def predict(self, query: str):
        query_engine = self.get_query_engine(
            # このデータは weave リポジトリの data/paul_graham にあります
            "data/paul_graham",
        )
        response = query_engine.query(query)
        return {"response": response.response}

weave.init("test-llamaindex-weave")

rag_pipeline = SimpleRAGPipeline()
response = rag_pipeline.predict("What did the author do growing up?")
print(response)
```

`weave.Model` を継承した `SimpleRAGPipeline` クラスには、この RAG パイプラインの重要なパラメーターがまとめられています。`query` メソッドを `weave.op()` でデコレートすると、トレースが有効になります。この構成にしておくことで、Weave 上で RAG パイプラインのさまざまな設定をバージョン管理、比較、評価できるようになります。

[<img src="https://mintcdn.com/coreweave-dbfa0e8d/3Dv_sw2eg8feUJlx/products/wandb/weave/_media/llamaindex_model.png?fit=max&auto=format&n=3Dv_sw2eg8feUJlx&q=85&s=2274e7eca16da733c8c632f50aa2604a" alt="llamaindex_model.png" width="3340" height="1866" data-path="products/wandb/weave/_media/llamaindex_model.png" />](https://forge.coreweave.com/wandb/wandbot/test-llamaindex-weave/weave/calls?filter=%7B%22traceRootsOnly%22%3Atrue%7D\&peekPath=%2Fwandbot%2Ftest-llamaindex-weave%2Fcalls%2Fa82afbf4-29a5-43cd-8c51-603350abeafd%3Ftracetree%3D1)

<h2 id="evaluate-with-weaveevaluation">
  `weave.Evaluation` で評価する
</h2>

このセクションでは、固定のデータセットでモデルのパフォーマンスを測定し、反復処理ごとの結果を定量的に比較する方法を説明します。

評価を使用すると、アプリケーションのパフォーマンスを測定できます。[`weave.Evaluation`](/ja/products/wandb/weave/guides/core-types/evaluations) クラスを使用すると、特定のタスクやデータセットに対するモデルのパフォーマンスを取得できるため、異なるモデルやアプリケーションの反復処理を比較できます。次の例では、前のセクションで作成したモデルを評価する方法を示します。

```python lines {25,32,36} theme={"system"}
import asyncio
from llama_index.core.evaluation import CorrectnessEvaluator

eval_examples = [
    {
        "id": "0",
        "query": "What programming language did Paul Graham learn to teach himself AI when he was in college?",
        "ground_truth": "Paul Graham learned Lisp to teach himself AI when he was in college.",
    },
    {
        "id": "1",
        "query": "What was the name of the startup Paul Graham co-founded that was eventually acquired by Yahoo?",
        "ground_truth": "The startup Paul Graham co-founded that was eventually acquired by Yahoo was called Viaweb.",
    },
    {
        "id": "2",
        "query": "What is the capital city of France?",
        "ground_truth": "I cannot answer this question because no information was provided in the text.",
    },
]

llm_judge = OpenAI(model="gpt-4", temperature=0.0)
evaluator = CorrectnessEvaluator(llm=llm_judge)

@weave.op()
def correctness_evaluator(query: str, ground_truth: str, output: dict):
    result = evaluator.evaluate(
        query=query, reference=ground_truth, response=output["response"]
    )
    return {"correctness": float(result.score)}

evaluation = weave.Evaluation(dataset=eval_examples, scorers=[correctness_evaluator])

rag_pipeline = SimpleRAGPipeline()

asyncio.run(evaluation.evaluate(rag_pipeline))
```

この評価は、前のセクションの例をベースにしています。`weave.Evaluation` で評価するには、評価用データセット、Scorer 関数、`weave.Model` が必要です。これら 3 つの主要なコンポーネントには、次の要件があります。

* 評価サンプルの dict のキーは、Scorer 関数の引数、および `weave.Model` の `predict` メソッドの引数と一致している必要があります。
* `weave.Model` には、`predict`、`infer`、`forward` のいずれかの名前のメソッドが必要です。トレースを有効にするには、このメソッドを `weave.op()` でデコレートする必要があります。
* Scorer 関数は `weave.op()` でデコレートし、名前付き引数として `output` を受け取る必要があります。

[<img src="https://mintcdn.com/coreweave-dbfa0e8d/3Dv_sw2eg8feUJlx/products/wandb/weave/_media/llamaindex_evaluation.png?fit=max&auto=format&n=3Dv_sw2eg8feUJlx&q=85&s=bd4ee5606ef48588b968d61bce95d1d2" alt="llamaindex_evaluation.png" width="3340" height="1866" data-path="products/wandb/weave/_media/llamaindex_evaluation.png" />](https://forge.coreweave.com/wandb/wandbot/llamaindex-weave/weave/calls?filter=%7B%22opVersionRefs%22%3A%5B%22weave%3A%2F%2F%2Fwandbot%2Fllamaindex-weave%2Fop%2FEvaluation.predict_and_score%3ANmwfShfFmgAhDGLXrF6Xn02T9MIAsCXBUcifCjyKpOM%22%5D%2C%22parentId%22%3A%2233491e66-b580-47fa-9d43-0cd6f1dc572a%22%7D\&peekPath=%2Fwandbot%2Fllamaindex-weave%2Fcalls%2F33491e66-b580-47fa-9d43-0cd6f1dc572a%3Ftracetree%3D1)

Weave を LlamaIndex と統合すると、LLM アプリケーションを包括的にログおよびモニタリングできます。これにより、評価を通じたデバッグやパフォーマンスの最適化を効率よく進められます。
