> ## Documentation Index
> Fetch the complete documentation index at: https://docs.coreweave.com/llms.txt
> Use this file to discover all available pages before exploring further.

# ストリーミング応答を有効にする

> Serverless Inference でストリーミング出力を有効にし、モデルの応答を段階的に受け取ります。

`stream` オプションを `true` に設定すると、モデルの応答がチャンクのストリームとして段階的に返されます。そのため、応答全体の完了を待たずに、届いた結果から順に表示できます。これは、モデルの出力生成に時間がかかる場合に便利です。

ホストされているすべてのモデルがストリーミング出力をサポートしています。[推論モデル](/ja/products/inference/serverless/response-settings/reasoning)ではストリーミングの使用を推奨します。非ストリーミングのリクエストでは、モデルが出力を開始するまでに時間がかかるとタイムアウトする可能性があるためです。

次のサンプルでは、チャット補完リクエストでストリーミングを有効にしています。

<Tabs>
  <Tab title="Python">
    ```python theme={"system"}
    import openai

    client = openai.OpenAI(
        base_url='https://api.inference.wandb.ai/v1',
        api_key="[YOUR-API-KEY]",  # https://forge.coreweave.com/settings でAPIキーを作成します
    )

    stream = client.chat.completions.create(
        model="openai/gpt-oss-120b",
        messages=[
            {"role": "user", "content": "Tell me a rambling joke"}
        ],
        stream=True,
    )

    for chunk in stream:
        if chunk.choices:
            print(chunk.choices[0].delta.content or "", end="", flush=True)
        else:
            print(chunk) # CompletionUsage オブジェクトを表示します
    ```
  </Tab>

  <Tab title="Bash">
    ```bash theme={"system"}
    curl https://api.inference.wandb.ai/v1/chat/completions \
      -H "Content-Type: application/json" \
      -H "Authorization: Bearer [YOUR-API-KEY]" \
      -d '{
        "model": "openai/gpt-oss-120b",
        "messages": [
          { "role": "user", "content": "Tell me a rambling joke" }
        ],
        "stream": true
      }'
    ```
  </Tab>
</Tabs>
