> ## Documentation Index
> Fetch the complete documentation index at: https://docs.coreweave.com/llms.txt
> Use this file to discover all available pages before exploring further.

# 추론 정보 보기

> Serverless Inference 응답에서 추론을 반환하고 보는 방법

추론 모델(예: [Google의 Gemma 4](https://huggingface.co/google/gemma-4-31B-it))은 최종 답변과 함께 추론 단계에 대한 정보를 반환합니다. 이 페이지에서는 Serverless Inference에서 추론을 지원하는 모델을 파악하는 방법, 응답에서 추론 출력을 찾는 위치, 추론을 켜고 끌 수 있는 모델의 경우 이를 제어하는 방법을 설명합니다. 모델의 중간 추론을 검사하거나 응답에 추론이 표시되는지 제어하려면 이 가이드를 사용하세요.

모델이 추론을 지원하는지 확인하려면 지원되는 모델 table 또는 UI의 카탈로그 페이지에 있는 **Supported Features** 섹션을 확인하세요.

추론 정보는 응답의 `reasoning` 필드에 나타납니다. 이 필드의 값은 추론을 지원하지 않는 모델의 응답에서는 `null`입니다.

<h2 id="supported-models-with-reasoning">
  추론을 지원하는 모델
</h2>

다음 표는 추론 출력을 반환할 수 있는 Serverless Inference 모델과 각 모델의 동작 방식을 보여줍니다.

* **항상 켜짐:** 모델이 항상 추론 출력을 반환하며, 이 기능은 비활성화할 수 없습니다.
* **기본적으로 활성화됨** / **기본적으로 비활성화됨:** 추론 출력을 켜거나 끌 수 있습니다. 표에는 설정을 지정하지 않았을 때 적용되는 기본값이 표시됩니다.
* **적응형; 기본적으로 모델이 선택:** 모델이 요청마다 추론 출력을 반환할지 결정합니다. 이 동작은 재정의할 수 있습니다.

| Model ID (for API usage) | 추론 지원 |
| - | - |
| `deepseek-ai/DeepSeek-V4.1-Flash` | 기본적으로 활성화됨 |
| `deepseek-ai/DeepSeek-V4-Flash` | 기본적으로 비활성화됨 |
| `deepseek-ai/DeepSeek-V4-Flash-0731` | 기본적으로 활성화됨 |
| `deepseek-ai/DeepSeek-V4-Pro` | 기본적으로 비활성화됨 |
| `deepseek-ai/DeepSeek-V4-Pro-0813` | 기본적으로 활성화됨 |
| `google/gemma-4-31B-it` | 기본적으로 비활성화됨 |
| `google/gemma-4-26B-A4B-it` | 기본적으로 비활성화됨 |
| `ibm-granite/granite-4.2-8b` | 기본적으로 활성화됨 |
| `MiniMaxAI/MiniMax-M3` | 적응형; 기본적으로 모델이 선택 |
| `moonshotai/Kimi-K2.7-Code` | 항상 켜짐 |
| `moonshotai/Kimi-K2.6` | 항상 켜짐 |
| `nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B` | 기본적으로 활성화됨 |
| `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B` | 기본적으로 활성화됨 |
| `openai/gpt-oss-120b` | 항상 켜짐 |
| `openai/gpt-oss-20b` | 항상 켜짐 |
| `Qwen/Qwen3.8-27B` | 기본적으로 활성화됨 |
| `Qwen/Qwen3.6-35B-A3B` | 기본적으로 활성화됨 |
| `Qwen/Qwen3.6-27B` | 기본적으로 활성화됨 |
| `Qwen/Qwen3.5-35B-A3B` | 기본적으로 활성화됨 |
| `zai-org/GLM-5.3-Flash` | 항상 켜짐 |
| `zai-org/GLM-5.2` | 기본적으로 활성화됨 |

<h3 id="models-with-always-on-reasoning">
  `Always on` 추론이 있는 모델
</h3>

이전 [지원되는 모델](#supported-models-with-reasoning) table에 `Always on`으로 나열된 모델은 항상 추론을 포함하며, 이를 비활성화할 수 없습니다.

<h3 id="disable-reasoning">
  추론 비활성화
</h3>

앞의 [지원되는 모델](#supported-models-with-reasoning) table에 `기본적으로 활성화됨`으로 표시된 모델은 추론을 비활성화하여 토큰 사용량을 줄이거나 응답을 단순화할 수 있습니다. 요청에서 추론을 사용하지 않으려면 `chat_template_kwargs`에서 `enable_thinking` 플래그를 `False`(Python) 또는 `false`(Bash)로 설정하세요. 요청이 완료되면 응답에 추론 콘텐츠가 포함되지 않습니다.

<Tabs>
  <Tab title="Python">
    ```python lines highlight={13-17} theme={"system"}
    import openai

    client = openai.OpenAI(
        base_url='https://api.inference.wandb.ai/v1',
        api_key="[YOUR-API-KEY]",  # https://forge.coreweave.com/settings 에서 API 키를 생성하세요
    )

    response = client.chat.completions.create(
        model="google/gemma-4-31B-it",
        messages=[
            {"role": "user", "content": "3.11 and 3.8, which is greater?"}
        ],
        extra_body={
            "chat_template_kwargs": {
                "enable_thinking": False
            }
        },
    )
    ```
  </Tab>

  <Tab title="Bash">
    ```bash lines highlight={9} theme={"system"}
    curl https://api.inference.wandb.ai/v1/chat/completions \
      -H "Content-Type: application/json" \
      -H "Authorization: Bearer [YOUR-API-KEY]" \
      -d '{
        "model": "google/gemma-4-31B-it",
        "messages": [
          { "role": "user", "content": "3.11 and 3.8, which is greater?" }
        ],
        "chat_template_kwargs": {"enable_thinking": false}
      }'
    ```
  </Tab>
</Tabs>

<h3 id="enable-reasoning">
  추론 활성화
</h3>

모델이 앞의 [지원되는 모델](#supported-models-with-reasoning) table에 `기본적으로 비활성화됨`으로 나열되어 있는 경우, 앞의 코드 스니펫에서 `enable_thinking` 플래그를 `True` (Python) 또는 `true` (Bash)로 설정하여 추론을 활성화할 수 있습니다.
