> ## Documentation Index
> Fetch the complete documentation index at: https://docs.coreweave.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Available models

> Browse the foundation models available through Serverless Inference


Serverless Inference provides access to several open source foundation models. Each model has different strengths and use cases.

## Generally available models

The following models are [generally available](/products/inference/serverless/lifecycle#model-lifecycle-stages):

| Model | Model ID (for API usage) | Type | Context Window | Parameters | Description |
| - | - | - | - | - | - |
| DeepSeek V4.1-Flash | `deepseek-ai/DeepSeek-V4.1-Flash` | Text, Vision | 1049k | 16B-552B (Active-Total) | DeepSeek V4.1 Flash is a multimodal MoE model for coding, reasoning, and agentic workloads with long contexts. |
| DeepSeek V4-Flash-0731 | `deepseek-ai/DeepSeek-V4-Flash-0731` | Text | 262k | 13B-284B (Active-Total) | DeepSeek V4-Flash-0731 is an MoE model great for coding, reasoning, and agentic workloads. |
| DeepSeek V4-Pro-0813 | `deepseek-ai/DeepSeek-V4-Pro-0813` | Text | 1049k | 49B-1.6T (Active-Total) | DeepSeek V4-Pro-0813 is a 1.6T-parameter MoE model excelling at advanced reasoning, coding, and complex agentic workloads. |
| DeepSeek V3.1 | `deepseek-ai/DeepSeek-V3.1` | Text | 161k | 37B-671B (Active-Total) | A large hybrid model that supports both thinking and non-thinking modes via prompt templates. |
| Google Gemma 4 31B | `google/gemma-4-31B-it` | Text, Vision | 262k | 31B (Total) | Gemma 4 31B Dense is designed for advanced reasoning, agentic workflows, and longer context and is natively trained on 140+ languages. |
| Google Gemma 4 26B A4B Instruct | `google/gemma-4-26B-A4B-it` | Text, Vision | 262k | 4B-26B (Active-Total) | Gemma 4 26B A4B is a multimodal MoE model with LoRA support and function calling for agentic workflows. |
| IBM Granite 4.2 8B | `ibm-granite/granite-4.2-8b` | Text | 131k | 8B (Total) | Granite 4.2 8B is an instruct model capable of enhanced tool calling, instruction following, and chat capabilities. |
| Meta Llama 3.3 70B | `meta-llama/Llama-3.3-70B-Instruct` | Text | 128k | 70B (Total) | Multilingual model excelling in conversational tasks, detailed instruction-following, and coding. |
| Meta Llama 3.1 8B | `meta-llama/Llama-3.1-8B-Instruct` | Text | 131k | 8B (Total) | Efficient conversational model optimized for responsive multilingual chatbot interactions. |
| MiniMax M3 | `MiniMaxAI/MiniMax-M3` | Text, Vision | 262k | 23B-428B (Active-Total) | MiniMax M3 is a multimodal MoE model with 23B active parameters optimized for coding and agentic workflows. |
| Moonshot AI Kimi K2.7 Code | `moonshotai/Kimi-K2.7-Code` | Text, Vision | 262k | 32B-1T (Active-Total) | Kimi K2.7 Code is a 1T-parameter MoE model with 32B active parameters purpose-built for long-horizon agentic coding and software engineering. |
| Moonshot AI Kimi K2.6 | `moonshotai/Kimi-K2.6` | Text, Vision | 262k | 32B-1T (Active-Total) | Kimi K2.6 is a multimodal Mixture-of-Experts language model featuring 32 billion activated parameters and a total of 1 trillion parameters. |
| NVIDIA Nemotron 3.5 Lightning | `nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B` | Text | 262k | 3B-30B (Active-Total) | Nemotron 3.5 Lightning is an MoE model built for fast, reliable agentic tasks across use cases such as financial services, cybersecurity, telecom, and retail. |
| NVIDIA Nemotron 3 Ultra | `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B` | Text | 262k | 55B-550B (Active-Total) | Nemotron 3 Ultra is a powerful MoE model designed for long-running agents across coding, deep research, and enterprise automation. |
| OpenAI GPT OSS 120B | `openai/gpt-oss-120b` | Text | 131k | 5.1B-117B (Active-Total) | Efficient Mixture-of-Experts model designed for high-reasoning, agentic and general-purpose use cases. |
| OpenAI GPT OSS 20B | `openai/gpt-oss-20b` | Text | 131k | 3.6B-20B (Active-Total) | Lower latency Mixture-of-Experts model trained on OpenAI's Harmony response format with reasoning capabilities. |
| Qwen3.8 27B | `Qwen/Qwen3.8-27B` | Text, Vision | 262k | 27B (Total) | Qwen3.8-27B is a dense multimodal model suited for coding, research, vision, and long-running agent tasks. |
| Qwen3.6 35B A3B | `Qwen/Qwen3.6-35B-A3B` | Text, Vision | 262k | 3B-35B (Active-Total) | Qwen3.6-35B-A3B is an MoE multimodal model with 262K context optimized for agentic coding workflows. |
| Z.AI GLM 5.3 Flash | `zai-org/GLM-5.3-Flash` | Text, Vision | 1049k | 18B-320B (Active-Total) | GLM-5.3-Flash is a natively multimodal model with 320B total parameters and 18B active parameters. |
| Z.AI GLM 5.2 | `zai-org/GLM-5.2` | Text | 1049k | 40B-744B (Active-Total) | GLM-5.2 is a Mixture-of-Experts language model featuring 40 billion activated parameters and a total of 744 billion parameters. |

## Experimental models

The following models are [experimental](/products/inference/serverless/lifecycle#model-lifecycle-stages):

*None currently*

## Deprecated models

The following models are [deprecated](/products/inference/serverless/lifecycle#model-lifecycle-stages):

| Model | Model ID (for API usage) | Type | Context Window | Parameters | Description |
| - | - | - | - | - | - |
| DeepSeek V4-Flash | `deepseek-ai/DeepSeek-V4-Flash` | Text | 1049k | 13B-284B (Active-Total) | DeepSeek V4-Flash is an MoE model with 1M context length great for coding, reasoning, and agentic workloads. |
| DeepSeek V4-Pro | `deepseek-ai/DeepSeek-V4-Pro` | Text | 1049k | 49B-1.6T (Active-Total) | DeepSeek V4-Pro is a 1.6T-parameter MoE model with 49B active parameters excelling at advanced reasoning, coding, and complex agentic workloads. |
| IBM Granite 4.1 8B | `ibm-granite/granite-4.1-8b` | Text | 131k | 8B (Total) | Granite 4.1 8B is a long-context instruct model capable of enhanced tool calling, instruction following, and chat capabilities. |
| JetBrains Mellum2 12B A2.5B | `JetBrains/Mellum2-12B-A2.5B-Instruct` | Text | 131k | 2.5B-12B (Active-Total) | Mellum2-12B-A2.5B-Instruct is a fast MoE model with 131K context built for coding, tool use, and low-latency AI workflows. |
| Meta Llama 3.1 70B | `meta-llama/Llama-3.1-70B-Instruct` | Text | 131k | 70B (Total) | Efficient conversational model optimized for responsive multilingual chatbot interactions. |
| OpenPipe Qwen3 14B Instruct | `OpenPipe/Qwen3-14B-Instruct` | Text | 32.8k | 14.8B (Total) | An efficient multilingual, dense, instruction-tuned model, optimized by OpenPipe for building agents with finetuning. |
| Qwen3.6 27B | `Qwen/Qwen3.6-27B` | Text, Vision | 262k | 27B (Total) | Qwen3.6-27B is a 27B dense multimodal model with 262K context built for flagship-level agentic coding. |
| Qwen3.5 35B A3B | `Qwen/Qwen3.5-35B-A3B` | Text, Vision | 262k | 3B-35B (Active-Total) | Qwen3.5-35B-A3B is an open-weights multimodal MoE model built for efficient, high-throughput inference across chat, reasoning, and agentic tasks. |
| Qwen3 30B A3B | `Qwen/Qwen3-30B-A3B-Instruct-2507` | Text | 262k | 3.3B-30.5B (Active-Total) | Qwen3-30B-A3B-Instruct-2507 is a 30.5B MoE instruction-tuned model with enhanced reasoning, coding, and long-context understanding. |

## Specify a model for inference

To specify a model when calling the API, use its `Model ID` from the preceding tables. For example:

```python theme={"system"}
response = client.chat.completions.create(
    model="meta-llama/Llama-3.1-8B-Instruct",
    messages=[...]
)
```

## Next steps

After you've chosen a model, continue with one of the following resources:

* Check [usage limits and pricing](/products/inference/serverless/usage-limits) for each model.
* See the [API reference](/products/inference/serverless/api-reference) for how to use these models.
* Try models in the [Playground](/products/inference/serverless/ui-guide).


## Related topics

- [Available models](/products/post-training/serverless-training/available-models.md)
