> ## Documentation Index
> Fetch the complete documentation index at: https://docs.coreweave.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Available models

> Browse the foundation models available for training with Serverless Training, including supported model variants and capabilities.

Serverless Training currently supports the following foundation models for both RL and SFT training.

To express interest in a particular model that isn't listed below, contact [support](mailto:forge-support@coreweave.com).

## Model catalog

| Model | Model ID (for API usage) | Type | Context Window | Parameters | Description |
| - | - | - | - | - | - |
| Google Gemma 4 26B A4B Instruct | `google/gemma-4-26B-A4B-it` | Text, Vision | 262k | 4B-26B (Active-Total) | Gemma 4 26B A4B is a multimodal MoE model with LoRA support and function calling for agentic workflows. |
| Meta Llama 3.1 8B | `meta-llama/Llama-3.1-8B-Instruct` | Text | 131k | 8B (Total) | Efficient conversational model optimized for responsive multilingual chatbot interactions. |
| NVIDIA Nemotron 3.5 Lightning | `nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B` | Text | 262k | 3B-30B (Active-Total) | Nemotron 3.5 Lightning is an MoE model built for fast, reliable agentic tasks across use cases such as financial services, cybersecurity, telecom, and retail. |
| OpenPipe Qwen3 14B Instruct | `OpenPipe/Qwen3-14B-Instruct` | Text | 32.8k | 14.8B (Total) | An efficient multilingual, dense, instruction-tuned model, optimized by OpenPipe for building agents with finetuning. |
| Qwen3.6 27B | `Qwen/Qwen3.6-27B` | Text, Vision | 262k | 27B (Total) | Qwen3.6-27B is a 27B dense multimodal model with 262K context built for flagship-level agentic coding. |
| Qwen3 30B A3B | `Qwen/Qwen3-30B-A3B-Instruct-2507` | Text | 262k | 3.3B-30.5B (Active-Total) | Qwen3-30B-A3B-Instruct-2507 is a 30.5B MoE instruction-tuned model with enhanced reasoning, coding, and long-context understanding. |
