Model catalog
| Model | Model ID (for API usage) | Type | Context Window | Parameters | Description |
|---|---|---|---|---|---|
| Google Gemma 4 26B A4B Instruct | google/gemma-4-26B-A4B-it | Text, Vision | 262k | 4B-26B (Active-Total) | Gemma 4 26B A4B is a multimodal MoE model with LoRA support and function calling for agentic workflows. |
| Meta Llama 3.1 8B | meta-llama/Llama-3.1-8B-Instruct | Text | 131k | 8B (Total) | Efficient conversational model optimized for responsive multilingual chatbot interactions. |
| NVIDIA Nemotron 3.5 Lightning | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B | Text | 262k | 3B-30B (Active-Total) | Nemotron 3.5 Lightning is an MoE model built for fast, reliable agentic tasks across use cases such as financial services, cybersecurity, telecom, and retail. |
| OpenPipe Qwen3 14B Instruct | OpenPipe/Qwen3-14B-Instruct | Text | 32.8k | 14.8B (Total) | An efficient multilingual, dense, instruction-tuned model, optimized by OpenPipe for building agents with finetuning. |
| Qwen3.6 27B | Qwen/Qwen3.6-27B | Text, Vision | 262k | 27B (Total) | Qwen3.6-27B is a 27B dense multimodal model with 262K context built for flagship-level agentic coding. |
| Qwen3 30B A3B | Qwen/Qwen3-30B-A3B-Instruct-2507 | Text | 262k | 3.3B-30.5B (Active-Total) | Qwen3-30B-A3B-Instruct-2507 is a 30.5B MoE instruction-tuned model with enhanced reasoning, coding, and long-context understanding. |