Frontier scale
240–590 GB mixture-of-experts models. Experts held in 768 GB of system memory, attention on the GPUs. This is the tier that needs purpose-built hardware.
Model catalogue
Open-weight models we can deploy on dedicated hardware — including mixture-of-experts models over half a terabyte in size. Every model here runs privately, with no prompt or output retention.
240–590 GB mixture-of-experts models. Experts held in 768 GB of system memory, attention on the GPUs. This is the tier that needs purpose-built hardware.
90–190 GB models living entirely in GPU memory across both cards. No system memory in the path, so throughput is substantially higher.
Up to 90 GB on one card, leaving the second free. The right tier for latency-sensitive work and for running several models side by side.
Tier one
Throughput is single-stream decode. Figures tagged measured were benchmarked on our hardware; the rest are engineering estimates from the same platform and are labelled as such.
| Model | Size | Active | Throughput | Best for |
|---|---|---|---|---|
| Ornith-1.0-397BMIT licence | 432 GB | 17B | 36.1 tok/s measured | Coding, agentic work — our speed leader |
| Kimi K2.7-CodeReasoning model | 554 GB | 32B | 18.96 tok/s measured | Highest-quality code generation |
| GLM-5.2 | 436 GB | 40B | 18.7 tok/s measured | General reasoning, agentic and SWE tasks |
| InklingApache-2.0 · multimodal | 587 GB | 41B | estimate on request | Vision plus language at frontier scale |
| Kimi K2.6 | 544 GB | 32B | 15–19 tok/s est. | General assistant, thinking and instant modes |
| DeepSeek-V3.2 | 535 GB | 37B | 9–13 tok/s est. | Reasoning and long-context analysis |
| Mistral-Large-3 | 537 GB | — | 6–10 tok/s est. | European-origin general model |
| Qwen3.5-397BApache-2.0 | 399 GB | 17B | 11–16 tok/s est. | Multilingual, long context |
| MiMo-V2-FlashMIT licence | 328 GB | 15B | estimate on request | Fast general-purpose reasoning |
| Hy3Apache-2.0 | 318 GB | 21B | estimate on request | General assistant |
| MiniMax-M3 | 247 GB | 23B | 17–24 tok/s est. | Fastest of the large hybrids |
| Nemotron-3-Ultra550B parameters | 554 GB | 55B | 3.7–5.3 tok/s est. | Batch and offline work where quality outranks speed |
Tier two
Fully GPU-resident across both cards. System memory is out of the path entirely, so these are markedly faster than the frontier tier and better suited to interactive and concurrent use.
| Model | Size | Type | Best for |
|---|---|---|---|
| DeepSeek-V4-Flash | 149 GB | Text | Very long context — up to 1M tokens |
| Meditron3-70B | 132 GB | Text | Medical and clinical domains |
| Qwen3-VL-235BApache-2.0 | 117 GB | Vision | Flagship multimodal — documents, images, charts |
| MiniMax-M2.7 | 101 GB | Text | General assistant |
| GLM-4.6VMIT licence | 98 GB | Vision | Image and document understanding |
Tier three
Fast, responsive, and cheap to run continuously. Several can be hosted simultaneously, or paired with a frontier model so a quick responder is always available.
| Model | Size | Type | Best for |
|---|---|---|---|
| Qwen3.5-122BApache-2.0 · NVFP4 | 79 GB | Text | Fast daily driver, strong general quality |
| Mistral-Medium-3.5 | 74 GB | Text | Balanced general assistant |
| InternVL3.5-38B | 72 GB | Vision | Detailed image analysis |
| Laguna-S-2.1OpenMDW · NVFP4 | 71 GB | Code | Fast agentic coding, always-on |
| Holo-3.1-35B | 66 GB | Text | Agentic workflows |
| Qwen-AgentWorld-35B | 65 GB | Agent | Tool use and multi-step automation |
| Qwen3-VL-32B | 63 GB | Vision | Multimodal at lower cost |
| XiYanSQL-32B | 62 GB | Text-to-SQL | Natural-language database querying |
| Gemma-4-31B | 59 GB | Text | Efficient general assistant |
| Tongyi-DeepResearch-30B | 57 GB | Research | Multi-step research and synthesis |
| Qwen3.6-27B | 51 GB | Text | Fast, low-cost general use |
| UI-TARS-1.5-7B | 31 GB | GUI agent | Screen understanding and UI automation |
| Hermes-4-14B | 28 GB | Text | Lightweight assistant |
| GLM-4.7-FlashNVFP4 | 18 GB | Agent | Best small agentic model — very fast |
Supporting models
The components most production systems need alongside a language model. These run continuously as sidecars without meaningfully affecting the main model's throughput.
| Model | Role | Best for |
|---|---|---|
| Qwen3-Embedding-8B | Embeddings | Semantic search and retrieval indexes |
| Qwen3-Reranker-4B | Reranking | Improving retrieval precision |
| Nemotron-3-Embed-1B | Embeddings | High-volume, low-cost indexing |
| Voxtral-Mini-4BApache-2.0 | Speech to text | Real-time transcription |
| Nemotron-3-Nano-Omni | Omni-input | Mixed text, image, and audio input |
| dots.ocr | Document OCR | Extracting structure from scans and PDFs |
On licensing. Every model listed here carries a licence permitting commercial use. We hold others in our archive that we deliberately do not offer — some carry non-commercial terms, and some originate from sanctioned entities. We confirm licence terms against your specific use case before anything reaches production, and we will tell you when a model you have asked for is one we will not deploy.
Send us the task. We will benchmark the candidates on your data and show you the methodology with the result.