Frontier scale
240–660 GB mixture-of-experts models. Experts held in 768 GB of system memory, attention on the GPUs. This is the tier that needs purpose-built hardware.
Model catalogue
Open-weight models we can deploy on dedicated hardware — including mixture-of-experts models over half a terabyte in size. Every model here runs privately, with no prompt or output retention.
240–660 GB mixture-of-experts models. Experts held in 768 GB of system memory, attention on the GPUs. This is the tier that needs purpose-built hardware.
90–190 GB models living entirely in GPU memory across both cards. No system memory in the path, so throughput is substantially higher.
Up to 90 GB on one card, leaving the second free. The right tier for latency-sensitive work and for running several models side by side.
Tier one
Throughput is single-stream decode. Figures tagged measured were benchmarked on our hardware; the rest are engineering estimates from the same platform and are labelled as such. Context is the window we actually serve at that speed, not the model's theoretical maximum — several are offered in both a faster and a longer-window configuration.
| Model | Size | Active | Context | Throughput | Best for |
|---|---|---|---|---|---|
| Ornith-1.0-397BMIT licence | 432 GB | 17B | 256K | 36.0 tok/s measured | Coding, agentic work — our speed leader |
| Kimi K2.7-CodeReasoning model | 554 GB | 32B | 256K | 22.3 tok/s measured | Highest-quality code generation |
| GLM-5.2 | 436 GB | 40B | 256K · 512K | 20.0 tok/s measured | General reasoning, agentic and SWE tasks |
| GLM-5.3Q6 — highest quality here | 618 GB | 40B | 128K · 512K | 16.2 tok/s measured | Our best general model when quality outranks speed |
| Kimi K32.8 trillion parameters | 662 GB | 50B | 1M | 9.7 tok/s measured | A one-million-token window — whole repositories and case files in a single prompt |
| Qwen3.5-397BApache-2.0 | 399 GB | 17B | 128K | 30.7 tok/s measured | Multilingual, long context |
| InklingApache-2.0 · multimodal | 587 GB | 41B | 128K | estimate on request | Vision plus language at frontier scale |
| MiMo-V2-FlashMIT licence | 328 GB | 15B | 128K | estimate on request | Fast general-purpose reasoning |
| Hy3Apache-2.0 | 318 GB | 21B | 128K | estimate on request | General assistant |
| MiniMax-M3 | 247 GB | 23B | 128K | 17–24 tok/s est. | Fastest of the large hybrids |
Tier two
Fully GPU-resident across both cards. System memory is out of the path entirely, so these are markedly faster than the frontier tier and better suited to interactive and concurrent use.
| Model | Size | Type | Throughput | Best for |
|---|---|---|---|---|
| Meditron3-70B | 132 GB | Text | on request | Medical and clinical domains |
| Qwen3-VL-235BApache-2.0 | 117 GB | Vision | on request | Flagship multimodal — documents, images, charts |
| MiniMax-M2.7 | 101 GB | Text | on request | General assistant |
| GLM-4.6VMIT licence | 98 GB | Vision | on request | Image and document understanding |
Tier three
Fast, responsive, and cheap to run continuously. Several can be hosted simultaneously, or paired with a frontier model so a quick responder is always available.
| Model | Size | Type | Throughput | Best for |
|---|---|---|---|---|
| Qwen3.5-122BApache-2.0 · NVFP4 | 79 GB | Text | 141.4 tok/s measured | Fast daily driver, strong general quality |
| Mistral-Medium-3.5 | 74 GB | Text | 14.6 tok/s measured | Balanced general assistant |
| InternVL3.5-38B | 72 GB | Vision | on request | Detailed image analysis |
| Laguna-S-2.1OpenMDW · NVFP4 | 71 GB | Code | 103.2 tok/s measured | Fast agentic coding, always-on |
| Holo-3.1-35B | 66 GB | Text | on request | Agentic workflows |
| Qwen-AgentWorld-35B | 65 GB | Simulation | on request | Language world model — predicts how an environment responds to an agent’s actions, for testing agents before they touch production |
| Qwen3-VL-32B | 63 GB | Vision | on request | Multimodal at lower cost |
| XiYanSQL-32B | 62 GB | Text-to-SQL | on request | Natural-language database querying |
| Gemma-4-31B | 59 GB | Text | on request | Efficient general assistant |
| Tongyi-DeepResearch-30B | 57 GB | Research | on request | Multi-step research and synthesis |
| Qwen3.6-35B | 51 GB | Text | 210.9 tok/s measured | Fast, low-cost general use |
| UI-TARS-1.5-7B | 31 GB | GUI agent | on request | Screen understanding and UI automation |
| Hermes-4-14B | 28 GB | Text | on request | Lightweight assistant |
| GLM-4.7-FlashNVFP4 | 18 GB | Agent | 153.4 tok/s measured | Best small agentic model — very fast |
Supporting models
The components most production systems need alongside a language model. These run continuously as sidecars without meaningfully affecting the main model's throughput.
| Model | Role | Best for |
|---|---|---|
| Qwen3-Embedding-8B | Embeddings | Semantic search and retrieval indexes |
| Qwen3-Reranker-4B | Reranking | Improving retrieval precision |
| Nemotron-3-Embed-1B | Embeddings | High-volume, low-cost indexing |
| Voxtral-Mini-4BApache-2.0 | Speech to text | Real-time transcription |
| Nemotron-3-Nano-Omni | Omni-input | Mixed text, image, and audio input |
| dots.ocr | Document OCR | Extracting structure from scans and PDFs |
On licensing. Every model listed here carries a licence permitting commercial use. We hold others in our archive that we deliberately do not offer — some carry non-commercial terms, and some originate from sanctioned entities. We confirm licence terms against your specific use case before anything reaches production, and we will tell you when a model you have asked for is one we will not deploy.
Send us the task. We will benchmark the candidates on your data and show you the methodology with the result.