Model catalogue

Models most providers
cannot run.

Open-weight models we can deploy on dedicated hardware — including mixture-of-experts models over half a terabyte in size. Every model here runs privately, with no prompt or output retention.

Tier one

Frontier scale

240–660 GB mixture-of-experts models. Experts held in 768 GB of system memory, attention on the GPUs. This is the tier that needs purpose-built hardware.

Tier two

Multi-GPU resident

90–190 GB models living entirely in GPU memory across both cards. No system memory in the path, so throughput is substantially higher.

Tier three

Single-GPU

Up to 90 GB on one card, leaving the second free. The right tier for latency-sensitive work and for running several models side by side.

Tier one

Frontier-scale models

Throughput is single-stream decode. Figures tagged measured were benchmarked on our hardware; the rest are engineering estimates from the same platform and are labelled as such. Context is the window we actually serve at that speed, not the model's theoretical maximum — several are offered in both a faster and a longer-window configuration.

ModelSizeActiveContextThroughputBest for
Ornith-1.0-397BMIT licence 432 GB17B 256K 36.0 tok/s measured Coding, agentic work — our speed leader
Kimi K2.7-CodeReasoning model 554 GB32B 256K 22.3 tok/s measured Highest-quality code generation
GLM-5.2 436 GB40B 256K · 512K 20.0 tok/s measured General reasoning, agentic and SWE tasks
GLM-5.3Q6 — highest quality here 618 GB40B 128K · 512K 16.2 tok/s measured Our best general model when quality outranks speed
Kimi K32.8 trillion parameters 662 GB50B 1M 9.7 tok/s measured A one-million-token window — whole repositories and case files in a single prompt
Qwen3.5-397BApache-2.0 399 GB17B 128K 30.7 tok/s measured Multilingual, long context
InklingApache-2.0 · multimodal 587 GB41B 128K estimate on request Vision plus language at frontier scale
MiMo-V2-FlashMIT licence 328 GB15B 128K estimate on request Fast general-purpose reasoning
Hy3Apache-2.0 318 GB21B 128K estimate on request General assistant
MiniMax-M3 247 GB23B 128K 17–24 tok/s est. Fastest of the large hybrids

Tier two

Multi-GPU resident

Fully GPU-resident across both cards. System memory is out of the path entirely, so these are markedly faster than the frontier tier and better suited to interactive and concurrent use.

ModelSizeTypeThroughputBest for
Meditron3-70B132 GBTexton requestMedical and clinical domains
Qwen3-VL-235BApache-2.0117 GBVisionon requestFlagship multimodal — documents, images, charts
MiniMax-M2.7101 GBTexton requestGeneral assistant
GLM-4.6VMIT licence98 GBVisionon requestImage and document understanding

Tier three

Single-GPU models

Fast, responsive, and cheap to run continuously. Several can be hosted simultaneously, or paired with a frontier model so a quick responder is always available.

ModelSizeTypeThroughputBest for
Qwen3.5-122BApache-2.0 · NVFP479 GBText141.4 tok/s measuredFast daily driver, strong general quality
Mistral-Medium-3.574 GBText14.6 tok/s measuredBalanced general assistant
InternVL3.5-38B72 GBVisionon requestDetailed image analysis
Laguna-S-2.1OpenMDW · NVFP471 GBCode103.2 tok/s measuredFast agentic coding, always-on
Holo-3.1-35B66 GBTexton requestAgentic workflows
Qwen-AgentWorld-35B65 GBSimulationon requestLanguage world model — predicts how an environment responds to an agent’s actions, for testing agents before they touch production
Qwen3-VL-32B63 GBVisionon requestMultimodal at lower cost
XiYanSQL-32B62 GBText-to-SQLon requestNatural-language database querying
Gemma-4-31B59 GBTexton requestEfficient general assistant
Tongyi-DeepResearch-30B57 GBResearchon requestMulti-step research and synthesis
Qwen3.6-35B51 GBText210.9 tok/s measuredFast, low-cost general use
UI-TARS-1.5-7B31 GBGUI agenton requestScreen understanding and UI automation
Hermes-4-14B28 GBTexton requestLightweight assistant
GLM-4.7-FlashNVFP418 GBAgent153.4 tok/s measuredBest small agentic model — very fast

Supporting models

Retrieval, speech, and document processing

The components most production systems need alongside a language model. These run continuously as sidecars without meaningfully affecting the main model's throughput.

ModelRoleBest for
Qwen3-Embedding-8BEmbeddingsSemantic search and retrieval indexes
Qwen3-Reranker-4BRerankingImproving retrieval precision
Nemotron-3-Embed-1BEmbeddingsHigh-volume, low-cost indexing
Voxtral-Mini-4BApache-2.0Speech to textReal-time transcription
Nemotron-3-Nano-OmniOmni-inputMixed text, image, and audio input
dots.ocrDocument OCRExtracting structure from scans and PDFs

On licensing. Every model listed here carries a licence permitting commercial use. We hold others in our archive that we deliberately do not offer — some carry non-commercial terms, and some originate from sanctioned entities. We confirm licence terms against your specific use case before anything reaches production, and we will tell you when a model you have asked for is one we will not deploy.

Not sure which model fits?

Send us the task. We will benchmark the candidates on your data and show you the methodology with the result.

Get in touch →