Infrastructure

We own the metal.
That changes the economics.

Consultancies that rent compute pass the meter on to you. We built a platform capable of running the largest open-weight models available, which means experiments stay affordable and your workload never shares a machine.

Platform

The machine

A single-node system built specifically for large mixture-of-experts models, where system memory capacity and bandwidth matter as much as GPU memory.

Processor
AMD EPYC 9755 — 128 cores / 256 threads, Zen 5
System memory
768 GB DDR5-5600, 12-channel — measured 408 GB/s copy bandwidth
Accelerators
2 × NVIDIA RTX PRO 6000 Blackwell — 192 GB combined GPU memory
Serving storage
12.8 TB enterprise NVMe — 5.5 GB/s write, 4.4 GB/s read
Archive
37 TB redundant array for model and dataset storage
Connectivity
Dual-WAN — fibre primary with independent 5G failover on a separate carrier
Facility
Hosted with conditioned power, dedicated cooling, and redundant connectivity

Measured performance

Numbers we ran ourselves

Single-stream decode throughput, measured on this hardware. These are not vendor figures and not estimates — they are what the machine does with these models loaded.

Ornith-1.0-397B 432 GB · MoE
36.0 tokens / second
Kimi K2.7-Code 554 GB · MoE
19.8 tokens / second
GLM-5.2 436 GB · MoE
18.0 tokens / second

Why the model sizes matter. A 554 GB model does not fit in any single machine's GPU memory — not ours, and not a rack of consumer cards. Running it at usable speed requires holding the model in system memory and moving the right parts to the GPUs at the right time. That is an engineering problem, and solving it is the reason these numbers exist.

Measured concurrency

What it does under load

Single-stream figures describe one user. This is the number that matters for a team or an API: total tokens per second with many people asking at once. Ramped in-house, identical prompts, zero errors at every level.

1 concurrent requestbaseline
193 tokens / second
8 concurrentsmall team
1,155 tokens / second
16 concurrent
2,043 tokens / second
32 concurrent
3,286 tokens / second
48 concurrentsustained peak
4,422 tokens / second

Latency holds while throughput scales. At 48 simultaneous requests, 95th-percentile time to first token was 0.03 seconds — indistinguishable from a single user. Aggregate throughput rises almost linearly to that point and then turns over, which is why capacity is sold below it rather than at it. Benchmarks use uniform prompts; real traffic is burstier, and quoting the peak would be quoting a number you would not reliably get.

Operating principles

How the platform is run

Tenancy

Single-tenant by allocation

Hosted workloads are assigned dedicated capacity rather than sharing a pool. Your throughput does not change because of what someone else is running.

Data

No prompt or output retention

Inference traffic is not logged to disk, retained, or used for any downstream purpose. Where you need an audit trail, it is built to your specification and stays under your control.

Licensing

Commercially licensed weights only

We deploy only models whose licences permit commercial use, and we confirm the licence terms for your specific use case before anything goes into production.

Need a model benchmarked on your workload?

Send us the task. We will measure it on this hardware and show you the methodology alongside the result.

Get in touch →