Performance
Real-world benchmarks.
Measured on production hardware. No synthetic scores. No cloud acceleration.
Consumer-grade GPU cluster · Up to 4× RTX 4090
Model Inference — Tokens / Second
Llama 3 8B (Q4)142 tok/s
Llama 3 70B (Q4)38 tok/s
Mistral 7B (Q4)161 tok/s
CodeLlama 13B (Q4)89 tok/s
Phi-3 Mini (Q4)184 tok/s
Measured with Ollama · Q4_K_M quantization · single concurrent request
Power Efficiency
85W
Idle Draw
Full system at rest
320W
Inference Load
Single model active
580W
Peak Load
All GPUs saturated
0.44
Tokens / Watt
Llama 3 8B Q4
Measured at wall with a kill-a-watt meter · ambient 22°C
Thermal Performance
GPU Idle Temp38°C
GPU Load Temp71°C
CPU Temp52°C
Throttle Threshold95°C
Ambient 22°C · sustained 30-min inference load · custom fan curve applied
Benchmarks represent real-world measurements on reference configurations. Actual results may vary by workload, quantization, and system configuration.