Performance

Real-world benchmarks.

Measured on production hardware. No synthetic scores. No cloud acceleration.

Consumer-grade GPU cluster · Up to 4× RTX 4090

Model Inference — Tokens / Second

Llama 3 8B (Q4)142 tok/s
Llama 3 70B (Q4)38 tok/s
Mistral 7B (Q4)161 tok/s
CodeLlama 13B (Q4)89 tok/s
Phi-3 Mini (Q4)184 tok/s

Measured with Ollama · Q4_K_M quantization · single concurrent request

Power Efficiency

85W

Idle Draw

Full system at rest

320W

Inference Load

Single model active

580W

Peak Load

All GPUs saturated

0.44

Tokens / Watt

Llama 3 8B Q4

Measured at wall with a kill-a-watt meter · ambient 22°C

Thermal Performance

GPU Idle Temp38°C
GPU Load Temp71°C
CPU Temp52°C
Throttle Threshold95°C

Ambient 22°C · sustained 30-min inference load · custom fan curve applied

Benchmarks represent real-world measurements on reference configurations. Actual results may vary by workload, quantization, and system configuration.

Ware.Systems

Private AI workstations built for professionals who refuse to compromise on security, performance, or sovereignty.

Our commitment to transparency and data protection

© 2026 Ware.Systems. All rights reserved. Private Property for Superintelligence