Performance Benchmarks
Explore Optimized AI models recipes across GPUs, frameworks,
and deployment configurations.
BENCHMARKED MODELS
REASONING
vLLM · 8× NVIDIA H100-80GB
Qwen3-235B-A22B
+138.7% token/s
+58% Cost Savings
Qwen3-235B-A22B · benchmark
REASONING
vLLM · 8× AMD Instinct MI325X
GLM-5.2
+26.8% token/s
GLM-5.2 · benchmark
MULTI-MODAL
vLLM · 1× NVIDIA H100-80GB
Gemma-4-31B-it
+66.0% token/s
+40% Cost Savings
Gemma-4-31B-it · benchmark
OPEN-WEIGHT
vLLM · 1× NVIDIA H100-80GB
GPT-OSS-120B
+24.5% token/s
+20% Cost Savings
GPT-OSS-120B · benchmark
PRODUCTION-READY PERFORMANCE
Benchmark with confidence. Deploy with clarity.
Use validated performance data to select the right model for your production workload.
Talk to Our Team
