llama.cpp Benchmark Report

Generated 2026-06-12 21:56:20 · 12 model(s) with data

System

GPUs
2 x NVIDIA RTX PRO 6000 Blackwell Workstation Edition
GPU driver
610.43.02
GPU power limit
550 W (set via nvidia-smi -pl; persistence mode on)
CPU
AMD Ryzen 9 9950X 16-Core Processor
System RAM
128 GiB (OS-visible: 123 GiB)
llama.cpp build
version: 9590 (d2462f8f7) built with GNU 14.2.0 for Linux x86_64

Benchmark configuration

GPU power limit (both GPUs)
550 W
Flash attention
on (`-fa 1`)
KV cache type (K and V)
q8_0 (`-ctk q8_0 -ctv q8_0`)
CPU threads
16 (`-t 16`)
GPU layers
all (`-ngl 999`)
Batch / micro-batch
llama-bench defaults (`-b 2048 -ub 512`)
Multi-GPU split (dual)
layer (`-sm layer`)
Repetitions
5 (`-r 5`)
Cool-down between models
300 s
nvidia-smi sampling
every 500 ms, scoped to the run's GPUs

Models under test

AliasHF repo : quantPlacementNote
devstral-small-2-24bunsloth/Devstral-Small-2507-GGUF:Q8_0gpu024B dense, Q8 ~26 GB (verify repo)
glm-4.6-355bunsloth/GLM-4.6-GGUF:UD-Q2_K_XLdual355B-A32B, Q2_K_XL ~135 GB, dual
gpt-oss-120bggml-org/gpt-oss-120b-GGUFgpu0MXFP4 native, ~60 GB, single card
hermes-4-70bunsloth/Hermes-4-70B-GGUF:Q8_0gpu070B dense (Llama 3.1), Q8 ~75 GB, single card
hermes-4.3-36bNousResearch/Hermes-4.3-36B-GGUF:Q8_0gpu036B dense (Seed-OSS-36B), Q8 ~38 GB, single card
minimax-m2.7unsloth/MiniMax-M2.7-GGUF:UD-Q4_K_XLdual~230B-A10B, Q4, dual (verify tag)
qwen2.5-coder-32bunsloth/Qwen2.5-Coder-32B-Instruct-GGUF:Q8_0gpu032B dense, Q8 ~35 GB
qwen3-235b-a22bunsloth/Qwen3-235B-A22B-Instruct-2507-GGUF:Q4_K_Mdual235B-A22B, Q4 ~133 GB, dual
qwen3-coder-480bunsloth/Qwen3-Coder-480B-A35B-Instruct-GGUF:UD-Q2_K_XLdual480B-A35B, Q2 ~180 GB, dual (tight)
qwen3-coder-next-80bunsloth/Qwen3-Coder-Next-GGUF:Q8_0gpu080B-A3B, Q8 ~85 GB, single card
qwen3.6-27bunsloth/Qwen3.6-27B-GGUF:Q8_0gpu027B dense, Q8 ~30 GB, single card
qwen3.6-35b-a3bunsloth/Qwen3.6-35B-A3B-GGUF:Q8_0gpu035B-A3B MoE, Q8 ~38 GB, single card

Prefill throughput (prompt processing)

Prefill 512 @ d0Prefill 4096 @ d00.00200040006000800010000tokens / secondqwen3.6-35b-a3bqwen3.6-35b-a3b — Prefill 512 @ d0: 99239923qwen3.6-35b-a3b — Prefill 4096 @ d0: 98199819gpt-oss-120bgpt-oss-120b — Prefill 512 @ d0: 94999499gpt-oss-120b — Prefill 4096 @ d0: 94469446devstral-small-2-24bdevstral-small-2-24b — Prefill 512 @ d0: 54805480devstral-small-2-24b — Prefill 4096 @ d0: 54285428qwen3-coder-next-80bqwen3-coder-next-80b — Prefill 512 @ d0: 48654865qwen3-coder-next-80b — Prefill 4096 @ d0: 48904890qwen3.6-27bqwen3.6-27b — Prefill 512 @ d0: 41024102qwen3.6-27b — Prefill 4096 @ d0: 41044104qwen2.5-coder-32bqwen2.5-coder-32b — Prefill 512 @ d0: 38383838qwen2.5-coder-32b — Prefill 4096 @ d0: 36633663minimax-m2.7minimax-m2.7 — Prefill 512 @ d0: 18251825minimax-m2.7 — Prefill 4096 @ d0: 36913691hermes-4.3-36bhermes-4.3-36b — Prefill 512 @ d0: 34703470hermes-4.3-36b — Prefill 4096 @ d0: 31163116qwen3-235b-a22bqwen3-235b-a22b — Prefill 512 @ d0: 19821982qwen3-235b-a22b — Prefill 4096 @ d0: 26022602hermes-4-70bhermes-4-70b — Prefill 512 @ d0: 18961896hermes-4-70b — Prefill 4096 @ d0: 18061806qwen3-coder-480bqwen3-coder-480b — Prefill 512 @ d0: 907907qwen3-coder-480b — Prefill 4096 @ d0: 14081408glm-4.6-355bglm-4.6-355b — Prefill 512 @ d0: 937937glm-4.6-355b — Prefill 4096 @ d0: 11471147

Decode throughput (token generation)

Decode 128 @ d0Decode 128 @ d163840.0050.0100150200250300tokens / secondgpt-oss-120bgpt-oss-120b — Decode 128 @ d0: 275275gpt-oss-120b — Decode 128 @ d16384: 235235qwen3.6-35b-a3bqwen3.6-35b-a3b — Decode 128 @ d0: 239239qwen3.6-35b-a3b — Decode 128 @ d16384: 225225qwen3-coder-next-80bqwen3-coder-next-80b — Decode 128 @ d0: 195195qwen3-coder-next-80b — Decode 128 @ d16384: 184184minimax-m2.7minimax-m2.7 — Decode 128 @ d0: 129129minimax-m2.7 — Decode 128 @ d16384: 89.789.7qwen3-235b-a22bqwen3-235b-a22b — Decode 128 @ d0: 83.983.9qwen3-235b-a22b — Decode 128 @ d16384: 53.653.6qwen3-coder-480bqwen3-coder-480b — Decode 128 @ d0: 73.473.4qwen3-coder-480b — Decode 128 @ d16384: 48.948.9devstral-small-2-24bdevstral-small-2-24b — Decode 128 @ d0: 61.461.4devstral-small-2-24b — Decode 128 @ d16384: 56.256.2glm-4.6-355bglm-4.6-355b — Decode 128 @ d0: 61.161.1glm-4.6-355b — Decode 128 @ d16384: 37.837.8qwen3.6-27bqwen3.6-27b — Decode 128 @ d0: 50.550.5qwen3.6-27b — Decode 128 @ d16384: 48.948.9qwen2.5-coder-32bqwen2.5-coder-32b — Decode 128 @ d0: 43.443.4qwen2.5-coder-32b — Decode 128 @ d16384: 38.638.6hermes-4.3-36bhermes-4.3-36b — Decode 128 @ d0: 39.539.5hermes-4.3-36b — Decode 128 @ d16384: 32.132.1hermes-4-70bhermes-4-70b — Decode 128 @ d0: 20.820.8hermes-4-70b — Decode 128 @ d16384: 18.418.4

Peak temperatures and power

Maxima are taken over the active benchmark window of each run; only GPUs that actively participated are shown, so single-card runs omit the idle second card. Peak power samples a few percent above the configured limit are normal: nvidia-smi reports instantaneous draw while NVIDIA's power controller regulates a time-averaged budget, so brief transients above the cap are expected and not a fault. Compare the average power chart against the limit instead.

Max core temperature

GPU 0GPU 10.0020.040.060.080.0degrees Cdevstral-small-2-24bdevstral-small-2-24b — GPU 0: 64.064.0glm-4.6-355bglm-4.6-355b — GPU 0: 77.077.0glm-4.6-355b — GPU 1: 70.070.0gpt-oss-120bgpt-oss-120b — GPU 0: 56.056.0hermes-4-70bhermes-4-70b — GPU 0: 77.077.0hermes-4.3-36bhermes-4.3-36b — GPU 0: 73.073.0minimax-m2.7minimax-m2.7 — GPU 0: 67.067.0minimax-m2.7 — GPU 1: 64.064.0qwen2.5-coder-32bqwen2.5-coder-32b — GPU 0: 69.069.0qwen3-235b-a22bqwen3-235b-a22b — GPU 0: 69.069.0qwen3-235b-a22b — GPU 1: 64.064.0qwen3-coder-480bqwen3-coder-480b — GPU 0: 78.078.0qwen3-coder-480b — GPU 1: 68.068.0qwen3-coder-next-80bqwen3-coder-next-80b — GPU 0: 56.056.0qwen3.6-27bqwen3.6-27b — GPU 0: 68.068.0qwen3.6-35b-a3bqwen3.6-35b-a3b — GPU 0: 56.056.0

Max power draw (instantaneous samples)

GPU 0GPU 10.00100200300400500600wattsdevstral-small-2-24bdevstral-small-2-24b — GPU 0: 556556glm-4.6-355bglm-4.6-355b — GPU 0: 462462glm-4.6-355b — GPU 1: 461461gpt-oss-120bgpt-oss-120b — GPU 0: 553553hermes-4-70bhermes-4-70b — GPU 0: 562562hermes-4.3-36bhermes-4.3-36b — GPU 0: 559559minimax-m2.7minimax-m2.7 — GPU 0: 554554minimax-m2.7 — GPU 1: 531531qwen2.5-coder-32bqwen2.5-coder-32b — GPU 0: 559559qwen3-235b-a22bqwen3-235b-a22b — GPU 0: 478478qwen3-235b-a22b — GPU 1: 438438qwen3-coder-480bqwen3-coder-480b — GPU 0: 533533qwen3-coder-480b — GPU 1: 441441qwen3-coder-next-80bqwen3-coder-next-80b — GPU 0: 464464qwen3.6-27bqwen3.6-27b — GPU 0: 558558qwen3.6-35b-a3bqwen3.6-35b-a3b — GPU 0: 542542

Average power draw (over the active window)

GPU 0GPU 10.00100200300400500wattsdevstral-small-2-24bdevstral-small-2-24b — GPU 0: 441441glm-4.6-355bglm-4.6-355b — GPU 0: 356356glm-4.6-355b — GPU 1: 361361gpt-oss-120bgpt-oss-120b — GPU 0: 351351hermes-4-70bhermes-4-70b — GPU 0: 485485hermes-4.3-36bhermes-4.3-36b — GPU 0: 480480minimax-m2.7minimax-m2.7 — GPU 0: 138138minimax-m2.7 — GPU 1: 172172qwen2.5-coder-32bqwen2.5-coder-32b — GPU 0: 465465qwen3-235b-a22bqwen3-235b-a22b — GPU 0: 358358qwen3-235b-a22b — GPU 1: 343343qwen3-coder-480bqwen3-coder-480b — GPU 0: 417417qwen3-coder-480b — GPU 1: 353353qwen3-coder-next-80bqwen3-coder-next-80b — GPU 0: 295295qwen3.6-27bqwen3.6-27b — GPU 0: 449449qwen3.6-35b-a3bqwen3.6-35b-a3b — GPU 0: 318318

Per-model thermals over time

Charts are trimmed to the active benchmark window: the model-loading phase and idle tails are cut (the amount removed is noted per model), so the measured runs fill the chart instead of being squeezed by minutes of GGUF loading. SM clock dropping during a run is the clearest sign of throttling. Power-cap limited is expected whenever the power limit is set below the card's maximum — the card is staying within budget, not faulting. Thermal / HW slowdown means the card hit a temperature or hardware limit and reduced clocks; investigate cooling and airflow.

devstral-small-2-24bGPU 0power-cap limited (expected at your set power limit)

-hf unsloth/Devstral-Small-2507-GGUF:Q8_0 | placement: gpu0

showing the 43s active window; 72s of model-load/idle trimmed from a 115s recording

Temperature (°C)
0.0010.020.030.040.050.060.070.00.0010.020.030.040.050.0°Cseconds (active window)GPU 0 core
Power draw (W)
0.001002003004005006007000.0010.020.030.040.050.0Wseconds (active window)550 W limitGPU 0
SM clock (MHz)
0.005001000150020002500300035000.0010.020.030.040.050.0MHzseconds (active window)GPU 0
glm-4.6-355bGPU 0 + 1power-cap limited (expected at your set power limit)

-hf unsloth/GLM-4.6-GGUF:UD-Q2_K_XL | placement: dual

showing the 108s active window; 604s of model-load/idle trimmed from a 712s recording

Temperature (°C)
0.0020.040.060.080.01000.0020.040.060.080.0100120°Cseconds (active window)GPU 0 coreGPU 1 core
Power draw (W)
0.001002003004005006000.0020.040.060.080.0100120Wseconds (active window)550 W limitGPU 0GPU 1
SM clock (MHz)
0.005001000150020002500300035000.0020.040.060.080.0100120MHzseconds (active window)GPU 0GPU 1
gpt-oss-120bGPU 0power-cap limited (expected at your set power limit)

-hf ggml-org/gpt-oss-120b-GGUF | placement: gpu0

showing the 21s active window; 187s of model-load/idle trimmed from a 208s recording

Temperature (°C)
0.0010.020.030.040.050.060.070.00.005.0010.015.020.025.0°Cseconds (active window)GPU 0 core
Power draw (W)
0.001002003004005006000.005.0010.015.020.025.0Wseconds (active window)550 W limitGPU 0
SM clock (MHz)
0.005001000150020002500300035000.005.0010.015.020.025.0MHzseconds (active window)GPU 0
hermes-4-70bGPU 0power-cap limited (expected at your set power limit)

-hf unsloth/Hermes-4-70B-GGUF:Q8_0 | placement: gpu0

showing the 124s active window; 228s of model-load/idle trimmed from a 351s recording

Temperature (°C)
0.0020.040.060.080.01000.0020.040.060.080.0100120140°Cseconds (active window)GPU 0 core
Power draw (W)
0.001002003004005006007000.0020.040.060.080.0100120140Wseconds (active window)550 W limitGPU 0
SM clock (MHz)
0.005001000150020002500300035000.0020.040.060.080.0100120140MHzseconds (active window)GPU 0
hermes-4.3-36bGPU 0power-cap limited (expected at your set power limit)

-hf NousResearch/Hermes-4.3-36B-GGUF:Q8_0 | placement: gpu0

showing the 76s active window; 115s of model-load/idle trimmed from a 190s recording

Temperature (°C)
0.0020.040.060.080.00.0020.040.060.080.0°Cseconds (active window)GPU 0 core
Power draw (W)
0.001002003004005006007000.0020.040.060.080.0Wseconds (active window)550 W limitGPU 0
SM clock (MHz)
0.005001000150020002500300035000.0020.040.060.080.0MHzseconds (active window)GPU 0
minimax-m2.7GPU 0 + 1power-cap limited (expected at your set power limit)

-hf unsloth/MiniMax-M2.7-GGUF:UD-Q4_K_XL | placement: dual

showing the 126s active window; 550s of model-load/idle trimmed from a 676s recording

Temperature (°C)
0.0010.020.030.040.050.060.070.080.00.0020.040.060.080.0100120140°Cseconds (active window)GPU 0 coreGPU 1 core
Power draw (W)
0.001002003004005006000.0020.040.060.080.0100120140Wseconds (active window)550 W limitGPU 0GPU 1
SM clock (MHz)
0.005001000150020002500300035000.0020.040.060.080.0100120140MHzseconds (active window)GPU 0GPU 1
qwen2.5-coder-32bGPU 0power-cap limited (expected at your set power limit)

-hf unsloth/Qwen2.5-Coder-32B-Instruct-GGUF:Q8_0 | placement: gpu0

showing the 63s active window; 101s of model-load/idle trimmed from a 164s recording

Temperature (°C)
0.0010.020.030.040.050.060.070.080.00.0010.020.030.040.050.060.070.0°Cseconds (active window)GPU 0 core
Power draw (W)
0.001002003004005006007000.0010.020.030.040.050.060.070.0Wseconds (active window)550 W limitGPU 0
SM clock (MHz)
0.005001000150020002500300035000.0010.020.030.040.050.060.070.0MHzseconds (active window)GPU 0
qwen3-235b-a22bGPU 0 + 1power-cap limited (expected at your set power limit)

-hf unsloth/Qwen3-235B-A22B-Instruct-2507-GGUF:Q4_K_M | placement: dual

showing the 58s active window; 683s of model-load/idle trimmed from a 741s recording

Temperature (°C)
0.0010.020.030.040.050.060.070.080.00.0010.020.030.040.050.060.0°Cseconds (active window)GPU 0 coreGPU 1 core
Power draw (W)
0.001002003004005006000.0010.020.030.040.050.060.0Wseconds (active window)550 W limitGPU 0GPU 1
SM clock (MHz)
0.005001000150020002500300035000.0010.020.030.040.050.060.0MHzseconds (active window)GPU 0GPU 1
qwen3-coder-480bGPU 0 + 1power-cap limited (expected at your set power limit)

-hf unsloth/Qwen3-Coder-480B-A35B-Instruct-GGUF:UD-Q2_K_XL | placement: dual

showing the 86s active window; 1113s of model-load/idle trimmed from a 1199s recording

Temperature (°C)
0.0020.040.060.080.01000.0020.040.060.080.0100°Cseconds (active window)GPU 0 coreGPU 1 core
Power draw (W)
0.001002003004005006000.0020.040.060.080.0100Wseconds (active window)550 W limitGPU 0GPU 1
SM clock (MHz)
0.005001000150020002500300035000.0020.040.060.080.0100MHzseconds (active window)GPU 0GPU 1
qwen3-coder-next-80bGPU 0no throttling

-hf unsloth/Qwen3-Coder-Next-GGUF:Q8_0 | placement: gpu0

showing the 34s active window; 250s of model-load/idle trimmed from a 283s recording

Temperature (°C)
0.0010.020.030.040.050.060.070.00.005.0010.015.020.025.030.035.0°Cseconds (active window)GPU 0 core
Power draw (W)
0.001002003004005006000.005.0010.015.020.025.030.035.0Wseconds (active window)550 W limitGPU 0
SM clock (MHz)
0.005001000150020002500300035000.005.0010.015.020.025.030.035.0MHzseconds (active window)GPU 0
qwen3.6-27bGPU 0power-cap limited (expected at your set power limit)

-hf unsloth/Qwen3.6-27B-GGUF:Q8_0 | placement: gpu0

showing the 51s active window; 85s of model-load/idle trimmed from a 135s recording

Temperature (°C)
0.0010.020.030.040.050.060.070.080.00.0010.020.030.040.050.060.0°Cseconds (active window)GPU 0 core
Power draw (W)
0.001002003004005006007000.0010.020.030.040.050.060.0Wseconds (active window)550 W limitGPU 0
SM clock (MHz)
0.005001000150020002500300035000.0010.020.030.040.050.060.0MHzseconds (active window)GPU 0
qwen3.6-35b-a3bGPU 0no throttling

-hf unsloth/Qwen3.6-35B-A3B-GGUF:Q8_0 | placement: gpu0

showing the 19s active window; 110s of model-load/idle trimmed from a 129s recording

Temperature (°C)
0.0010.020.030.040.050.060.070.00.005.0010.015.020.0°Cseconds (active window)GPU 0 core
Power draw (W)
0.001002003004005006000.005.0010.015.020.0Wseconds (active window)550 W limitGPU 0
SM clock (MHz)
0.005001000150020002500300035000.005.0010.015.020.0MHzseconds (active window)GPU 0

Summary table

ModelQuantGPUsPrefill 512Prefill 4096Decode @d0Decode @d16384t/s per W (avg)Avg power WPeak power WMax VRAM GiBMax core °CThrottled
devstral-small-2-24bQ8_005480542861.456.20.1444155625.264.0power cap
glm-4.6-355bUD-Q2_K_XL0+1937114761.137.80.0971746213077.0power cap
gpt-oss-120b0949994462752350.7835155359.956.0power cap
hermes-4-70bQ8_001896180620.818.40.0448556273.077.0power cap
hermes-4.3-36bQ8_003470311639.532.10.0848055938.673.0power cap
minimax-m2.7UD-Q4_K_XL0+11825369112989.70.4230955413567.0power cap
qwen2.5-coder-32bQ8_003838366343.438.60.0946555935.369.0power cap
qwen3-235b-a22bQ4_K_M0+11982260283.953.60.1270147813669.0power cap
qwen3-coder-480bUD-Q2_K_XL0+1907140873.448.90.1077053317278.0power cap
qwen3-coder-next-80bQ8_00486548901951840.6629546480.056.0none
qwen3.6-27bQ8_004102410450.548.90.1144955827.368.0power cap
qwen3.6-35b-a3bQ8_00992398192392250.7531854235.356.0none