llama.cpp Benchmark Report

Generated 2026-06-12 16:32:01 · 12 model(s) with data

System

GPUs
2 x NVIDIA RTX PRO 6000 Blackwell Workstation Edition
GPU driver
610.43.02
GPU power limit
450 W (set via nvidia-smi -pl; persistence mode on)
CPU
AMD Ryzen 9 9950X 16-Core Processor
System RAM
128 GiB (OS-visible: 123 GiB)
llama.cpp build
version: 9590 (d2462f8f7) built with GNU 14.2.0 for Linux x86_64

Benchmark configuration

GPU power limit (both GPUs)
450 W
Flash attention
on (`-fa 1`)
KV cache type (K and V)
q8_0 (`-ctk q8_0 -ctv q8_0`)
CPU threads
16 (`-t 16`)
GPU layers
all (`-ngl 999`)
Batch / micro-batch
llama-bench defaults (`-b 2048 -ub 512`)
Multi-GPU split (dual)
layer (`-sm layer`)
Repetitions
5 (`-r 5`)
Cool-down between models
300 s
nvidia-smi sampling
every 500 ms, scoped to the run's GPUs

Models under test

AliasHF repo : quantPlacementNote
devstral-small-2-24bunsloth/Devstral-Small-2507-GGUF:Q8_0gpu024B dense, Q8 ~26 GB (verify repo)
glm-4.6-355bunsloth/GLM-4.6-GGUF:UD-Q2_K_XLdual355B-A32B, Q2_K_XL ~135 GB, dual
gpt-oss-120bggml-org/gpt-oss-120b-GGUFgpu0MXFP4 native, ~60 GB, single card
hermes-4-70bunsloth/Hermes-4-70B-GGUF:Q8_0gpu070B dense (Llama 3.1), Q8 ~75 GB, single card
hermes-4.3-36bNousResearch/Hermes-4.3-36B-GGUF:Q8_0gpu036B dense (Seed-OSS-36B), Q8 ~38 GB, single card
minimax-m2.7unsloth/MiniMax-M2.7-GGUF:UD-Q4_K_XLdual~230B-A10B, Q4, dual (verify tag)
qwen2.5-coder-32bunsloth/Qwen2.5-Coder-32B-Instruct-GGUF:Q8_0gpu032B dense, Q8 ~35 GB
qwen3-235b-a22bunsloth/Qwen3-235B-A22B-Instruct-2507-GGUF:Q4_K_Mdual235B-A22B, Q4 ~133 GB, dual
qwen3-coder-480bunsloth/Qwen3-Coder-480B-A35B-Instruct-GGUF:UD-Q2_K_XLdual480B-A35B, Q2 ~180 GB, dual (tight)
qwen3-coder-next-80bunsloth/Qwen3-Coder-Next-GGUF:Q8_0gpu080B-A3B, Q8 ~85 GB, single card
qwen3.6-27bunsloth/Qwen3.6-27B-GGUF:Q8_0gpu027B dense, Q8 ~30 GB, single card
qwen3.6-35b-a3bunsloth/Qwen3.6-35B-A3B-GGUF:Q8_0gpu035B-A3B MoE, Q8 ~38 GB, single card

Prefill throughput (prompt processing)

Prefill 512 @ d0Prefill 4096 @ d00.00200040006000800010000tokens / secondqwen3.6-35b-a3bqwen3.6-35b-a3b — Prefill 512 @ d0: 96209620qwen3.6-35b-a3b — Prefill 4096 @ d0: 88888888gpt-oss-120bgpt-oss-120b — Prefill 512 @ d0: 90809080gpt-oss-120b — Prefill 4096 @ d0: 85128512devstral-small-2-24bdevstral-small-2-24b — Prefill 512 @ d0: 49284928devstral-small-2-24b — Prefill 4096 @ d0: 47024702qwen3-coder-next-80bqwen3-coder-next-80b — Prefill 512 @ d0: 48664866qwen3-coder-next-80b — Prefill 4096 @ d0: 48824882qwen3.6-27bqwen3.6-27b — Prefill 512 @ d0: 36733673qwen3.6-27b — Prefill 4096 @ d0: 35913591qwen2.5-coder-32bqwen2.5-coder-32b — Prefill 512 @ d0: 34013401qwen2.5-coder-32b — Prefill 4096 @ d0: 31863186minimax-m2.7minimax-m2.7 — Prefill 512 @ d0: 17861786minimax-m2.7 — Prefill 4096 @ d0: 32633263hermes-4.3-36bhermes-4.3-36b — Prefill 512 @ d0: 30223022hermes-4.3-36b — Prefill 4096 @ d0: 27072707qwen3-235b-a22bqwen3-235b-a22b — Prefill 512 @ d0: 19071907qwen3-235b-a22b — Prefill 4096 @ d0: 25302530hermes-4-70bhermes-4-70b — Prefill 512 @ d0: 16281628hermes-4-70b — Prefill 4096 @ d0: 15541554qwen3-coder-480bqwen3-coder-480b — Prefill 512 @ d0: 830830qwen3-coder-480b — Prefill 4096 @ d0: 12821282glm-4.6-355bglm-4.6-355b — Prefill 512 @ d0: 858858glm-4.6-355b — Prefill 4096 @ d0: 10451045

Decode throughput (token generation)

Decode 128 @ d0Decode 128 @ d163840.0050.0100150200250300tokens / secondgpt-oss-120bgpt-oss-120b — Decode 128 @ d0: 274274gpt-oss-120b — Decode 128 @ d16384: 231231qwen3.6-35b-a3bqwen3.6-35b-a3b — Decode 128 @ d0: 239239qwen3.6-35b-a3b — Decode 128 @ d16384: 225225qwen3-coder-next-80bqwen3-coder-next-80b — Decode 128 @ d0: 195195qwen3-coder-next-80b — Decode 128 @ d16384: 184184minimax-m2.7minimax-m2.7 — Decode 128 @ d0: 130130minimax-m2.7 — Decode 128 @ d16384: 89.789.7qwen3-235b-a22bqwen3-235b-a22b — Decode 128 @ d0: 83.883.8qwen3-235b-a22b — Decode 128 @ d16384: 53.553.5qwen3-coder-480bqwen3-coder-480b — Decode 128 @ d0: 73.373.3qwen3-coder-480b — Decode 128 @ d16384: 49.049.0devstral-small-2-24bdevstral-small-2-24b — Decode 128 @ d0: 61.461.4devstral-small-2-24b — Decode 128 @ d16384: 55.855.8glm-4.6-355bglm-4.6-355b — Decode 128 @ d0: 61.061.0glm-4.6-355b — Decode 128 @ d16384: 37.937.9qwen3.6-27bqwen3.6-27b — Decode 128 @ d0: 50.550.5qwen3.6-27b — Decode 128 @ d16384: 48.848.8qwen2.5-coder-32bqwen2.5-coder-32b — Decode 128 @ d0: 43.443.4qwen2.5-coder-32b — Decode 128 @ d16384: 38.038.0hermes-4.3-36bhermes-4.3-36b — Decode 128 @ d0: 39.539.5hermes-4.3-36b — Decode 128 @ d16384: 31.031.0hermes-4-70bhermes-4-70b — Decode 128 @ d0: 20.820.8hermes-4-70b — Decode 128 @ d16384: 18.118.1

Peak temperatures and power

Maxima are taken over the active benchmark window of each run; only GPUs that actively participated are shown, so single-card runs omit the idle second card. Peak power samples a few percent above the configured limit are normal: nvidia-smi reports instantaneous draw while NVIDIA's power controller regulates a time-averaged budget, so brief transients above the cap are expected and not a fault. Compare the average power chart against the limit instead.

Max core temperature

GPU 0GPU 10.0010.020.030.040.050.060.070.080.0degrees Cdevstral-small-2-24bdevstral-small-2-24b — GPU 0: 60.060.0glm-4.6-355bglm-4.6-355b — GPU 0: 73.073.0glm-4.6-355b — GPU 1: 65.065.0gpt-oss-120bgpt-oss-120b — GPU 0: 52.052.0hermes-4-70bhermes-4-70b — GPU 0: 72.072.0hermes-4.3-36bhermes-4.3-36b — GPU 0: 66.066.0minimax-m2.7minimax-m2.7 — GPU 0: 63.063.0minimax-m2.7 — GPU 1: 60.060.0qwen2.5-coder-32bqwen2.5-coder-32b — GPU 0: 64.064.0qwen3-235b-a22bqwen3-235b-a22b — GPU 0: 67.067.0qwen3-235b-a22b — GPU 1: 61.061.0qwen3-coder-480bqwen3-coder-480b — GPU 0: 74.074.0qwen3-coder-480b — GPU 1: 64.064.0qwen3-coder-next-80bqwen3-coder-next-80b — GPU 0: 55.055.0qwen3.6-27bqwen3.6-27b — GPU 0: 64.064.0qwen3.6-35b-a3bqwen3.6-35b-a3b — GPU 0: 53.053.0

Max power draw (instantaneous samples)

GPU 0GPU 10.00100200300400500wattsdevstral-small-2-24bdevstral-small-2-24b — GPU 0: 462462glm-4.6-355bglm-4.6-355b — GPU 0: 403403glm-4.6-355b — GPU 1: 406406gpt-oss-120bgpt-oss-120b — GPU 0: 459459hermes-4-70bhermes-4-70b — GPU 0: 469469hermes-4.3-36bhermes-4.3-36b — GPU 0: 456456minimax-m2.7minimax-m2.7 — GPU 0: 466466minimax-m2.7 — GPU 1: 466466qwen2.5-coder-32bqwen2.5-coder-32b — GPU 0: 467467qwen3-235b-a22bqwen3-235b-a22b — GPU 0: 437437qwen3-235b-a22b — GPU 1: 405405qwen3-coder-480bqwen3-coder-480b — GPU 0: 466466qwen3-coder-480b — GPU 1: 392392qwen3-coder-next-80bqwen3-coder-next-80b — GPU 0: 452452qwen3.6-27bqwen3.6-27b — GPU 0: 461461qwen3.6-35b-a3bqwen3.6-35b-a3b — GPU 0: 459459

Average power draw (over the active window)

GPU 0GPU 10.00100200300400500wattsdevstral-small-2-24bdevstral-small-2-24b — GPU 0: 395395glm-4.6-355bglm-4.6-355b — GPU 0: 212212glm-4.6-355b — GPU 1: 216216gpt-oss-120bgpt-oss-120b — GPU 0: 308308hermes-4-70bhermes-4-70b — GPU 0: 417417hermes-4.3-36bhermes-4.3-36b — GPU 0: 412412minimax-m2.7minimax-m2.7 — GPU 0: 345345minimax-m2.7 — GPU 1: 349349qwen2.5-coder-32bqwen2.5-coder-32b — GPU 0: 409409qwen3-235b-a22bqwen3-235b-a22b — GPU 0: 344344qwen3-235b-a22b — GPU 1: 331331qwen3-coder-480bqwen3-coder-480b — GPU 0: 379379qwen3-coder-480b — GPU 1: 324324qwen3-coder-next-80bqwen3-coder-next-80b — GPU 0: 296296qwen3.6-27bqwen3.6-27b — GPU 0: 408408qwen3.6-35b-a3bqwen3.6-35b-a3b — GPU 0: 297297

Per-model thermals over time

Charts are trimmed to the active benchmark window: the model-loading phase and idle tails are cut (the amount removed is noted per model), so the measured runs fill the chart instead of being squeezed by minutes of GGUF loading. SM clock dropping during a run is the clearest sign of throttling. Power-cap limited is expected whenever the power limit is set below the card's maximum — the card is staying within budget, not faulting. Thermal / HW slowdown means the card hit a temperature or hardware limit and reduced clocks; investigate cooling and airflow.

devstral-small-2-24bGPU 0power-cap limited (expected at your set power limit)

-hf unsloth/Devstral-Small-2507-GGUF:Q8_0 | placement: gpu0

showing the 45s active window; 72s of model-load/idle trimmed from a 117s recording

Temperature (°C)
0.0010.020.030.040.050.060.070.00.0010.020.030.040.050.0°Cseconds (active window)GPU 0 core
Power draw (W)
0.001002003004005000.0010.020.030.040.050.0Wseconds (active window)450 W limitGPU 0
SM clock (MHz)
0.005001000150020002500300035000.0010.020.030.040.050.0MHzseconds (active window)GPU 0
glm-4.6-355bGPU 0 + 1power-cap limited (expected at your set power limit)

-hf unsloth/GLM-4.6-GGUF:UD-Q2_K_XL | placement: dual

showing the 195s active window; 453s of model-load/idle trimmed from a 648s recording

Temperature (°C)
0.0020.040.060.080.00.0050.0100150200°Cseconds (active window)GPU 0 coreGPU 1 core
Power draw (W)
0.001002003004005000.0050.0100150200Wseconds (active window)450 W limitGPU 0GPU 1
SM clock (MHz)
0.005001000150020002500300035000.0050.0100150200MHzseconds (active window)GPU 0GPU 1
gpt-oss-120bGPU 0power-cap limited (expected at your set power limit)

-hf ggml-org/gpt-oss-120b-GGUF | placement: gpu0

showing the 22s active window; 184s of model-load/idle trimmed from a 206s recording

Temperature (°C)
0.0010.020.030.040.050.060.00.005.0010.015.020.025.0°Cseconds (active window)GPU 0 core
Power draw (W)
0.001002003004005000.005.0010.015.020.025.0Wseconds (active window)450 W limitGPU 0
SM clock (MHz)
0.005001000150020002500300035000.005.0010.015.020.025.0MHzseconds (active window)GPU 0
hermes-4-70bGPU 0power-cap limited (expected at your set power limit)

-hf unsloth/Hermes-4-70B-GGUF:Q8_0 | placement: gpu0

showing the 131s active window; 228s of model-load/idle trimmed from a 359s recording

Temperature (°C)
0.0020.040.060.080.00.0020.040.060.080.0100120140°Cseconds (active window)GPU 0 core
Power draw (W)
0.001002003004005006000.0020.040.060.080.0100120140Wseconds (active window)450 W limitGPU 0
SM clock (MHz)
0.005001000150020002500300035000.0020.040.060.080.0100120140MHzseconds (active window)GPU 0
hermes-4.3-36bGPU 0power-cap limited (expected at your set power limit)

-hf NousResearch/Hermes-4.3-36B-GGUF:Q8_0 | placement: gpu0

showing the 81s active window; 114s of model-load/idle trimmed from a 195s recording

Temperature (°C)
0.0010.020.030.040.050.060.070.080.00.0020.040.060.080.0100°Cseconds (active window)GPU 0 core
Power draw (W)
0.001002003004005000.0020.040.060.080.0100Wseconds (active window)450 W limitGPU 0
SM clock (MHz)
0.005001000150020002500300035000.0020.040.060.080.0100MHzseconds (active window)GPU 0
minimax-m2.7GPU 0 + 1power-cap limited (expected at your set power limit)

-hf unsloth/MiniMax-M2.7-GGUF:UD-Q4_K_XL | placement: dual

showing the 44s active window; 641s of model-load/idle trimmed from a 684s recording

Temperature (°C)
0.0010.020.030.040.050.060.070.00.0010.020.030.040.050.0°Cseconds (active window)GPU 0 coreGPU 1 core
Power draw (W)
0.001002003004005006000.0010.020.030.040.050.0Wseconds (active window)450 W limitGPU 0GPU 1
SM clock (MHz)
0.005001000150020002500300035000.0010.020.030.040.050.0MHzseconds (active window)GPU 0GPU 1
qwen2.5-coder-32bGPU 0power-cap limited (expected at your set power limit)

-hf unsloth/Qwen2.5-Coder-32B-Instruct-GGUF:Q8_0 | placement: gpu0

showing the 66s active window; 101s of model-load/idle trimmed from a 167s recording

Temperature (°C)
0.0010.020.030.040.050.060.070.00.0010.020.030.040.050.060.070.0°Cseconds (active window)GPU 0 core
Power draw (W)
0.001002003004005006000.0010.020.030.040.050.060.070.0Wseconds (active window)450 W limitGPU 0
SM clock (MHz)
0.005001000150020002500300035000.0010.020.030.040.050.060.070.0MHzseconds (active window)GPU 0
qwen3-235b-a22bGPU 0 + 1power-cap limited (expected at your set power limit)

-hf unsloth/Qwen3-235B-A22B-Instruct-2507-GGUF:Q4_K_M | placement: dual

showing the 60s active window; 864s of model-load/idle trimmed from a 924s recording

Temperature (°C)
0.0010.020.030.040.050.060.070.080.00.0010.020.030.040.050.060.0°Cseconds (active window)GPU 0 coreGPU 1 core
Power draw (W)
0.001002003004005000.0010.020.030.040.050.060.0Wseconds (active window)450 W limitGPU 0GPU 1
SM clock (MHz)
0.005001000150020002500300035000.0010.020.030.040.050.060.0MHzseconds (active window)GPU 0GPU 1
qwen3-coder-480bGPU 0 + 1power-cap limited (expected at your set power limit)

-hf unsloth/Qwen3-Coder-480B-A35B-Instruct-GGUF:UD-Q2_K_XL | placement: dual

showing the 92s active window; 1111s of model-load/idle trimmed from a 1202s recording

Temperature (°C)
0.0020.040.060.080.00.0020.040.060.080.0100°Cseconds (active window)GPU 0 coreGPU 1 core
Power draw (W)
0.001002003004005006000.0020.040.060.080.0100Wseconds (active window)450 W limitGPU 0GPU 1
SM clock (MHz)
0.005001000150020002500300035000.0020.040.060.080.0100MHzseconds (active window)GPU 0GPU 1
qwen3-coder-next-80bGPU 0power-cap limited (expected at your set power limit)

-hf unsloth/Qwen3-Coder-Next-GGUF:Q8_0 | placement: gpu0

showing the 33s active window; 212s of model-load/idle trimmed from a 245s recording

Temperature (°C)
0.0010.020.030.040.050.060.00.005.0010.015.020.025.030.035.0°Cseconds (active window)GPU 0 core
Power draw (W)
0.001002003004005000.005.0010.015.020.025.030.035.0Wseconds (active window)450 W limitGPU 0
SM clock (MHz)
0.005001000150020002500300035000.005.0010.015.020.025.030.035.0MHzseconds (active window)GPU 0
qwen3.6-27bGPU 0power-cap limited (expected at your set power limit)

-hf unsloth/Qwen3.6-27B-GGUF:Q8_0 | placement: gpu0

showing the 53s active window; 85s of model-load/idle trimmed from a 137s recording

Temperature (°C)
0.0010.020.030.040.050.060.070.00.0010.020.030.040.050.060.0°Cseconds (active window)GPU 0 core
Power draw (W)
0.001002003004005000.0010.020.030.040.050.060.0Wseconds (active window)450 W limitGPU 0
SM clock (MHz)
0.005001000150020002500300035000.0010.020.030.040.050.060.0MHzseconds (active window)GPU 0
qwen3.6-35b-a3bGPU 0power-cap limited (expected at your set power limit)

-hf unsloth/Qwen3.6-35B-A3B-GGUF:Q8_0 | placement: gpu0

showing the 20s active window; 110s of model-load/idle trimmed from a 130s recording

Temperature (°C)
0.0010.020.030.040.050.060.00.005.0010.015.020.025.0°Cseconds (active window)GPU 0 core
Power draw (W)
0.001002003004005000.005.0010.015.020.025.0Wseconds (active window)450 W limitGPU 0
SM clock (MHz)
0.005001000150020002500300035000.005.0010.015.020.025.0MHzseconds (active window)GPU 0

Summary table

ModelQuantGPUsPrefill 512Prefill 4096Decode @d0Decode @d16384t/s per W (avg)Avg power WPeak power WMax VRAM GiBMax core °CThrottled
devstral-small-2-24bQ8_004928470261.455.80.1639546225.260.0power cap
glm-4.6-355bUD-Q2_K_XL0+1858104561.037.90.1442840613073.0power cap
gpt-oss-120b0908085122742310.8930845959.952.0power cap
hermes-4-70bQ8_001628155420.818.10.0541746973.072.0power cap
hermes-4.3-36bQ8_003022270739.531.00.1041245638.666.0power cap
minimax-m2.7UD-Q4_K_XL0+11786326313089.70.1969446613563.0power cap
qwen2.5-coder-32bQ8_003401318643.438.00.1140946735.364.0power cap
qwen3-235b-a22bQ4_K_M0+11907253083.853.50.1267443713667.0power cap
qwen3-coder-480bUD-Q2_K_XL0+1830128273.349.00.1070346617274.0power cap
qwen3-coder-next-80bQ8_00486648821951840.6629645280.055.0power cap
qwen3.6-27bQ8_003673359150.548.80.1240846127.364.0power cap
qwen3.6-35b-a3bQ8_00962088882392250.8029745935.353.0power cap