Generated 2026-06-13 10:35:20 · 12 model(s) with data
| Alias | HF repo : quant | Placement | Note |
|---|---|---|---|
| devstral-small-2-24b | unsloth/Devstral-Small-2507-GGUF:Q8_0 | gpu0 | 24B dense, Q8 ~26 GB (verify repo) |
| glm-4.6-355b | unsloth/GLM-4.6-GGUF:UD-Q2_K_XL | dual | 355B-A32B, Q2_K_XL ~135 GB, dual |
| gpt-oss-120b | ggml-org/gpt-oss-120b-GGUF | gpu0 | MXFP4 native, ~60 GB, single card |
| hermes-4-70b | unsloth/Hermes-4-70B-GGUF:Q8_0 | gpu0 | 70B dense (Llama 3.1), Q8 ~75 GB, single card |
| hermes-4.3-36b | NousResearch/Hermes-4.3-36B-GGUF:Q8_0 | gpu0 | 36B dense (Seed-OSS-36B), Q8 ~38 GB, single card |
| minimax-m2.7 | unsloth/MiniMax-M2.7-GGUF:UD-Q4_K_XL | dual | ~230B-A10B, Q4, dual (verify tag) |
| qwen2.5-coder-32b | unsloth/Qwen2.5-Coder-32B-Instruct-GGUF:Q8_0 | gpu0 | 32B dense, Q8 ~35 GB |
| qwen3-235b-a22b | unsloth/Qwen3-235B-A22B-Instruct-2507-GGUF:Q4_K_M | dual | 235B-A22B, Q4 ~133 GB, dual |
| qwen3-coder-480b | unsloth/Qwen3-Coder-480B-A35B-Instruct-GGUF:UD-Q2_K_XL | dual | 480B-A35B, Q2 ~180 GB, dual (tight) |
| qwen3-coder-next-80b | unsloth/Qwen3-Coder-Next-GGUF:Q8_0 | gpu0 | 80B-A3B, Q8 ~85 GB, single card |
| qwen3.6-27b | unsloth/Qwen3.6-27B-GGUF:Q8_0 | gpu0 | 27B dense, Q8 ~30 GB, single card |
| qwen3.6-35b-a3b | unsloth/Qwen3.6-35B-A3B-GGUF:Q8_0 | gpu0 | 35B-A3B MoE, Q8 ~38 GB, single card |
Maxima are taken over the active benchmark window of each run; only GPUs that actively participated are shown, so single-card runs omit the idle second card. Peak power samples a few percent above the configured limit are normal: nvidia-smi reports instantaneous draw while NVIDIA's power controller regulates a time-averaged budget, so brief transients above the cap are expected and not a fault. Compare the average power chart against the limit instead.
Charts are trimmed to the active benchmark window: the model-loading phase and idle tails are cut (the amount removed is noted per model), so the measured runs fill the chart instead of being squeezed by minutes of GGUF loading. SM clock dropping during a run is the clearest sign of throttling. Power-cap limited is expected whenever the power limit is set below the card's maximum — the card is staying within budget, not faulting. Thermal / HW slowdown means the card hit a temperature or hardware limit and reduced clocks; investigate cooling and airflow.
| Model | Quant | GPUs | Prefill 512 | Prefill 4096 | Decode @d0 | Decode @d16384 | t/s per W (avg) | Avg power W | Peak power W | Max VRAM GiB | Max core °C | Throttled |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| devstral-small-2-24b | Q8_0 | 0 | 5844 | 5718 | 61.4 | 56.2 | 0.13 | 460 | 603 | 25.2 | 65.0 | power cap |
| glm-4.6-355b | UD-Q2_K_XL | 0+1 | 960 | 1178 | 61.1 | 37.8 | 0.12 | 498 | 496 | 130 | 78.0 | power cap |
| gpt-oss-120b | – | 0 | 9505 | 9541 | 275 | 235 | 0.84 | 327 | 582 | 59.9 | 56.0 | none |
| hermes-4-70b | Q8_0 | 0 | 2006 | 1915 | 20.8 | 18.4 | 0.04 | 501 | 607 | 73.0 | 79.0 | power cap |
| hermes-4.3-36b | Q8_0 | 0 | 3639 | 3282 | 39.5 | 32.1 | 0.08 | 492 | 605 | 38.6 | 74.0 | power cap |
| minimax-m2.7 | UD-Q4_K_XL | 0+1 | 1837 | 3727 | 129 | 89.6 | 0.17 | 764 | 569 | 135 | 68.0 | power cap |
| qwen2.5-coder-32b | Q8_0 | 0 | 3977 | 3850 | 43.4 | 38.6 | 0.09 | 474 | 607 | 35.3 | 71.0 | power cap |
| qwen3-235b-a22b | Q4_K_M | 0+1 | 1977 | 2613 | 83.9 | 53.6 | 0.12 | 711 | 478 | 136 | 68.0 | power cap |
| qwen3-coder-480b | UD-Q2_K_XL | 0+1 | 928 | 1444 | 73.4 | 48.9 | 0.09 | 798 | 558 | 172 | 81.0 | power cap |
| qwen3-coder-next-80b | Q8_0 | 0 | 4869 | 4887 | 195 | 183 | 0.64 | 307 | 467 | 80.0 | 56.0 | none |
| qwen3.6-27b | Q8_0 | 0 | 4276 | 4299 | 50.5 | 48.9 | 0.11 | 466 | 603 | 27.3 | 70.0 | power cap |
| qwen3.6-35b-a3b | Q8_0 | 0 | 9919 | 9810 | 239 | 225 | 0.77 | 311 | 540 | 35.3 | 56.0 | none |