I’m trying to decide whether I want to switch to Qwen3.8. I’m guessing that the performance for the FP16 will be deadly slow. How does FP8 run? I’ve seen a MTP version on Hugging Face. Should I just wait until they create a MoE model? I’m loving Qwen3.6 MoE, the performance isn’t exactly fast but I can tolerate it. Looking for people who must try the new things to see if I want to try this new thing.
I get 8.6 t/s on the strix halo using this command line for llama.cpp:
llama-server -ngl 99 --flash-attn auto --cache-type-k q8_0 --cache-type-v q8_0 -c 65536 --parallel 1 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --presence-penalty 0.0 --repeat-penalty 1.0 --jinja -hf unsloth/Qwen3.8-27B-GGUF:UD-Q6_K_XL
This compares to 21.8 t/s on one Radeon 9700 AI card on the same system using a similar command line.
Per Simon Willison’s notes, it can do a lot of reasoning before responding, he suggests reducing the reasoning_effort, which I’ve done by adding --reasoning-budget 8192 --chat-template-kwargs '{"reasoning_effort":"low"}' for llama.cpp at cli (maybe you need that in your jinja template).
Simon Wilson was talking about the just released qwen 3.8 flash next model, not the 27b model I tested earlier, so I downloaded and tested that one as well.
As always, the tokens per second value varies quite a bit depending on the size of the context window you want. I tend to prefer the larger(ish) context windows and a little more thinking as you can see in my initial post in this thread, using the 64KiB context window.
For the strix halo 128GiB alone, the highest quant I could fit was the Q4. Larger sizes were getting killed by the OOM handler.
With this command line, I got ~16 tokens/second.
llama-server -ngl 99 --flash-attn auto --cache-type-k q8_0 --cache-type-v q8_0 -c 32768 --parallel 1 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --presence-penalty 0.0 --repeat-penalty 1.0 --jinja --reasoning-budget 8192 --chat-template-kwargs '{"reasoning_effort":"medium"}' -hf unsloth/Qwen3.8-Flash-Next-GGUF:UD-IQ4_XS
Just by changing the context window size down to 16KiB, -c 16384 I was able to get ~21 t/s.
With the 2 extra AMD 9700 cards attached to my system, I was (barely) able to fit the Q6 model with a 64K context window. This got me ~17 Tokens/Second with hopefully better quality.
For reference: on a strix point (not halo), I got around 6 t/s for the 4b quant of Qwen3.8-27B, using lemonade rocm llama build and 8k context.
Will try 27B and the flash version on a halo soon.
I don’t see mention of flash attention in the piece I linked or the Qwen 3.8 27B release page it also talks about. I don’t know where you got that idea.
I think you’re supposed to limit cache so the 51gb of ngram lookup stays on disk. I am on vacation without my Strix Halo but I was going to test IQ5_NL with a limited cache. Didn’t dive into how they build that ngram lookup to see whether it can be compressed or quantized.
TBH while this seems to be made for these unified memory machines, if it’s going to get 20-30t/s, I’m not really sure it’s going to be a great experience for agentic coding. I’m considering coughing up money for 4 V100s or 2 MI210s if the “as good as Opus 4.6” benchmark result seems to hold up.
Yeah, I just went back and looked, and I have no clue where I got the idea either. It mentions 27b all over the article.
I guess I got distracted.
Sorry for my confusion.
OP here, I’m happy for anyone to post any type of results with any version of Qwen 3.8. I’m probably going to wait and see if Qwen posts a MoE before I actually jump. I get twitchy with fewer than about 30t/s.
Just getting into local LLMs with the FWD after playing on FW12
on FWD $ amd-ttm
Current TTM pages limit: 25165824 pages (96.00 GB)
Total system memory: 122.83 GB
under Unsloth Desktop loaded unsloth/Qwen3.8-27B-GGUF:Q8_0
Prompt eval2.89s
Prompt speed305.7 tok/s
Generation87.67s
Speed 14.8 tok/s
Tokens1,295
First token2.89s
Cache hits2,368
Total1259.80s
Chunks1148
i was getting more with Q4 under llama-server routing to pi but there are so many variables I’m learning about so will post more stats as i test
Update: Unsloth Dynamic 3 BF16 gguf runs about 6.4 t/s
Been test driving 3.8 Flash on a subscription for coding and it’s been surprisingly satisfactory. I realized today that pipeline parallelism might be doable with the PCI-e slot (since it’s well under 1MB/token) for highly sparse MoE.
It’s not really clear how much cache miss hard drive reads for the ngram table would slow this down but I’m now doing some card shopping, perhaps an old MI100 would be perfect.
It also occurred to me that the KV cache on 3.8 flash next is very small, seems like a feature that will come to the <70b class, so perhaps a MI100 32gb would be quite future proof.
this is run on 2 different llama versions on standard FW desktop (debian/MXlinux ver 13.6)
ROCm, it is using pre-release b10713
Vulkan, it is using EngramHalo.cpp b10710
these set of run, power limit set = 60w
ROCm ave TCTL reporting T = approx 70-75C, Vulkan T = approx 55C
| model | size | params | backend | ngl | n_ubatch | fa | dev | lm | test | t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | -------: | --: | ------------ | ---------: | --------------: | -------------------: |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | ROCm | 256 | 160 | 1 | ROCm0 | none | pp512 @ d131 | 226.22 ± 0.17 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | ROCm | 256 | 160 | 1 | ROCm0 | none | tg128 @ d131 | 7.64 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | ROCm | 256 | 168 | 1 | ROCm0 | none | pp512 @ d131 | 195.50 ± 2.47 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | ROCm | 256 | 168 | 1 | ROCm0 | none | tg128 @ d131 | 7.64 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | ROCm | 256 | 176 | 1 | ROCm0 | none | pp512 @ d131 | 214.46 ± 0.05 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | ROCm | 256 | 176 | 1 | ROCm0 | none | tg128 @ d131 | 7.64 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | ROCm | 256 | 184 | 1 | ROCm0 | none | pp512 @ d131 | 214.07 ± 0.04 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | ROCm | 256 | 184 | 1 | ROCm0 | none | tg128 @ d131 | 7.64 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | ROCm | 256 | 192 | 1 | ROCm0 | none | pp512 @ d131 | 225.82 ± 5.87 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | ROCm | 256 | 192 | 1 | ROCm0 | none | tg128 @ d131 | 7.64 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | ROCm | 256 | 200 | 1 | ROCm0 | none | pp512 @ d131 | 207.83 ± 0.07 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | ROCm | 256 | 200 | 1 | ROCm0 | none | tg128 @ d131 | 7.64 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | ROCm | 256 | 208 | 1 | ROCm0 | none | pp512 @ d131 | 211.65 ± 0.30 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | ROCm | 256 | 208 | 1 | ROCm0 | none | tg128 @ d131 | 7.64 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | ROCm | 256 | 216 | 1 | ROCm0 | none | pp512 @ d131 | 221.99 ± 5.94 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | ROCm | 256 | 216 | 1 | ROCm0 | none | tg128 @ d131 | 7.64 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | ROCm | 256 | 224 | 1 | ROCm0 | none | pp512 @ d131 | 224.52 ± 0.12 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | ROCm | 256 | 224 | 1 | ROCm0 | none | tg128 @ d131 | 7.64 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | ROCm | 256 | 232 | 1 | ROCm0 | none | pp512 @ d131 | 208.63 ± 0.97 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | ROCm | 256 | 232 | 1 | ROCm0 | none | tg128 @ d131 | 7.64 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | ROCm | 256 | 240 | 1 | ROCm0 | none | pp512 @ d131 | 207.54 ± 0.71 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | ROCm | 256 | 240 | 1 | ROCm0 | none | tg128 @ d131 | 7.64 ± 0.00 |
| qwen4exp A3B Q3_K - Medium | 83.80 GiB | 176.94 B | ROCm | 256 | 192 | 1 | ROCm0 | none | pp512 @ d131 | 261.51 ± 7.18 |
| qwen4exp A3B Q3_K - Medium | 83.80 GiB | 176.94 B | ROCm | 256 | 192 | 1 | ROCm0 | none | tg128 @ d131 | 16.98 ± 0.55 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 16 | 0 | Vulkan0 | none | pp512 @ d131 | 31.90 ± 2.70 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 16 | 0 | Vulkan0 | none | tg128 @ d131 | 7.66 ± 0.01 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 32 | 0 | Vulkan0 | none | pp512 @ d131 | 64.74 ± 2.59 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 32 | 0 | Vulkan0 | none | tg128 @ d131 | 7.71 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 48 | 0 | Vulkan0 | none | pp512 @ d131 | 84.32 ± 0.62 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 48 | 0 | Vulkan0 | none | tg128 @ d131 | 7.71 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 64 | 0 | Vulkan0 | none | pp512 @ d131 | 117.17 ± 0.54 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 64 | 0 | Vulkan0 | none | tg128 @ d131 | 7.71 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 80 | 0 | Vulkan0 | none | pp512 @ d131 | 100.06 ± 0.67 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 80 | 0 | Vulkan0 | none | tg128 @ d131 | 7.70 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 96 | 0 | Vulkan0 | none | pp512 @ d131 | 110.26 ± 0.25 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 96 | 0 | Vulkan0 | none | tg128 @ d131 | 7.70 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 112 | 0 | Vulkan0 | none | pp512 @ d131 | 127.75 ± 0.21 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 112 | 0 | Vulkan0 | none | tg128 @ d131 | 7.70 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 128 | 0 | Vulkan0 | none | pp512 @ d131 | 147.92 ± 0.88 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 128 | 0 | Vulkan0 | none | tg128 @ d131 | 7.70 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 144 | 0 | Vulkan0 | none | pp512 @ d131 | 113.19 ± 0.45 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 144 | 0 | Vulkan0 | none | tg128 @ d131 | 7.70 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 160 | 0 | Vulkan0 | none | pp512 @ d131 | 110.88 ± 0.57 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 160 | 0 | Vulkan0 | none | tg128 @ d131 | 7.70 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 176 | 0 | Vulkan0 | none | pp512 @ d131 | 130.05 ± 0.43 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 176 | 0 | Vulkan0 | none | tg128 @ d131 | 7.70 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 192 | 0 | Vulkan0 | none | pp512 @ d131 | 141.63 ± 0.58 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 192 | 0 | Vulkan0 | none | tg128 @ d131 | 7.70 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 208 | 0 | Vulkan0 | none | pp512 @ d131 | 149.89 ± 0.85 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 208 | 0 | Vulkan0 | none | tg128 @ d131 | 7.70 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 224 | 0 | Vulkan0 | none | pp512 @ d131 | 145.24 ± 0.52 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 224 | 0 | Vulkan0 | none | tg128 @ d131 | 7.69 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 240 | 0 | Vulkan0 | none | pp512 @ d131 | 150.87 ± 0.36 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 240 | 0 | Vulkan0 | none | tg128 @ d131 | 7.70 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 256 | 0 | Vulkan0 | none | pp512 @ d131 | 177.99 ± 0.74 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 256 | 0 | Vulkan0 | none | tg128 @ d131 | 7.70 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 272 | 0 | Vulkan0 | none | pp512 @ d131 | 152.84 ± 0.68 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 272 | 0 | Vulkan0 | none | tg128 @ d131 | 7.70 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 288 | 0 | Vulkan0 | none | pp512 @ d131 | 147.24 ± 1.27 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 288 | 0 | Vulkan0 | none | tg128 @ d131 | 7.70 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 16 | 0 | Vulkan0 | none | pp512 @ d131 | 32.44 ± 2.79 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 16 | 0 | Vulkan0 | none | tg128 @ d131 | 7.69 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 16 | 1 | Vulkan0 | none | pp512 @ d131 | 31.32 ± 1.39 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 16 | 1 | Vulkan0 | none | tg128 @ d131 | 7.69 ± 0.02 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 256 | 0 | Vulkan0 | none | pp512 @ d131 | 177.36 ± 0.78 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 256 | 0 | Vulkan0 | none | tg128 @ d131 | 7.69 ± 0.01 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 256 | 1 | Vulkan0 | none | pp512 @ d131 | 178.47 ± 0.73 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 256 | 1 | Vulkan0 | none | tg128 @ d131 | 7.69 ± 0.02 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 512 | 0 | Vulkan0 | none | pp512 @ d131 | 184.49 ± 0.17 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 512 | 0 | Vulkan0 | none | tg128 @ d131 | 7.67 ± 0.01 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 512 | 1 | Vulkan0 | none | pp512 @ d131 | 187.82 ± 0.21 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 512 | 1 | Vulkan0 | none | tg128 @ d131 | 7.70 ± 0.02 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 640 | 0 | Vulkan0 | none | pp512 @ d131 | 180.54 ± 0.51 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 640 | 0 | Vulkan0 | none | tg128 @ d131 | 7.65 ± 0.03 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 640 | 1 | Vulkan0 | none | pp512 @ d131 | 186.58 ± 2.03 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 640 | 1 | Vulkan0 | none | tg128 @ d131 | 7.67 ± 0.01 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 1024 | 0 | Vulkan0 | none | pp512 @ d131 | 188.14 ± 0.20 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 1024 | 0 | Vulkan0 | none | tg128 @ d131 | 7.67 ± 0.02 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 1024 | 1 | Vulkan0 | none | pp512 @ d131 | 194.75 ± 0.41 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 1024 | 1 | Vulkan0 | none | tg128 @ d131 | 7.68 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 2048 | 0 | Vulkan0 | none | pp512 @ d131 | 191.12 ± 0.65 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 2048 | 0 | Vulkan0 | none | tg128 @ d131 | 7.67 ± 0.01 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 2048 | 1 | Vulkan0 | none | pp512 @ d131 | 196.41 ± 0.15 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 2048 | 1 | Vulkan0 | none | tg128 @ d131 | 7.71 ± 0.02 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 16 | 1 | Vulkan0 | none | pp512 @ d131 | 29.77 ± 0.05 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 16 | 1 | Vulkan0 | none | tg128 @ d131 | 7.72 ± 0.01 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 3072 | 1 | Vulkan0 | none | pp512 @ d131 | 196.41 ± 0.37 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 3072 | 1 | Vulkan0 | none | tg128 @ d131 | 7.71 ± 0.01 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 4096 | 1 | Vulkan0 | none | pp512 @ d131 | 195.20 ± 0.04 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 4096 | 1 | Vulkan0 | none | tg128 @ d131 | 7.70 ± 0.02 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 5120 | 1 | Vulkan0 | none | pp512 @ d131 | 194.68 ± 1.52 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 5120 | 1 | Vulkan0 | none | tg128 @ d131 | 7.71 ± 0.01 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 16 | 1 | Vulkan0 | none | pp512 @ d131 | 31.92 ± 0.20 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 16 | 1 | Vulkan0 | none | tg128 @ d131 | 7.71 ± 0.01 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 2816 | 1 | Vulkan0 | none | pp512 @ d131 | 194.82 ± 1.37 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 2816 | 1 | Vulkan0 | none | tg128 @ d131 | 7.70 ± 0.00 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 3072 | 1 | Vulkan0 | none | pp512 @ d131 | 194.67 ± 0.52 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 3072 | 1 | Vulkan0 | none | tg128 @ d131 | 7.72 ± 0.01 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 3328 | 1 | Vulkan0 | none | pp512 @ d131 | 195.96 ± 0.03 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 3328 | 1 | Vulkan0 | none | tg128 @ d131 | 7.70 ± 0.03 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 3584 | 1 | Vulkan0 | none | pp512 @ d131 | 195.83 ± 0.18 |
| qwen35 27B Q8_0 | 26.62 GiB | 26.90 B | Vulkan | 256 | 3584 | 1 | Vulkan0 | none | tg128 @ d131 | 7.72 ± 0.00 |
| qwen4exp A3B Q3_K - Medium | 83.80 GiB | 176.94 B | Vulkan | 256 | 32 | 1 | Vulkan0 | none | pp512 @ d131 | 100.52 ± 0.82 |
| qwen4exp A3B Q3_K - Medium | 83.80 GiB | 176.94 B | Vulkan | 256 | 32 | 1 | Vulkan0 | none | tg128 @ d131 | 26.70 ± 0.21 |
| qwen4exp A3B Q3_K - Medium | 83.80 GiB | 176.94 B | Vulkan | 256 | 32 | 0 | Vulkan0 | none | pp512 @ d131 | 100.85 ± 0.43 |
| qwen4exp A3B Q3_K - Medium | 83.80 GiB | 176.94 B | Vulkan | 256 | 32 | 0 | Vulkan0 | none | tg128 @ d131 | 26.44 ± 0.03 |
| qwen4exp A3B Q3_K - Medium | 83.80 GiB | 176.94 B | Vulkan | 256 | 3328 | 1 | Vulkan0 | none | pp512 @ d131 | 324.80 ± 0.56 |
| qwen4exp A3B Q3_K - Medium | 83.80 GiB | 176.94 B | Vulkan | 256 | 3328 | 1 | Vulkan0 | none | tg128 @ d131 | 26.72 ± 0.04 |
| qwen4exp A3B Q3_K - Medium | 83.80 GiB | 176.94 B | Vulkan | 256 | 3328 | 0 | Vulkan0 | none | pp512 @ d131 | 318.37 ± 2.97 |
| qwen4exp A3B Q3_K - Medium | 83.80 GiB | 176.94 B | Vulkan | 256 | 3328 | 0 | Vulkan0 | none | tg128 @ d131 | 26.37 ± 0.04 |
this below set, power limit = 140w, TCTL limit = 98C
| model | size | params | backend | ngl | n_ubatch | fa | dev | lm | test | t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | -------: | --: | ------------ | ---------: | --------------: | -------------------: |
| qwen4exp A3B Q3_K - Medium | 83.80 GiB | 176.94 B | ROCm | 256 | 32 | 1 | ROCm0 | none | pp512 @ d131 | 125.43 ± 3.10 |
| qwen4exp A3B Q3_K - Medium | 83.80 GiB | 176.94 B | ROCm | 256 | 32 | 1 | ROCm0 | none | tg128 @ d131 | 19.27 ± 0.19 |
| qwen4exp A3B Q3_K - Medium | 83.80 GiB | 176.94 B | ROCm | 256 | 192 | 1 | ROCm0 | none | pp512 @ d131 | 291.36 ± 1.54 |
| qwen4exp A3B Q3_K - Medium | 83.80 GiB | 176.94 B | ROCm | 256 | 192 | 1 | ROCm0 | none | tg128 @ d131 | 19.34 ± 0.06 |
| qwen4exp A3B Q3_K - Medium | 83.80 GiB | 176.94 B | Vulkan | 256 | 32 | 1 | Vulkan0 | none | pp512 @ d131 | 114.43 ± 0.45 |
| qwen4exp A3B Q3_K - Medium | 83.80 GiB | 176.94 B | Vulkan | 256 | 32 | 1 | Vulkan0 | none | tg128 @ d131 | 27.92 ± 0.01 |
| qwen4exp A3B Q3_K - Medium | 83.80 GiB | 176.94 B | Vulkan | 256 | 32 | 0 | Vulkan0 | none | pp512 @ d131 | 115.14 ± 0.38 |
| qwen4exp A3B Q3_K - Medium | 83.80 GiB | 176.94 B | Vulkan | 256 | 32 | 0 | Vulkan0 | none | tg128 @ d131 | 27.45 ± 0.17 |
| qwen4exp A3B Q3_K - Medium | 83.80 GiB | 176.94 B | Vulkan | 256 | 3328 | 1 | Vulkan0 | none | pp512 @ d131 | 414.25 ± 1.47 |
| qwen4exp A3B Q3_K - Medium | 83.80 GiB | 176.94 B | Vulkan | 256 | 3328 | 1 | Vulkan0 | none | tg128 @ d131 | 27.82 ± 0.01 |
| qwen4exp A3B Q3_K - Medium | 83.80 GiB | 176.94 B | Vulkan | 256 | 3328 | 0 | Vulkan0 | none | pp512 @ d131 | 410.30 ± 2.71 |
| qwen4exp A3B Q3_K - Medium | 83.80 GiB | 176.94 B | Vulkan | 256 | 3328 | 0 | Vulkan0 | none | tg128 @ d131 | 27.49 ± 0.05 |
ROCm TCTL observed peak = 89C / power = 88w
Vulkan observed peak = 73C / power = 115w
vulkan uses more power, and runs cooler ![]()
eg :
./llama-bench --load-mode none -ngl 256 -dev ROCm0 -d 131 -r 3 -fa 1 -ub 32,192 -m / abc.gguf
Qwen 4?
beats me, thats what llama says
(Qwen3.8-Flash-Next-UD-Q3_K_XL.gguf)
| qwen3.8-flash-next UD-IQ4_XS unsloth | 405.19 (pp512 t/s) | 27.47 (tg128 t/s) |
|---|---|---|
| Qwen3.8-27B Q8_0 unsloth, non-MTP | 245.63 | 7.83 |
|---|---|---|
Latest llama.cpp.
Flash: exec llama-server -m /data/models/Qwen3.8-Flash-Next-UD-IQ4_XS/UD-IQ4_XS/Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf
-ngl 999 --jinja -c 65536 -ub 2048 -b 2048
–temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.0
–host 0.0.0.0 --port 8210 “$@”
Dense: exec llama-server -m /data/models/Qwen3.8-27B/Qwen3.8-27B-Q8_0.gguf
-ngl 999 -fa on --jinja -c 32768 -ub 2048 -b 2048
–cache-type-k q8_0 --cache-type-v q8_0
–temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --presence-penalty 0.0
–host 0.0.0.0 --port 8214 “$@”
3 bit seems aggressive, how does the output look?
I ordered a 32gb V100 and will try some layer split stuff that AI suggested when I get back from vacation. I believe this should be able to ~2x the speed on 3.8 flash, worst case probably 1.4x ish. 3.8 flash hosted by Aliyun seems pretty good at coding/algorithm design, but when I ask it to examine the llama.cpp repo code it seems to have trouble searching through it and finding relevant stuff. Probably good enough for local use as backup though.
well it did resolve some fluid dynamics problems
–load-mode none -ngl all --repeat_penalty 1.02 --temp 0.2 -fa 1 --top-p 0.4 --top-k 18 --chat-template-kwargs ‘{“reasoning_effort”: “medium”}’
eg : Ld = (64/Re)*(1-y1)+(0.0032+0.221/Re^0.237)*(y1-y3)+y2*0.11*E^0.25
did a few variants, comparing it with claude and chatgpt. its impressive. i even tried to “emulate” a FW desktop heatsink dissipation.
but i have no way of checking the answers
Okay, I’m using a Lemonade server, and I gave Dense, 27B, Q8 a short prompt, "Create a new pun that will play on grammar ambiguity similar in structure like, “Time flies like an arrow, fruit flies like a banana.” model=Qwen3.8-27B-GGUF-Q8_0, tokens=12510 (in=699, out=11811), ttft=3.281s, tps=13.74 BTW, I’ve never actually been impressed by any of the actual result puns created with this prompt, "Bells beat like hearts, heart beats like bells." but some of the discarded puns in the reasoning are pretty good as idea generators for me to use.