I want to share my benchmarks (the hardware is not Framework laptop but mini pc and it is closest I can try that should be the same in terms of performance as top Framework 13 Ryzen AI hx370).
As I have only 32GB RAM on the device I can not really try something like Qwen3-Next-80b but I tried multiple different llms that I use currently often on my pre AI Ryzen framework.
I checked multiple quantization and as you can see it’s even possible to converse with 32b dense model (a bit slow but doable).
In general comparing to my current setup it is usually x2 faster tokens per second and x3 faster prompt tokens process speed.
AMD Ryzen AI hx 370 32gb (NOT ACTUAL Framework laptop)
Fedora 43 live (Gnome)
TLDR; You can expect about x2 performance boost switching from Framework 13 top non AI AMD to hx 370, but for workflows that requires a lot of context in the beginning of the first message - the delay is very noticable and just about 3 times shorter than with previous generation.
llama-bench (vulkan)
| model | size | params | backend | ngl | test | t/s |
|---|---|---|---|---|---|---|
| qwen3 0.6B BF16 | 1.11 GiB | 596.05 M | Vulkan | 99 | pp512 | 1863.63 ± 10.14 |
| qwen3 0.6B BF16 | 1.11 GiB | 596.05 M | Vulkan | 99 | tg128 | 58.44 ± 0.39 |
| qwen3 1.7B Q8_0 | 1.70 GiB | 1.72 B | Vulkan | 99 | pp512 | 1566.35 ± 10.67 |
| qwen3 1.7B Q8_0 | 1.70 GiB | 1.72 B | Vulkan | 99 | tg128 | 40.08 ± 0.48 |
| phi3 14B Q8_0 | 14.51 GiB | 14.66 B | Vulkan | 99 | pp512 | 235.90 ± 0.39 |
| phi3 14B Q8_0 | 14.51 GiB | 14.66 B | Vulkan | 99 | tg128 | 5.38 ± 0.00 |
| mistral3 3B Q4_K - Medium | 1.99 GiB | 3.43 B | Vulkan | 99 | pp512 | 862.70 ± 1.07 |
| mistral3 3B Q4_K - Medium | 1.99 GiB | 3.43 B | Vulkan | 99 | tg128 | 33.03 ± 0.27 |
| phi3 3B Q8_0 | 3.80 GiB | 3.84 B | Vulkan | 99 | pp512 | 759.54 ± 2.64 |
| phi3 3B Q8_0 | 3.80 GiB | 3.84 B | Vulkan | 99 | tg128 | 18.83 ± 0.04 |
| mistral3 8B Q4_K - Medium | 4.83 GiB | 8.49 B | Vulkan | 99 | pp512 | 356.36 ± 0.22 |
| mistral3 8B Q4_K - Medium | 4.83 GiB | 8.49 B | Vulkan | 99 | tg128 | 15.43 ± 0.06 |
| mistral3 14B Q8_0 | 13.37 GiB | 13.51 B | Vulkan | 99 | pp512 | 196.30 ± 1.02 |
| mistral3 14B Q8_0 | 13.37 GiB | 13.51 B | Vulkan | 99 | tg128 | 5.87 ± 0.01 |
| gpt-oss 20B Q4_K - Medium | 10.81 GiB | 20.91 B | Vulkan | 99 | pp512 | 449.16 ± 5.27 |
| gpt-oss 20B Q4_K - Medium | 10.81 GiB | 20.91 B | Vulkan | 99 | tg128 | 31.20 ± 0.13 |
| llama 9B Q4_K - EuroLLM | 5.20 GiB | 9.15 B | Vulkan | 99 | pp512 | 341.60 ± 0.86 |
| llama 9B Q4_K - EuroLLM | 5.20 GiB | 9.15 B | Vulkan | 99 | tg128 | 14.42 ± 0.01 |
| mistral3 24B Q4_0 Devstral S 2 | 12.56 GiB | 23.57 B | Vulkan | 99 | pp512 | 138.50 ± 0.23 |
| mistral3 24B Q4_0 Devstral S 2 | 12.56 GiB | 23.57 B | Vulkan | 99 | tg128 | 6.28 ± 0.01 |
| deepseek2 16B Q4_K - Medium | 9.65 GiB | 15.71 B | Vulkan | 99 | pp512 | 614.76 ± 4.81 |
| deepseek2 16B Q4_K - Medium | 9.65 GiB | 15.71 B | Vulkan | 99 | tg128 | 39.80 ± 0.09 |
| EuroMoE 2.6Ba0.6b F16 | 4.87 GiB | 2.61 B | Vulkan | 99 | pp512 | 1774.07 ± 75.99 |
| EuroMoE 2.6Ba0.6b F16 | 4.87 GiB | 2.61 B | Vulkan | 99 | tg128 | 68.07 ± 0.57 |
| gemma3 4B Q8_0 | 3.84 GiB | 3.88 B | Vulkan | 99 | pp512 | 826.68 ± 0.95 |
| gemma3 4B Q8_0 | 3.84 GiB | 3.88 B | Vulkan | 99 | tg128 | 19.00 ± 0.04 |
| apertus 8B Q4_K - Medium | 4.70 GiB | 8.05 B | Vulkan | 99 | pp512 | 378.87 ± 0.66 |
| apertus 8B Q4_K - Medium | 4.70 GiB | 8.05 B | Vulkan | 99 | tg128 | 16.37 ± 0.03 |
| llama 13B Q4_K - Medium | 6.96 GiB | 12.25 B | Vulkan | 99 | pp512 | 241.41 ± 0.40 |
| llama 13B Q4_K - Medium | 6.96 GiB | 12.25 B | Vulkan | 99 | tg128 | 10.98 ± 0.04 |
| qwen2 14B Q4_K - Medium | 8.37 GiB | 14.77 B | Vulkan | 99 | pp512 | 198.06 ± 0.41 |
| qwen2 14B Q4_K - Medium | 8.37 GiB | 14.77 B | Vulkan | 99 | tg128 | 9.04 ± 0.00 |
llama-bench (cpu)
| model | size | params | backend | threads | test | t/s |
|---|---|---|---|---|---|---|
| qwen3 0.6B BF16 | 1.11 GiB | 596.05 M | CPU | 12 | pp512 | 877.56 ± 1.24 |
| qwen3 0.6B BF16 | 1.11 GiB | 596.05 M | CPU | 12 | tg128 | 55.54 ± 0.04 |
| qwen3 1.7B Q8_0 | 1.70 GiB | 1.72 B | CPU | 12 | pp512 | 354.10 ± 0.53 |
| qwen3 1.7B Q8_0 | 1.70 GiB | 1.72 B | CPU | 12 | tg128 | 36.88 ± 0.03 |
| phi3 14B Q8_0 | 14.51 GiB | 14.66 B | CPU | 12 | pp512 | 40.87 ± 0.58 |
| phi3 14B Q8_0 | 14.51 GiB | 14.66 B | CPU | 12 | tg128 | 5.11 ± 0.00 |
| mistral3 3B Q4_K - Medium | 1.99 GiB | 3.43 B | CPU | 12 | pp512 | 190.58 ± 5.36 |
| mistral3 3B Q4_K - Medium | 1.99 GiB | 3.43 B | CPU | 12 | tg128 | 31.77 ± 0.55 |
| phi3 3B Q8_0 | 3.80 GiB | 3.84 B | CPU | 12 | pp512 | 163.31 ± 10.99 |
| phi3 3B Q8_0 | 3.80 GiB | 3.84 B | CPU | 12 | tg128 | 17.97 ± 0.05 |
| mistral3 8B Q4_K - Medium | 4.83 GiB | 8.49 B | CPU | 12 | pp512 | 82.04 ± 2.28 |
| mistral3 8B Q4_K - Medium | 4.83 GiB | 8.49 B | CPU | 12 | tg128 | 14.75 ± 0.03 |
| mistral3 14B Q8_0 | 13.37 GiB | 13.51 B | CPU | 12 | pp512 | 45.60 ± 0.57 |
| mistral3 14B Q8_0 | 13.37 GiB | 13.51 B | CPU | 12 | tg128 | 5.56 ± 0.01 |
| gpt-oss 20B Q4_K - Medium | 10.81 GiB | 20.91 B | CPU | 12 | pp512 | 99.13 ± 0.14 |
| gpt-oss 20B Q4_K - Medium | 10.81 GiB | 20.91 B | CPU | 12 | tg128 | 26.78 ± 0.02 |
| llama 9B Q4_K - EuroLLM | 5.20 GiB | 9.15 B | CPU | 12 | pp512 | 74.93 ± 2.21 |
| llama 9B Q4_K - EuroLLM | 5.20 GiB | 9.15 B | CPU | 12 | tg128 | 13.55 ± 0.00 |
| mistral3 24B Q4_0 Devstral S 2 | 12.56 GiB | 23.57 B | CPU | 12 | pp512 | 32.60 ± 0.06 |
| mistral3 24B Q4_0 Devstral S 2 | 12.56 GiB | 23.57 B | CPU | 12 | tg128 | 5.78 ± 0.01 |
| deepseek2 16B Q4_K - Medium | 9.65 GiB | 15.71 B | CPU | 12 | pp512 | 179.17 ± 7.12 |
| deepseek2 16B Q4_K - Medium | 9.65 GiB | 15.71 B | CPU | 12 | tg128 | 36.28 ± 0.05 |
| EuroMoE 2.6Ba0.6b F16 | 4.87 GiB | 2.61 B | CPU | 12 | pp512 | 542.98 ± 1.27 |
| EuroMoE 2.6Ba0.6b F16 | 4.87 GiB | 2.61 B | CPU | 12 | tg128 | 63.30 ± 0.16 |
| gemma3 4B Q8_0 | 3.84 GiB | 3.88 B | CPU | 12 | pp512 | 168.55 ± 6.55 |
| gemma3 4B Q8_0 | 3.84 GiB | 3.88 B | CPU | 12 | tg128 | 15.80 ± 0.00 |
| apertus 8B Q4_K - Medium | 4.70 GiB | 8.05 B | CPU | 12 | pp512 | 79.43 ± 1.62 |
| apertus 8B Q4_K - Medium | 4.70 GiB | 8.05 B | CPU | 12 | tg128 | 14.22 ± 0.00 |
| llama 13B Q4_K - Medium | 6.96 GiB | 12.25 B | CPU | 12 | pp512 | 57.09 ± 0.67 |
| llama 13B Q4_K - Medium | 6.96 GiB | 12.25 B | CPU | 12 | tg128 | 10.24 ± 0.00 |
| qwen2 14B Q4_K - Medium | 8.37 GiB | 14.77 B | CPU | 12 | pp512 | 46.60 ± 0.51 |
| qwen2 14B Q4_K - Medium | 8.37 GiB | 14.77 B | CPU | 12 | tg128 | 8.53 ± 0.01 |
llama-cli
Backend: vulcan
prompt: Hello! Who are you?
Context: 512
build : b7988-c03a5a46f
model : Qwen3-32B-Q4_K_M.gguf
[ Prompt: 11.9 t/s | Generation: 3.9 t/s ]