Framework 13 + Ryzen AI + Linux Distro + LLM

I want to share my benchmarks (the hardware is not Framework laptop but mini pc and it is closest I can try that should be the same in terms of performance as top Framework 13 Ryzen AI hx370).

As I have only 32GB RAM on the device I can not really try something like Qwen3-Next-80b but I tried multiple different llms that I use currently often on my pre AI Ryzen framework.

I checked multiple quantization and as you can see it’s even possible to converse with 32b dense model (a bit slow but doable).

In general comparing to my current setup it is usually x2 faster tokens per second and x3 faster prompt tokens process speed.

AMD Ryzen AI hx 370 32gb (NOT ACTUAL Framework laptop)
Fedora 43 live (Gnome)

TLDR; You can expect about x2 performance boost switching from Framework 13 top non AI AMD to hx 370, but for workflows that requires a lot of context in the beginning of the first message - the delay is very noticable and just about 3 times shorter than with previous generation.

llama-bench (vulkan)

model size params backend ngl test t/s
qwen3 0.6B BF16 1.11 GiB 596.05 M Vulkan 99 pp512 1863.63 ± 10.14
qwen3 0.6B BF16 1.11 GiB 596.05 M Vulkan 99 tg128 58.44 ± 0.39
qwen3 1.7B Q8_0 1.70 GiB 1.72 B Vulkan 99 pp512 1566.35 ± 10.67
qwen3 1.7B Q8_0 1.70 GiB 1.72 B Vulkan 99 tg128 40.08 ± 0.48
phi3 14B Q8_0 14.51 GiB 14.66 B Vulkan 99 pp512 235.90 ± 0.39
phi3 14B Q8_0 14.51 GiB 14.66 B Vulkan 99 tg128 5.38 ± 0.00
mistral3 3B Q4_K - Medium 1.99 GiB 3.43 B Vulkan 99 pp512 862.70 ± 1.07
mistral3 3B Q4_K - Medium 1.99 GiB 3.43 B Vulkan 99 tg128 33.03 ± 0.27
phi3 3B Q8_0 3.80 GiB 3.84 B Vulkan 99 pp512 759.54 ± 2.64
phi3 3B Q8_0 3.80 GiB 3.84 B Vulkan 99 tg128 18.83 ± 0.04
mistral3 8B Q4_K - Medium 4.83 GiB 8.49 B Vulkan 99 pp512 356.36 ± 0.22
mistral3 8B Q4_K - Medium 4.83 GiB 8.49 B Vulkan 99 tg128 15.43 ± 0.06
mistral3 14B Q8_0 13.37 GiB 13.51 B Vulkan 99 pp512 196.30 ± 1.02
mistral3 14B Q8_0 13.37 GiB 13.51 B Vulkan 99 tg128 5.87 ± 0.01
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B Vulkan 99 pp512 449.16 ± 5.27
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B Vulkan 99 tg128 31.20 ± 0.13
llama 9B Q4_K - EuroLLM 5.20 GiB 9.15 B Vulkan 99 pp512 341.60 ± 0.86
llama 9B Q4_K - EuroLLM 5.20 GiB 9.15 B Vulkan 99 tg128 14.42 ± 0.01
mistral3 24B Q4_0 Devstral S 2 12.56 GiB 23.57 B Vulkan 99 pp512 138.50 ± 0.23
mistral3 24B Q4_0 Devstral S 2 12.56 GiB 23.57 B Vulkan 99 tg128 6.28 ± 0.01
deepseek2 16B Q4_K - Medium 9.65 GiB 15.71 B Vulkan 99 pp512 614.76 ± 4.81
deepseek2 16B Q4_K - Medium 9.65 GiB 15.71 B Vulkan 99 tg128 39.80 ± 0.09
EuroMoE 2.6Ba0.6b F16 4.87 GiB 2.61 B Vulkan 99 pp512 1774.07 ± 75.99
EuroMoE 2.6Ba0.6b F16 4.87 GiB 2.61 B Vulkan 99 tg128 68.07 ± 0.57
gemma3 4B Q8_0 3.84 GiB 3.88 B Vulkan 99 pp512 826.68 ± 0.95
gemma3 4B Q8_0 3.84 GiB 3.88 B Vulkan 99 tg128 19.00 ± 0.04
apertus 8B Q4_K - Medium 4.70 GiB 8.05 B Vulkan 99 pp512 378.87 ± 0.66
apertus 8B Q4_K - Medium 4.70 GiB 8.05 B Vulkan 99 tg128 16.37 ± 0.03
llama 13B Q4_K - Medium 6.96 GiB 12.25 B Vulkan 99 pp512 241.41 ± 0.40
llama 13B Q4_K - Medium 6.96 GiB 12.25 B Vulkan 99 tg128 10.98 ± 0.04
qwen2 14B Q4_K - Medium 8.37 GiB 14.77 B Vulkan 99 pp512 198.06 ± 0.41
qwen2 14B Q4_K - Medium 8.37 GiB 14.77 B Vulkan 99 tg128 9.04 ± 0.00

llama-bench (cpu)

model size params backend threads test t/s
qwen3 0.6B BF16 1.11 GiB 596.05 M CPU 12 pp512 877.56 ± 1.24
qwen3 0.6B BF16 1.11 GiB 596.05 M CPU 12 tg128 55.54 ± 0.04
qwen3 1.7B Q8_0 1.70 GiB 1.72 B CPU 12 pp512 354.10 ± 0.53
qwen3 1.7B Q8_0 1.70 GiB 1.72 B CPU 12 tg128 36.88 ± 0.03
phi3 14B Q8_0 14.51 GiB 14.66 B CPU 12 pp512 40.87 ± 0.58
phi3 14B Q8_0 14.51 GiB 14.66 B CPU 12 tg128 5.11 ± 0.00
mistral3 3B Q4_K - Medium 1.99 GiB 3.43 B CPU 12 pp512 190.58 ± 5.36
mistral3 3B Q4_K - Medium 1.99 GiB 3.43 B CPU 12 tg128 31.77 ± 0.55
phi3 3B Q8_0 3.80 GiB 3.84 B CPU 12 pp512 163.31 ± 10.99
phi3 3B Q8_0 3.80 GiB 3.84 B CPU 12 tg128 17.97 ± 0.05
mistral3 8B Q4_K - Medium 4.83 GiB 8.49 B CPU 12 pp512 82.04 ± 2.28
mistral3 8B Q4_K - Medium 4.83 GiB 8.49 B CPU 12 tg128 14.75 ± 0.03
mistral3 14B Q8_0 13.37 GiB 13.51 B CPU 12 pp512 45.60 ± 0.57
mistral3 14B Q8_0 13.37 GiB 13.51 B CPU 12 tg128 5.56 ± 0.01
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B CPU 12 pp512 99.13 ± 0.14
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B CPU 12 tg128 26.78 ± 0.02
llama 9B Q4_K - EuroLLM 5.20 GiB 9.15 B CPU 12 pp512 74.93 ± 2.21
llama 9B Q4_K - EuroLLM 5.20 GiB 9.15 B CPU 12 tg128 13.55 ± 0.00
mistral3 24B Q4_0 Devstral S 2 12.56 GiB 23.57 B CPU 12 pp512 32.60 ± 0.06
mistral3 24B Q4_0 Devstral S 2 12.56 GiB 23.57 B CPU 12 tg128 5.78 ± 0.01
deepseek2 16B Q4_K - Medium 9.65 GiB 15.71 B CPU 12 pp512 179.17 ± 7.12
deepseek2 16B Q4_K - Medium 9.65 GiB 15.71 B CPU 12 tg128 36.28 ± 0.05
EuroMoE 2.6Ba0.6b F16 4.87 GiB 2.61 B CPU 12 pp512 542.98 ± 1.27
EuroMoE 2.6Ba0.6b F16 4.87 GiB 2.61 B CPU 12 tg128 63.30 ± 0.16
gemma3 4B Q8_0 3.84 GiB 3.88 B CPU 12 pp512 168.55 ± 6.55
gemma3 4B Q8_0 3.84 GiB 3.88 B CPU 12 tg128 15.80 ± 0.00
apertus 8B Q4_K - Medium 4.70 GiB 8.05 B CPU 12 pp512 79.43 ± 1.62
apertus 8B Q4_K - Medium 4.70 GiB 8.05 B CPU 12 tg128 14.22 ± 0.00
llama 13B Q4_K - Medium 6.96 GiB 12.25 B CPU 12 pp512 57.09 ± 0.67
llama 13B Q4_K - Medium 6.96 GiB 12.25 B CPU 12 tg128 10.24 ± 0.00
qwen2 14B Q4_K - Medium 8.37 GiB 14.77 B CPU 12 pp512 46.60 ± 0.51
qwen2 14B Q4_K - Medium 8.37 GiB 14.77 B CPU 12 tg128 8.53 ± 0.01

llama-cli

Backend: vulcan
prompt: Hello! Who are you?
Context: 512
build : b7988-c03a5a46f
model : Qwen3-32B-Q4_K_M.gguf

[ Prompt: 11.9 t/s | Generation: 3.9 t/s ]