GLM 5.2 on Dual Strix Halo (256GB): Worth it?

This video covers my experiments running the new GLM 5.2 on a dual Strix Halo setup (2 Framework Desktops). I cover what quantizations you can run and the speed you get in terms of token generation and prompt processing, comparing it with DeepSeek V4 Flash. I also ran SWE Bench verified mini to check the actual performance at coding tasks.

6 Likes

Betteridge’s Law strikes again! Answer: no, … not really …

It does work! That is very impressive @kyuz0. But wow, so slow.

Thanks for showing us, hope you’ll mention your future work in these forums too.

1 Like

Thank you Donato for very interesting and carefully conducted experiments!

On a subject, though… No, not very practical, at least for current generation of models. I really hope that next year we’d have heterogeneous models that would combine speculative/diffusion drafter coupled with autoregressive validator, that would bring decent inference speeds to a consumer grade infrastructure.

It’s already a thing, but just not a mainstream one.