Replying to @⁨leanleft@lemmy.ml⁩

Even “big” open source models like DSV4 and Ling/Ring are very efficient. They’re big, but (seemingly) sparser than US models, so they’re cheap.

They run surprisingly well with hybrid CPU+GPU inference on desktops. And thats not even getting into the efficient attention mechanisms.

I can run DSV4 Flash, barely quantized, with ~1M context on my Ryzen desktop at ~11 tokens/s. If you told me that two years ago, I would not have believed you.

Edited ⁨⁨Aug⁩ ⁨17⁩, ⁨2026⁩, ⁨22:28⁩⁩en