← Back to post

Edit history

Most recent

Even “big” open source models like DSV4 and Ling/Ring are very efficient. They’re big, but (seemingly) sparser than US models, so they’re cheap.

They run surprisingly well with hybrid CPU+GPU inference on desktops. And thats not even getting into the efficient attention mechanisms.

I can run DSV4 Flash, barely quantized, with ~1M context on my Ryzen desktop at ~11 tokens/s. If you told me that two years ago, I would not have believed you.

Edited

Even “big” open source models like DSV4 are very efficient. They’re big, but (seemingly) sparser than US models, so they’re cheap.

They run surprisingly well with hybrid CPU+GPU inference on desktops. And thats not even getting into the efficient attention mechanisms.

I can run DSV4 Flash, barely quantized, with ~1M context on my Ryzen desktop at ~11 tokens/s. If you told me that two years ago, I would not have believed you.

Edited

Even “big” open source models like DSV4 are very efficient. They’re big, but (seemingly) sparser than US models, so they’re cheap.

They run surprisingly well with hybrid CPU+GPU inference on desktops.

Original

Even “big” open source models like DSV4 are very efficient. They’re big, but (seemingly) sparser than US models, so they’re cheap and even run surprisingly well with hybrid CPU+GPU inference on a desktop.