← Back to post

Edit history

Most recent

First of all, I mean zero offense with any purchase decision. A 5090 is very good.

…But if I were paying that kind of money, I’d probably get a 4090 and a new motherboard/CPU instead. Maybe a used DRR4 threadripper system.

Hybrid (CPU + GPU) inference is where it’s at these days. It opens up a whole world of huge MoE models, whereas on an 5090 you are stuck with Qwen 27B.

Having a fast CPU, with lots of RAM channels, with full PCIe bandwidth is much more important for that than having a 5090, where a 3090 or 4090 will get the job done.

Even if pure speed is your primary concern, you can tune a sparse 120B model (like Laguna) to run almost as fast as Qwen 27B on a 5090, and get at-least-good results.

It’s more finicky and involved, though. For sure.

Running an LLM on a 5090 is a task. Hybrid CPU + GPU inference is a hobby.

Edited

First of all, I mean zero offense with any purchase decision. A 5090 is very good.

…But if I were paying that kind of money, I’d probably get a 4090 and a new motherboard/CPU instead. Maybe a used DRR4 threadripper system.

Hybrid (CPU + GPU) inference is where it’s at these days. It opens up a whole world of huge MoE models, whereas on an 5090 you are stuck with Qwen 27B.

Having a fast CPU, with lots of RAM channels, with full PCIe bandwidth is much more important for that than having a 5090, where a 3090 or 4090 will get the job done.

Even if pure speed is your primary concern, you can tune a sparse 120B model (like Laguna) to run almost as fast as Qwen 27B on a 5090, and get at-lasts-good results.

It’s more finicky and involved, though. For sure. Running an LLM on a 5090 is a task, hybrid CPU + GPU inference is a hobby.

Edited

First of all, I mean zero offense with any purchase decision. A 5090 is very good.

…But if I were paying that kind of money, I’d probably get a 4090 and a new motherboard/CPU instead. Maybe a used DRR4 threadripper system.

Hybrid (CPU + GPU) inference is where it’s at these days. It opens up a whole world of huge MoE models, whereas on an 5090 you are stuck with Qwen 27B.

Having a fast CPU, with lots of RAM channels, with full PCIe bandwidth is much more important for that than having a 5090 instead of a 4090, or something.

Even if pure speed is your primary concern, you can tune a sparse 120B model (like Laguna) to run almost as fast as Qwen 27B on a 5090, and get at-lasts-good results.

It’s more finicky and involved, though. For sure. Running an LLM on a 5090 is a task, hybrid CPU + GPU inference is a hobby.

Edited

First of all, I mean zero offense with any purchase decision. A 5090 is very good.

…But if I were paying that kind of money, I’d probably get a 4090 and a new motherboard/CPU instead. Maybe a used DRR4 threadripper system.

Hybrid (CPU + GPU) inference is where it’s at these days. It opens up a whole world of huge MoE models, whereas on an 5090 you are stuck with Qwen 27B.

Having a fast CPU, with lots of RAM channels, with full PCIe bandwidth is much more important for that than having a 5090 instead of a 4090, or something.

Edited

First of all, I mean zero offense with any purchase decision. A 5090 is very good.

…But if I were paying that kind of money, I’d probably get a 4090 and a new motherboard/CPU instead. Maybe a used DRR4 threadripper system.

Hybrid (CPU + GPU) inference is where it’s at these days. It opens up a whole world of huge MoE models. And having a gast CPU with full PCIe bandwidth is much more important for that than having a 5090 instead of a 4090 or something.

Original

First of all, I mean zero offense with any purchase decision. A 5090 is very good.

…But if I were paying that kind of money, I’d probably get a 4090 and a new motherboard/CPU instead. Maybe a used DRR4 threadripper system.

Hybrid (CPU + GPU) inference is where it’s at these days. It opens up a whole world of huge MoE models. And having a big CPU is much more important for that than having a 5090 instead of a 4090 or something.