← Back to post

Edit history

Most recent

Be aware that you pay a significant performance penalty for going over occulink.

The hit modest for fully offloaded dense models (like Qwen 27B), but dramatic for hybrid inference of big MoEs.

Even my old 3090 got a noticeable performance gain going from a PCIe 3.0 x16 riser to a PCIe 4.0 one.

Edited

Be aware that you pay a significant performance penalty for going over occulink.

The hit modest for fully offloaded dense models (like Qwen 27B), but dramatic for hybrid inference of big MoEs.

Even my old 3090 got a big jump going from a PCIe 3.0 x16 riser to a PCIe 4.0 one.

Edited

Be aware that you pay a significant performance penalty for going over occulink.

The hit modest for fully offloaded dense models (like Qwen 27B), but dramatic for hybrid inference of big MoEs.

Even my old 3090 got a big prompt-processing jump going from a PCIe 3.0 x16 riser to a PCIe 4.0 one.

Original

Be aware that you pay a significant performance penalty for going over occulink.

Even my old 3090 got a huge jump going from a PCIe 3.0 x16 riser to a PCIe 4.0 one.