PPlamenu
HomeTrendingLive feedsPeopleGroupsRulesStaff
Sign in
PPlamenu
HomeTrendingLive feedsPeopleGroupsRulesStaff
Sign in

posted in Technology

eicker@eicker@lemmy.world
⁨8⁩d

OpenAI Hoarding Tens Of Thousands Of Apple Mac mini And Mac Studio Devices, As ASUS And MSI Burn Through Their Entire First Batch Of NVIDIA RTX Spark Chip And Beg For More.

wccftech.com/openai-hoarding-tens-of-thousands-of-apple-mac-mini-and-mac-studio-devices-as-asus-and-msi-burn-through-their-entire-first-batch-of-nvidia-rtx-spark-chip-and-beg-for-more/
1510331
Open original page
BoostsQuotesFavs
deleted@deleted@lemmy.world
edited⁨8⁩d

Replying to @⁨eicker@lemmy.world⁩

Local 27b models are good enough for most tasks.

Can’t wait to buy one of these from Ebay for 10% of the price next year.

4000
Open original page
BoostsQuotesFavs
Lydia_K@Lydia_K@lemmy.world
⁨8⁩d

Replying to @⁨deleted@lemmy.world⁩

github.com/…/atomic-llama-cpp-turboquant

I’m running gwen 3.6 with 131k context window on a 3090, it’s fast enough and about as good as pay to play Claude at work.

2000
Open original page
BoostsQuotesFavs
ArchAengelus@ArchAengelus@lemmy.dbzer0.com
⁨8⁩d

Replying to @⁨Lydia_K@lemmy.world⁩

Upgrade that to 3.8 as soon as your hardware allows (and your use case makes sense). 3.8 is quite a bit more rational.

⁨Aug⁩ ⁨31⁩, ⁨2026⁩, ⁨15:03⁩en
1000
Open original page
BoostsQuotesFavs
Lydia_K@Lydia_K@lemmy.world
⁨8⁩d

Replying to @⁨ArchAengelus@lemmy.dbzer0.com⁩

I plan to once there is a version with turboquant and MTP as that huge context window is key.

0000
Open original page
BoostsQuotesFavs