PPlamenu
HomeTrendingLive feedsPeopleGroupsRulesStaff
Sign in
PPlamenu
HomeTrendingLive feedsPeopleGroupsRulesStaff
Sign in
Wanderer atop the sea of clouds or whatever@WandererUber@poa.st
⁨3⁩d
https://huggingface.co/Akahsizrr/fuse-1-Lite

:Pogey:
1000
Open original page
BoostsQuotesFavs
lain@lain@lain.com
⁨3⁩d

Replying to @⁨WandererUber@poa.st⁩

@WandererUber https://en.wikipedia.org/wiki/Lottery_ticket_hypothesis
en.wikipedia.orgLottery ticket hypothesis - Wikipedia
1000
Open original page
BoostsQuotesFavs
Wanderer atop the sea of clouds or whatever@WandererUber@poa.st
⁨3⁩d

Replying to @⁨lain@lain.com⁩

@lain is that even related? They already know which "subnetwork" is "tuned to similar performance", by way of expert activation. You can isolate which experts are activated for coding, no?
They just took those and put them in a smaller model to reduce overhead...
1000
Open original page
BoostsQuotesFavs
lain@lain@lain.com
⁨3⁩d

Replying to @⁨WandererUber@poa.st⁩

@WandererUber every expert in a3b is 3b. so a 2.6b model can't just take the experts. they must have derived more fundamental experts / pruned the useless nodes to achieve this.
⁨Aug⁩ ⁨12⁩, ⁨2026⁩, ⁨14:02⁩
1000
Open original page
BoostsQuotesFavs
Wanderer atop the sea of clouds or whatever@WandererUber@poa.st
⁨3⁩d

Replying to @⁨lain@lain.com⁩

@lain
1000
Open original page
BoostsQuotesFavs
lain@lain@lain.com
⁨3⁩d

Replying to @⁨WandererUber@poa.st⁩

@WandererUber ah, i misread it! very interesting idea!
0000
Open original page
BoostsQuotesFavs