@lain is that even related? They already know which "subnetwork" is "tuned to similar performance", by way of expert activation. You can isolate which experts are activated for coding, no?
They just took those and put them in a smaller model to reduce overhead...
They just took those and put them in a smaller model to reduce overhead...
Replying to @WandererUber@poa.st
@WandererUber every expert in a3b is 3b. so a 2.6b model can't just take the experts. they must have derived more fundamental experts / pruned the useless nodes to achieve this.
Replying to @WandererUber@poa.st
@WandererUber ah, i misread it! very interesting idea!


