Bonsai? Or whatever it’s called? It’s a con, so far; it’s not better than smaller models quantized to 3-4 bits.
I love, love the idea of bitnet, but it only seems to work with models trained from scratch, which no one has done at scale yet.
Bonsai? Or whatever it’s called? It’s a con, so far; it’s not better than smaller models quantized to 3-4 bits.
I love, love the idea of bitnet, but it only seems to work with models trained from scratch, which no one has done at scale yet.
Bonsai? Or whatever it’s called? It’s a con, so far; it’s not better than smaller models quantized to 3-4 bits.