story time: tiramisu (ft. quantization)
when I finished training my 10M gpt model, the weights were 40MB.
i needed it to load in the browser.
so i cut it down to 11MB.
now it loads in seconds. no retraining necessary
here's how: 🧵
story time: tiramisu (ft. matmul)
my matrix multiply took 14 minutes to train a neural network.
i was able to cut it down to 1 min 40.
this is what the naive implementation looks like 🧵
10.8M parameters, 6 transformer layers, 8 attention heads, 512-dim embeddings, gpt-2 architecture.
trained on free T4 GPU via @kaggle (ty </3)
C++ engine is compiled directly to WebAssembly via Emscripten.
it's not coherent, but it's trying really really hard.
"I will hom?"