Models using low-rank bottlenecks for the embeddings or the output head. See https://www.gilesthomas.com/2026/10/low-rank-vocab-matrices
-
gpjt/8xa100m40-lore-1-baseline-disabled
Text Generation • 0.2B • Updated • 138 -
gpjt/8xa100m40-lore-10-both-corrected-smart-init
Text Generation • 0.1B • Updated • 135 -
gpjt/8xa100m40-lore-11-output-only-corrected-smart-init
Text Generation • 0.1B • Updated • 146 -
gpjt/8xa100m40-lore-2-both-ordinary-init
Text Generation • 0.1B • Updated • 159