I trained a 102M recursive BitNet-v2 model from scratch: 64K context, trained on less than 5B tokens

DISCLAIMER: the post and the model card was made with the assist of AI. Hiya, I’m releasing Recursive BitNet N-Gram 102M, a small experiment combining ternary weights, shared transformer layers, and hashed n-gram embeddings, trained with a whooping budget of 100€ Weights, inference code, model card, and evaluation results on Hugging Face ( huggingface.co/n00nehere/recursive-bitnet-ngram-102M-64K-instruct-preview ) All model weights started from random initialization. I reused the Cosmo2 tokenizer with its 49,152-token vocabulary. The architecture is a decoder-only transformer with a few additions: • 102.28M unique parameters, hidden size 1,024, grouped-query attention with 16 query heads and 4 KV heads, and squared-ReLU feed-forward layers. • Six transformer blocks run twice, giving twelve effective layer applications. Both passes share weights, with separate KV caches at each effective depth. • Ternary BitLinear projections use scaled weights from {-1, 0, +1} during the forward pass. Training retains FP32 master weights and BF16 activations. • Causal token 2-, 3-, and 4-gram embeddings are hashed into small lookup tables and added to the input embeddings. This branch adds about 1.6M parameters. • 65,536-token training context during the continuation phase. Training happened in two stages: first the backbone, then continued training after adding the n-gram branch. Stage Hardware Context Input tokens processed ━━━━━━━━━━━━━━━━━━━━━━━━ Scratch backbone 4× NVIDIA B300 2K → 4K 3.487B ──────────────────────── N-gram continuation 1× NVIDIA B300 64K 1.216B ──────────────────────── Total 4.703B Both stages used AdamW with an effective batch of 131,072 input tokens. The 4.703B counter includes repeated examples, masked prompts, and padding. About 2.781B positions contributed supervised targets. Roughly one in eight continuation updates used synthetic memory episodes spanning distances up to 60K tokens. Source manifests and dataset license details are in the model card. For benchmarks, I used lm-eval 0.4.12, zero-shot, full available splits, plain-text prompts, and a 2,048-token scoring context: 20,465 questions overall. Benchmark Metric Score ━━━━━━━━━━━━━━━ ARC-Easy acc 40.11% ─────────────── ARC-Challenge acc_norm 23.63% ─────────────── PIQA acc 57.62% ─────────────── WinoGrande acc 51.22% ─────────────── OpenBookQA acc_norm 25.60% ─────────────── BoolQ acc 58.65% ─────────────── HellaSwag acc_norm 27.27% The unweighted average is 40.59%. The released checkpoint was selected by held-out loss across ten saved snapshots. A couple of interesting ablations: one, two, and three recursive passes scored 39.19%, 40.59%, and 39.67%, respectively. The release includes 409.14MB FP32 master weights and a 124.24MB packed export. Packing stores five ternary values per byte, while embeddings and other tensors remain floating point. The benchmark table uses the reference weights. The packed export passed numerical and generation checks. This is an experimental preview with modest results but still better than what i initially expected. All feedbacks are more than welcome :)

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论