Update #5: Post training yandex/AliceAI-80B-A3B [instruct!] from scratch

Last update for those following: Project in a sentence: An instruct finetune of ALiceAI-80B-A3B-Base capable of agentic work and conversation. I'm creating a shallow distill of qwen 3.8 27b on medium to teach the model chain of thought reasoning and conversation. All training is done locally on 3, 32gb v100s. Additionally, all the training data is being generated locally on said V100s via sftmill. Up to this point I've been doing training runs and live-streaming the progress. Last update explained underfitting and next steps. Training has begun again! I've synthesized about 5M more tokens for the SFT, this time across a much larger general instruct trajectory to try to reduce the underfitting. Dropped the learning rate about 4x over my original LoRA adapter. I'm live streaming training again: geological-estimate-fifth-pct.trycloudflare.com This one should last 12-14 hours, and I plan to run another epoch if this isn't sufficient. Stay tuned! Thanks for following along.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论