The First ChineseBabyLM Challenge: training data-efficient and cognitively plausible language models for Chinese

arXiv:2607.10745v1 Announce Type: new Abstract: This paper describes the first ChineseBabyLM challenge, which will be held in the 2026 NLPCC conference. The challenge calls for researchers to train language models from scratch with 100 million Chinese tokens and evaluates the models on 3 tracks of tasks: NLU, cognitive alignment and Hanzi knowledge. There is no restriction on tokenizer, model architecture and the number of training epochs. Details of the challenge can be found inchinese-babylm.github.io.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论