Cleo: trying to fit full analyst behavior in a 2B model [P]
Hello all! Half of all industrial "chatbots" are just text-to-SQL models in a trenchcoat (and the other half RAG!). I wanted to explore just how small you could make these models if you trained, evaluated, and ran inference in the exact same structured harness, leading to Cleo: a Qwen3.5-2B-Base finetune. Currently, some features of cleo that are only possible/useful in a unified hardel are: Training on the exact same gather, repair, and answer contract it uses at inference time Searching over candidate que
评论
?
参与讨论