Fine-tuning Cactus Needle 2 can match DeepSeek v4 on the specific task
Hey LocalLlama, Henry from Cactus here! When we trained Needle 2, I had a strict rule to not expose the model to any data sample that remotely felt like these benchmarks. It seemed over-the-top, but benchmarks are easy to overfit around, yet struggle in the wild, especially since we quantize to 2-bit. Numbers would look flashy but real-world intent-matching goes to s**t. However, we exposed a playground: cactuscompute.com/needle for users to figure out what works for them, then fine-tune with the python library: github.com/cactus-compute/needle if problem space is constrained. Fine-tuning completely happens on your Mac/PC in minutes. To demonstrate this, when we fine-tuned for each task, Needle 2 surpassed DeepSeek v4 Flash on those specific tasks. Ofc, general-purpose models carry the burden of learning broad linguistic distributions and these results must be interpreted with that in mind. Products running Needle 2 in production almost always does this. We have seen Needle 2 fine-tuned for a wide range of task and we'd love to get everyone's thoughts on this. Please setup good evals to avoid overfitting. We are working on novel Needle fine-tuning algorithms anyway but launching later. Please share reservations, genuine criticism!