Agentic behavior

Since around a year ago, my business built our own agent software running on our FreeBSD servers. We selected to use the Anthropic models through their API and kept adding new models as they came out. The agent could route different types of task to different models. For basic tasks like reading, minimal text editing, and similar, it mostly selected Haiku. For planning, it used Opus, and later Fable. For larger workloads, it used Sonnet. Roughly: · Easy tasks: Haiku (with a 25% boost to Sonnet) · Ordinary tasks: Sonnet · Complex tasks: Opus / Fable It was fully functional. We ran it like this all the time, only adding support for new models and occasionally extending our tools. But the API cost was kind of skyrocketing, and I have to say, Anthropic models are pretty bad for agentic work, especially low-level C and assembly. Since a few weeks ago, we tried to replace the Anthropic API with DeepSeek, not for everything, it was more like a rest to get a better idea how the DeepSeek models are handling our type of work. We have moved more and more to run via the DeepSeek models and right now, it almost feels like we’re running and using the models for free. It costs almost nothing. DeepSeek’s agentic behavior is much better then I expected, and the result from the DeepSeek Flash v4/v4.1 models is on pair or slightly better than identical work done with the Anthropic Sonnet 5/5.5 model. Another issue with Anthropic, was the fact that they quite often flagged our work as a breach of some policy they had set up. While DeepSeek just gets it done without any policies refusal work... Another very big issue with the Anthropic models are the fact that they are very, I should say extremely verbose. Their models output huge amounts of text, relative to the actual task. We’ve run tests where the system prompt, tools, and complete agent software were identical between DeepSeek and Anthropic. In every single case, DeepSeek gave a much better result. I know that other people does not directly share our view regarding DeepSeek, or how worthless the Anthropic models actually are, compare to their cost. You are paying a premium that does not exist. It almost feels like we were tricked into using Anthropic because it’s a Western company, while DeepSeek is from China. TL;DR: DeepSeek has much better agentic behavior than the Anthropic models, and its cost is basically free compared to Anthropic.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论