High Token Usage of flash 0731 with Hermes
Anyone else experiencing unusually high token usage with Hermes Agent + DeepSeek Flash-0731? Hi everyone, I’m curious if anyone else has experienced something similar while using Hermes Agent . I’ve been using Hermes actively for about 2 months with the DeepSeek V4-Pro API , and my usage typically cost me around $25/month . After the release of Flash-0731 , I switched to it because I expected similar or better results at roughly one-third of the price . Instead, I’m seeing extremely high token usage . In one day alone, it burned through around $12 , even though I wasn’t doing anything particularly intensive. What changed? Model: DeepSeek V4-Pro → Flash-0731 Reasoning: Initially: Max Then switched to Low Then tried None Still ended up spending around $5 on just a few very simple tasks Hermes: I didn’t change anything else on the Hermes side. I’m basically just switching the reasoning mode and starting new sessions. Below is the usage statistics. Has anyone else experienced similar behavior with Hermes + DeepSeek , particularly with Flash-0731? Does anyone have an idea what could cause this kind of token usage or where I should start looking? My first thought was context length , but I’m not convinced that’s the main issue. Nothing changed on my side that should cause significantly larger contexts, and even starting completely new sessions doesn’t seem to help much. Obviously, larger contexts can lead to higher token usage , but that seems more like a consequence of the problem rather than the cause. Would really appreciate any insights or suggestions on what I should check. Usage over the last 7 days Usage over the last 30 days