37 billion tokens, $404 - Super Impressed! DeepSeek Flash
I wanted to share this because I've been really impressed with DeepSeek Flash, and I don't see many people talking about how well it holds up in agent setups. I run an orchestrator across a bunch of different models for my projects. Flash started out as one small piece of that and has turned into the thing doing most of the heavy lifting. Right now I've got anywhere from 60 to 100 Flash agents running at any given time, working alongside my other models on several large projects. They're making a big dent in the backlog. The screenshot shows my last 30 days: 37 billion tokens and 233k requests for $404. Most of that came in the last few days, once I really started scaling it up. That works out to about a penny per million tokens. It's not perfect. Some tasks are too much for it, and that's what the orchestrator handles. When Flash can't get something done, the orchestrator kicks it up to a stronger model. That happens a lot less than I expected. Flash does the bulk of the work, and the expensive models only step in for the hard stuff. I'd hate to see what this month would have cost if I'd run everything through the usual frontier models. The savings have been huge. If you're building anything agent-heavy and haven't tried putting Flash at the base of your stack, give it a shot. Happy to answer questions about the setup. preview.redd.it/h0kt2yirmish1.png