$1 per 100M deepseek claims, real math or just cache hits and off peak

We keep seeing the $1 per 100m deepseek screenshots too and the math never works at face value tbh. here is what that number actually seems to be made of (lmk if im wrong): cache hits doing the heavy lifting. a repeat context turn bills a fraction of a fresh read so anyone reusing big prefixes gets ratios that look fake next to cold start usage off peak window on direct. deepseek direct cuts input rates off peak, so the same tokens cost different amounts depending when you run, while flat hosts like deepinfra or others bill the same rate around the clock. screenshots never show the timestamp. input heavy mixes. 100m of cached input reads is not 100m of fresh output writes, output bills way higher and agent loops are output heavy. the headline number hides the mix. router vs direct cut. going thru a router adds its own fee and can lose cache hits along the way so the same usage bills more than hitting the model endpoint direct. stale screenshots. model Ids and rates moved since spring and the old v4 flash names resolve to v4.1 flash now (v4 pro was set to re route as well but they walked it back and kept it as is) so an old screenshot and todays bill can be different models so the $1 screenshots are real but narrow, cached input off peak with the right mix, not everyday agent usage. for normal use the boring levers are pinning big prefixes for cache and picking flat per token billing over quota walls, direct from the maker or wherever the rate stays flat. what raitos are u actually seeing day to day, anyone close to those screenshots on real workloads.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论