I compiled LLM inference pricing across 7 providers — the caching numbers are surprising(spreadsheet included) [R]
I've been comparing GPU/LLM providers for a side project and ended up with way too many browser tabs and spreadsheets. So I decided to pull the public pricing data into one sheet and compare it side by side. A quick disclaimer: this is not benchmark data . I didn't run latency tests or throughput measurements. Everything comes from public pricing pages and APIs (OpenRouter, DeepSeek, Together AI, Fireworks, Groq, etc.). The spreadsheet currently tracks: Input/output token pricing Context windows Cached inpu
评论
?
参与讨论