I spent $100 benchmarking GPT-4o, Claude Opus 4, and DeepSeek V4 on 100 real-world prompts — here are the results
I run a small SaaS and my AI bill was getting out of hand. So I ran a controlled benchmark: same 100 prompts (coding, writing, analysis, translation) across 3 models. Price context: Model Input / 1M tokens Output / 1M tokens OpenAI GPT-4o $2.50 $10.00 Claude Opus 4 $15.00 $75.00 DeepSeek V4 Pro $0.30 $0.60 Results summary (coding tasks, 50 prompts): GPT-4o: passed 43/50, avg quality score 7.8/10 Opus 4: passed 46/50, avg quality score 8.4/10 DeepSeek V4: passed 44/50, avg quality score 7.9/10 The kicker: De
评论
?
参与讨论