DeepSeek V4 Pro scored 89.1 at $0.0017/call in our 17-model, 4-test benchmark
We've been running a private benchmark suite for a few months, testing models on strategic reasoning, advisory quality, long-form analytical production, and adversarial critique. 17 models, 4 test batteries, scored against a reference answer by a separate model. This consolidates all of it into one picture. Blend = 30% strategic reasoning (4 scenarios: frame-breaking, multi-dimensional review, channel coordination, portfolio prioritization) + 25% advisory quality (6 prompts: triage, architecture risk, blindspots, client advisory, routing) + 25% long-form analytical production (structured section generation from scratch, scored on rigor, depth, voice) + 20% critical review (structural audit + adversarial CTO critique). Weights redistributed for models not tested on all four. Rank Model Strategic Advisory Writing Review Blend - Fable 5 (ref) 100 - - - 100 1 GPT-5.6 Sol high 100 96.7 92.8 95.6 96.5 2 GPT-5.6 Terra max - - 97.2 94.8 96.1 3 Qwen 3.8 Max 92.3 98.3 - - 95.0 4 GPT-5.6 Sol xhigh 90.0 - 95.2 96.3 93.8 5 Opus 4.8 87.0 96.7 93.6 93.2 92.3 6 Grok 4.5 96.0 98.3 88.0 83.2 92.0 7 GLM-5.2 84.5 98.3 - 93.2 91.4 8 Kimi K3 93.0 95.0 88.3 86.7 91.1 9 Sonnet 4.6 - - 91.0 - 91.0 10 DeepSeek V4 Pro 86.0 86.7 94.0 90.5 89.1 11 GPT-5.5 85.3 - 92.8 89.9 89.0 12 Muse Spark 1.1 - 98.3 82.0 78.3 86.8 13 Qwen 3.7 Max - - 86.4 85.4 85.9 14 Sonnet 5 - - 84.0 87.2 85.4 15 Qwen 3.7 Plus 81.5 - - - 81.5 16 Gemini 3.5 Flash 76.0 - 81.8 83.4 79.9 17 MiniMax M3 58.0 - - - 58.0 Dashes = not tested on that dimension. Fable 5 is the reference answer used for scoring, not a contestant. Caveats: n=1 per test per model, directional only. Qwen 3.8 was blind-validated by GLM-5.2 (different model family); other models were judged by Opus 4.8 - cross-judge comparison is approximate within ±3-5 pts. GPT-5.6 variants are separate rows because they behave as different models in practice. Strategic reasoning and advisory quality are domain-specific to strategic analysis; says nothing about coding, vision, or long-context work. This is private operator benchmarking, not a scientific ranking.