Claude Opus 5.5 placed 3rd of 14 models at rebuilding photos in Blender, 1 point off first on the desk. Plus a caching mistake worth knowing about
I'm building a photo-to-Blender tool and benchmarked 14 models as its agent: look at a photo, write and run Blender Python, render, compare, repeat. Hard caps of 20 minutes, $4 and 60 requests per scene. Scoring is deterministic code, not an LLM judge. How Claude did (average of three photos, 0–100): Model Score Cost per attempt Notes Claude Opus 5.5 60 $1.19 3rd overall; desk 61 vs GPT-6 Astra's 62 Claude Sonnet 5.5 56 $0.50 quickest on the board, about 5 min a scene Claude Fable 5.1 48 $3.92 Claude Opus 5 45 $3.92 best on the shopfront (56), worst on the car (34) Top of the board: GPT-6 Astra 66 ($3.91), GPT-6.1 Sol 61 ($0.36). The 5.5 models ran later, in Claude Code. As a check I ran GPT-6 Astra the same way and it scored 63 instead of 66, so read the gap with that in mind. The caching mistake. In my first runs of Fable 5.1 and Opus 5, the gateway didn't add Anthropic's cache-control markers. 0% of the context came from cache, against 96% for the OpenAI models, so every request paid full input price for the whole conversation again. Those runs hit the $4 cap after 8 to 18 requests; Astra averaged about 36. With the markers added they ran at 96–97% cached, and those are the numbers above. The uncached attempts cost about $16.39 and aren't in the ranking. If you call Claude through your own proxy or gateway, check your cache hit rate. One run per model per photo. Write-up with every render and the scoring details: kaloyan.blog/ai-models-rebuild-a-photo-in-blender