I gave GPT-5.6 Sol, Claude Opus 4.8, and Grok 4.5 the same 100 frontend briefs—here are all 300 results

After generating enough websites with coding models, I started noticing that each model seemed to reach for the same handful of visual ideas. A single impressive screenshot can’t tell you whether that’s actually true, so I tried testing it at a larger scale. I gave GPT-5.6 Sol, Claude Opus 4.8, and Grok 4.5 the same 100 frontend design briefs. They covered unrelated categories including architecture, deep tech, skincare, streetwear, and coffee, producing 300 websites in total. I put every result into a benchmark called Sitegeist. You can compare the three models on the same brief, or look across one model’s work and see its recurring visual fingerprints: typography, hero layouts, color choices, geometric elements, information density, and overall composition. This isn’t an attempt to pretend that design quality or originality can be reduced to one perfectly objective score. The useful part is the scale and consistency of the experiment: the same 100 tasks, given to three models, with all 300 outputs available instead of a few cherry-picked examples. You can explore them here: sitegeist.kian.im Disclosure: I built the benchmark. preview.redd.it/ugrvdu6kindh1.png