The Frontier AEO Tracker: What Astra Chooses (and every other frontier model, and what you can do about it)
Naive autoresearch investment in our AEO have yielded impressive ROI, and so naturally it was time to take it seriously. We were inspired by What Claude Code Actually Chooses, and decided to extend/adjust it to our tastes.
After a few billion tokens of prototyping, aligning, and scaling pipelines, here’s the Latent Space Frontier AEO tracker. Our methodology extends AmplifyingAI’s to run 6 prompt variations over 7 models (search on) in 161 categories, from coding agents to AI podcasts to AI Sandboxes to Managed Databases to ASR models to even oddball categories like Angel investors and Corporate spend and Payroll software.
Answer extraction was done by Astra, and scored for a proprietary AEO score that gives weight to first choices, alternative choices, mentions, but also negative weights to mild and strong anti-recommendations (which are rare, but do happen). Because we know you’ll want it, we also extracted the top cited sources which influence Agent recommendations, as well as an analysis of top failures.
Basic Results
Here are the most dominant products (in their categories) in the world:
There are some familiar names in there — opening up the natural question of contamination, which we have checked. Since we have nothing to hide, every prompt and answer pair is inspectable.
However, bias does exist - when models are asked for coding agent recommendations, Fable/Opus like Claude Code and Sol/Astra like Codex and Grok loves Cursor and Muse loves Muse Code and SWE-1.7 loves Devin and so on. I wonder why. You can see other “soft biases” emerge too…
That said there are notable examples of GPT models recommending Claude, a laudable nonbias:
There are 28 categories (out of our total 161) which have a universally dominant primary choice - among all surveyed frontier models.
There are a lot more “close contests” and “always the vibesmaid, never the vibe” categories which should be key AEO battlegrounds.
Sol vs Astra, Opus vs Fable
New pretrains for new model classes represents a new opportunity to check in on what the labs are moving towards in their data and RL priorities, and to check in on whether startups’ investments in AEO are paying off. We prepared special reports analyzing our rankings, observing VERY consequential flips in model choices between model generations from the same lab.
We have separate Opus→Fable and Sol→Astra summary pages. For some flips, we highlighted a neutral analysis of what competitors did better in each scenario.
Efficiency vs Confidence, and Recommendation Sourcing
One of our most surprising findings between Sol→Astra and Opus→Fable is that Anthropic seems to be biasing their models to searching more sources (Sol median of 9 sources, vs Astra median of 5, vs Opus median of 11 sources, vs Fable of 15). Astra seems to be just generally a lot more “confident”, or “efficient”, depending how you look at it - Astra is FAR less likely to change its mind when you lightly paraphrase your question. This makes the value of AEO itself rise as choice randomness declines.
Sources analysis also somewhat strongly predicts what the labs do prioritize vs don’t.
However the sample size is small here and only represents what we can scrape from attempted toolcalls, not the pretrain dataset. What we CAN validate is that AEO practices measured by Ora and Vercel, like markdown content-negotiation, are real and failures discourage models from reading your content.
Just for fun
Here are the top Angels in the world according to LLMs (some dedupes left to do…).
See more
We also made a little family feud type game where you can see if your priors align with the data. Fun!
We are open to further suggestions and business enquiries to develop this if it is of interest. Ping @latentspacepod or email business@latent.space (we have a business manager now! woo!)
As we note in our methodology post, we did try VERY hard to include Gemini/Antigravity, GLM/Zcode, and DeepSeek/DeepCode, but errors and rate limits made them untenable to include in this first run analysis. Please let us know how to raise limits if you represent these companies.