Flippin’ heck
Agentic workflows are increasingly used in investment applications. That is a good thing because they can significantly enhance productivity for analysts and portfolio managers. But it also increases the possibility that AI hallucinations go unnoticed because no human is checking the output of every single step in the agentic workflow.
A team of researchers from the US, the UK, China, JP Morgan, BlackRock, and State Street tried to find out how big the risk is that agentic AI workflows mess up the investment process. They did this by splitting a typical investment process into two separate steps.
First, they asked ChatGPT, Gemini, Qwen and DeepSeek to summarise the Management Discussion and Analysis (MD&A) section from the quarterly reports of the 100 largest US stocks in 2025 Q1 to Q3. They asked the models to come up with the 20 most important points and preserve the decision direction and confidence as close as possible (the note contains the exact prompts, in case you are interested). Then they gave this summary to Gemini and asked it to make a buy/hold/sell recommendation on the stock. They also gave the full MD&A discussion to Gemini and asked it to do the same thing. As a control experiment, they did the same with the transcripts of earnings calls from these periods.
You’d expect Gemini to come up with the same buy/hold/sell recommendation for each stock, no matter whether it used the whole discussion as input or the summary from AI. If the recommendation differs, it gives you a measure of the reliability (or lack thereof) of an agentic workflow that reads company statements and formulates a recommendation autonomously.
Now, because Gemini, like all generative AI, sometimes hallucinates, you can expect that there will be differences in the recommendations between the two approaches. For Gemini 3.1 Flash, the model used to make the decision was about 11% for the MD&A sections and 9% for the earnings calls, meaning that if you re-run the model on the full input, it will come up with a different buy/hold/sell answer in about 11% or 9% of the cases.
If you think that this is a high number and surely Google’s most advanced model isn’t going to hallucinate that often, I have news for you. The New York Times recently tested Google’s AI Overviews on its search page. They found that with Gemini 2, the results were accurate 85% of the time, and with Gemini 3, 91% of the time. Let me rephrase this. Even with Gemini 3, 9% of the answers you get in Google’s AI Overview are plain wrong and hallucinated. And even if the AI Overview provided the correct answer, it either didn’t provide a source or provided the wrong source that had nothing to do with the answer in 56% of the cases.
So, getting the recommendation wrong in one of ten stocks is par for the course, even in the world’s most advanced models.
But here is what really happened. When given the 20-point summary, the eventual recommendation to buy/hold/sell the stock differed far more often from the recommendation when using the whole doc. The rate at which the recommendation between the two approaches flipped (from buy to sell or sell to buy) ranged from one in four to one in three.
I don’t know about you, but I would not trust an AI where the recommendation to buy or sell a stock changes on the input for a quarter to a third of the stocks in my portfolio. How can you rely on a tool that is that random?
The decision changes after compressing the content