Can ChatGPT Forecast Stock Price Movements?
Financial markets process an enormous volume of corporate news every day. Earnings announcements, management changes, clinical trial results, insider transactions, partnerships, and regulatory developments can all affect a company’s value. The challenge is not simply identifying whether a headline sounds positive or negative. Investors must understand its economic implications, anticipate how other market participants will interpret it, and determine whether the information has already been incorporated into the stock price.
Large language models were not originally designed to forecast financial markets. Yet their ability to interpret context and synthesize complex information raises an important question: Can a general-purpose model such as ChatGPT anticipate stock-price reactions to corporate news?
Can ChatGPT Forecast Stock Price Movements?
- Alejandro Lopez-Lira, Yuehua Tang
- Journal of Financial Economics, 2026
- A version of this paper can be found here
- Want to read our summaries of academic finance papers? Check out our Academic Research Insight category
Key Academic Insights
GPT-4 can interpret the economic meaning of financial news
The authors ask GPT-4 to classify each corporate headline as positive, negative, or uncertain for the company’s short-term stock price. The model receives only the company name and headline, without numerical market data, analyst forecasts, or financial fine-tuning. Its task therefore requires more than conventional sentiment analysis: it must interpret the economic context of the announcement and determine how investors are likely to respond.
The test is genuinely out of sample
To limit look-ahead bias and the possibility that the model memorized historical market reactions, the study examines headlines published from October 2021 through May 2024, while the GPT-4 version used by the authors had a September 2021 knowledge cutoff. The final dataset contains 159,137 firm-headline-date observations covering 4,123 U.S. companies.
GPT-4 accurately predicts the immediate market reaction
GPT-4’s classifications align strongly with the direction of the market’s initial response. The daily long-short portfolio based on overnight headlines has a 93.3% hit rate for the immediate reaction, while the corresponding hit rate for intraday news is 88.8%. Stocks associated with positive overnight news rise by an average of 1.27% between the previous close and the market opening, while stocks associated with negative news fall by 1.79%. The initial reactions are even larger for intraday announcements.
Predicting the initial response is not the same as earning a tradable return
The exceptionally high initial-reaction hit rates describe price movements occurring as the news reaches the market. Unless an investor had advance or unusually rapid access to the information, those returns would generally not be available for trading. The economically relevant question is therefore whether prices continue moving after the announcement, once ordinary investors have an opportunity to act.
Stock prices continue drifting in the predicted direction
Following the initial reaction, prices tend to move in the direction identified by GPT-4 for another one to two trading days. A daily rebalanced strategy that buys stocks with positive overnight headlines and shorts those with negative headlines produces an average pre-cost return of 34 basis points per day, with a 58% hit rate and an annualized Sharpe ratio of 2.97. For intraday news, the corresponding figures are 50 basis points, 55%, and 2.63.
Negative news contains more predictability
The post-announcement drift is substantially stronger following negative news. For overnight headlines, the long portfolio earns an average of 8 basis points per day with a Sharpe ratio of 0.78, while the short portfolio earns 26 basis points with a Sharpe ratio of 2.01. This asymmetry is consistent with limits to arbitrage because negative views are more difficult and costly to express through short selling.
Smaller stocks exhibit greater underreaction
GPT-4’s score predicts subsequent returns more strongly among smaller companies. Small stocks tend to receive less analyst attention, have lower liquidity, and impose higher trading costs, making it more difficult for sophisticated investors to correct mispricing immediately. The result supports the interpretation that the model is identifying delayed information incorporation rather than a conventional risk premium.
Transaction costs substantially reduce the apparent profits
The headline strategy requires extremely high turnover. Before transaction costs, the overnight long-short portfolio produces a cumulative return of approximately 700% over the sample. With assumed round-trip trading costs of 5 basis points, the cumulative return remains above 300%; at 10 basis points, it remains above 100%; and at 20 basis points, the strategy becomes unprofitable. These estimates do not necessarily capture the full price impact, borrowing costs, operational constraints, or scalability challenges an investor would face.
The predictive opportunity appears to be declining
The annualized Sharpe ratio of the overnight strategy falls from 6.54 in the fourth quarter of 2021 to 3.68 in 2022, 2.33 in 2023, and 1.22 from January through May 2024. The authors interpret this decline as suggestive—not conclusive—evidence that broader LLM adoption is accelerating the incorporation of news into market prices. Other changes in market conditions could also contribute to the pattern.
Practical Applications for Investment Advisors
Use LLMs to improve information processing, not as automatic trading systems
The strongest evidence in the paper concerns GPT-4’s ability to interpret the economic implications of corporate news. Advisors may therefore find LLMs most useful as research assistants that organize information, identify potentially material developments, and explain why a headline could matter. Turning those assessments into an automated strategy introduces an entirely different set of execution, governance, and risk-management challenges.
Distinguish forecasting accuracy from investable performance
A model may correctly predict a market reaction without giving ordinary investors an opportunity to earn the associated return. Much of GPT-4’s highest forecasting accuracy relates to the initial response that occurs before a feasible trade can be placed. Any investment evaluation should focus on post-signal returns after realistic delays, transaction costs, borrowing expenses, turnover, and price impact.
Focus research attention where information is hardest to process
The results suggest that LLMs may add the most value when news requires contextual interpretation rather than the mechanical extraction of a number. Insider transactions, specialized conference announcements, scientific information, and strategically framed corporate disclosures may warrant greater analytical attention than transparent and standardized announcements that markets process quickly.
Expect AI advantages to decay
A publicly known signal rarely remains profitable indefinitely. As investment firms adopt similar tools, competition accelerates price discovery and compresses returns. Advisors should be skeptical of backtests that assume a historical LLM signal can be implemented unchanged in the future, particularly when the model, the investor population, and market microstructure are evolving rapidly.
How to Explain This to Clients
“ChatGPT does not possess a crystal ball that tells us where stocks will trade next month or next year. What this study shows is narrower and more interesting. When GPT-4 was given a company name and a news headline, it was often able to understand whether the announcement was economically positive or negative. Its interpretation closely matched the market’s immediate reaction, even though the model had not been specifically trained to forecast stock prices. Prices also continued moving in the same direction for a short period, particularly after negative news and among smaller companies. This suggests that markets sometimes need time to fully understand complicated information. However, the trading results were highly sensitive to transaction costs, required substantial turnover, and weakened as AI adoption increased. The practical lesson is not that investors should blindly trade every ChatGPT opinion. It is that advanced language models may help investors process complex information more efficiently—and that widespread use of these tools may make markets more efficient over time.”
The Most Important Chart from the Paper
This figure shows the performance of trading strategies based on GPT-4’s assessment scores of overnight news, without considering transaction costs. If a piece of news is released before 9 a.m. on a trading day, we enter the position at the market opening and exit at the close of the same day. If the news is announced after the market closes, we assume we enter the position at the next opening price and exit at the close of the next trading day. All the strategies are rebalanced daily. The green line corresponds to an equal-weighted portfolio that buys companies with good news, according to ChatGPT 4.
The red line corresponds to an equal-weighted portfolio that short-sells companies with bad news, according to ChatGPT 4. The blue line corresponds to an equal-weighted longshort portfolio that buys companies with good news and short-sells companies with bad news, according to ChatGPT 4. The grey line corresponds to a value-weighted market portfolio

The results are hypothetical results and are NOT an indicator of future results and do NOT represent returns that any investor actually attained. Indexes are unmanaged and do not reflect management or trading fees, and one cannot invest directly in an index.
Abstract
We document the capability of large language models (LLMs) like ChatGPT to predict stock market reactions from news headlines without direct financial training. Using post-knowledge-cutoff headlines, GPT-4 captures initial market responses, achieving approximately 90% portfolio-day hit rates for the non-tradable initial reaction. GPT-4 scores also significantly predict the subsequent drift, especially for small stocks and negative news. Forecasting ability generally increases with model size, suggesting that financial reasoning is an emerging capacity of complex LLMs. Strategy returns decline as LLM adoption rises, consistent with improved price efficiency. To rationalize these findings, we develop a theoretical model that incorporates LLM technology, information-processing capacity constraints, underreaction, and limits to arbitrage.
was originally published at Alpha Architect. Please read the Alpha Architect disclosures at your convenience.