Most AI Stock-Picking Funds Lagged the S&P 500—AIEQ Survived

Most publicly available funds in a historical study of AI-assisted stock selection failed to beat the S&P 500 over their respective lifetimes through 2023. That finding remains a serious warning, but it is not a live ranking of every fund using artificial intelligence in 2026.
The clearest update is that AIEQ, the earliest fund in the study’s fully AI group, did not disappear. As of August 2026, the official AIEQ fund page listed about $123.7 million in net assets, 160 holdings and a 0.75% expense ratio; its NAV return was 130.23% cumulatively and 9.96% annualized since inception through July 31, 2026. Those are positive absolute returns, but the page does not provide a matching since-inception S&P 500 comparison, so the figures cannot establish that AIEQ has repaired its historical benchmark deficit.
What the original comparison actually found
Gary N. Smith and Sam Wyatt examined funds launched from October 2017 onward that allowed an AI system either to make investment decisions or to participate in stock selection. Their sample contained 11 fully AI funds and 43 partly AI funds, with performance measured from each product’s launch until December 31, 2023, or its earlier closure.
In the authors’ published account of the fund results, all 11 fully AI products trailed the S&P 500 and six lost money. Their average annual return was −1.8%, versus 7.6% for the index. Among the 43 partly AI funds, only 10 beat the benchmark; their average annual return was 7.11%, compared with 12.43% for the S&P 500. Six fully AI funds and 25 partly AI funds had closed by the time of the analysis.
AIEQ illustrates the historical gap in a single product. Through the end of 2023, the study calculated a 63% cumulative total return for the fund, against 108% for the S&P 500 over the corresponding period. MIND, another fully AI product, closed in 2022 after recording a −12% cumulative return while the comparison index gained 65% over its relevant window.
The research subsequently moved beyond its initial opinion-article presentation. EBSCO’s journal record lists Smith and Wyatt’s “The Disappointing Performance of AI-Powered Funds” in the 2025 volume of The Journal of Investing, with DOI 10.3905/joi.2025.1.346.
Why AIEQ’s survival does not overturn the study
A fund can make money and still disappoint investors relative to a cheaper or simpler alternative. Absolute return asks whether an investment gained value; benchmark-relative return asks whether it justified its strategy, costs and risks against an appropriate comparison. The two questions should not be collapsed into one headline number.
AIEQ’s current record shows why dates matter. Its newer gains extend the endpoint beyond the study’s December 2023 cutoff, while the historical comparison remains valid for the period the researchers measured. Declaring either a permanent failure or a comeback would require recalculating both AIEQ and the same total-return benchmark through an identical date, including fund expenses and distributions.
Survival is also a commercial result, not proof of investment superiority. A product may remain viable because it retains sufficient assets, distribution and investor demand even when it has not beaten a broad index. Conversely, closure can reflect low assets or an issuer’s product strategy as well as poor returns; the closure count alone does not isolate AI as the cause.
“AI fund” describes several different products
The study concerned funds that used AI in the investment process. That category should not be confused with an AI-themed ETF that simply owns semiconductor, software or automation companies. A technology-themed portfolio can rise because its holdings prosper even if no algorithm selects them, while an AI-selected portfolio can own companies outside the technology sector.
There is another important split inside AI-powered management. In a fully AI fund, the system is presented as making portfolio decisions without discretionary human intervention. In a partly AI fund, software may rank securities or generate signals while a human adviser retains authority over the final portfolio. A result for the combined human-machine process cannot be attributed to the model alone.
Chatbot stock suggestions are further removed from the evidence. The fund study evaluated investable products with observable return histories; it did not test a general-purpose language model responding to one investor’s prompt. Screenshots of a model naming several successful stocks omit allocation, timing, rejected picks, transaction costs and the losses that may have occurred elsewhere in the portfolio.
What the evidence can and cannot prove
The results show that attaching AI to a fund did not reliably produce market-beating returns in this sample. They do not prove that every machine-learning strategy must fail, that the S&P 500 is the correct benchmark for every mandate, or that algorithms contributed nothing to risk control and operations.
The comparison also joins funds launched on different dates. Lifetime returns are appropriate for judging each product’s public record, but the market conditions experienced by a 2017 launch were not identical to those faced by a later entrant. Differences in geography, market capitalization, portfolio constraints, fees and trading activity can also affect the gap against a large-cap US index.
Most importantly, poor fund-level performance does not by itself identify a single technical cause. A model may overfit historical patterns, react slowly to structural change or optimize the wrong objective, but a fund can also suffer from its mandate, implementation costs or human overrides. The observed result belongs to the complete investment product.
How to evaluate the next AI stock-picking claim
For investors—and for creators presenting AI-generated market content—the useful question is not whether a model produced an impressive pick. It is whether the complete, investable strategy generated better risk-adjusted results after costs than a relevant alternative.
- Match the dates. Compare the strategy and benchmark over exactly the same interval, preferably using total returns that include distributions.
- Include the full record. Closed funds, abandoned experiments and losing selections matter; highlighting survivors creates a misleading sample.
- Separate tests from deployed money. A backtest, simulated portfolio and live fund are different levels of evidence.
- Identify human discretion. Readers should know whether AI chooses securities, supplies research signals or merely supports a manager who makes the decisions.
- Account for costs and risk. Fees, turnover, taxes, volatility and drawdowns can change the practical value of a headline return.
The updated record therefore supports a narrower conclusion than either hype or ridicule suggests. Most funds in the measured sample lagged their benchmark, and many closed; meanwhile, AIEQ survived and accumulated a positive return. That combination is not a contradiction—it is a reminder that profitability, outperformance and product survival are three different tests.
Also read:
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.