Frontier Models Are No Stock Pickers
Published: September 14, 2026
Since the end of April, Flexreport Finance has run four live stock-picking strategies. Each has two books built from the same shortlist on the same morning. A mechanical book holds the top-ranked names from a scorer of well-documented stock factors. Its LLM twin is chosen by a frontier model, OpenAI's GPT-5.4 at its highest reasoning setting, which reads the same candidates along with every score and all of our research on them.
Through September 11, the mechanical book beat its LLM twin in every strategy, by 4.5 to 13.3 percentage points. Three of the four LLM books trailed the S&P 500, while three of the four mechanical books beat it. Momentum was the weakest strategy for both books, because the price signals it relies on had no edge this summer, and the model did not rescue it.
A clear story has emerged that cutting-edge LLMs are just not very good at stock picking, despite excelling in other areas such as coding and research. In our study, the shortfall comes from which stocks the model chose. It largely ignored the screen's ranking, re-selected names it already owned, and drifted toward larger, more heavily covered companies with better stories. This article will examine these shortcomings in greater detail and argue that the issues run deeper than methodology.
Key findings
- 4 of 4 strategies. The mechanical book beat its LLM twin in every strategy since inception, by +9.9 pts (SMID fundamentals), +4.5 (value & quality), +13.3 (multi-signal) and +4.7 (momentum).
- 1 of 4 LLM books beat the S&P. Only the LLM value & quality book beat the S&P 500. Three of the four mechanical books did; momentum lagged in both books.
- 2% of 1,515 rationales. The model saw each candidate's rank and factor scores but mentioned the screen's rank or score in 2% of its written justifications. In SMID it picked bottom-quartile names as often as top-quartile ones.
- 75–96% re-pick rate. When the screen dropped a stock the model already held, the model kept it 75–96% of the time. A comparable name it did not already own was picked 12–45% of the time.
- p = 0.16 pooled weekly gap. The direction is consistent, but twenty weeks is short: across strategies the LLM books lagged by 44 bp a week. That is not yet statistically conclusive; a bootstrap puts a 94% probability on the cumulative gap being negative.
The methodology
Every strategy follows the same steps:
- Universe: U.S.-listed common stocks that are actively trading and domestic, excluding ETFs and funds, with a 30-day average daily dollar volume of at least $1 million.
- Ranking: a strategy-specific scorer ranks the universe.
- Shortlist: the top 2N names form the shortlist.
- Holdings: each book holds N of them until the next weekly rebalance, generally Monday around midday Eastern time.
The mechanical book holds ranks 1 through N, equally weighted, with no discretion. The LLM book is chosen by GPT-5.4 (snapshot 2026-03-05, reasoning effort xhigh) in one call per strategy per week, each running 70,000 to 125,000 input tokens. The model receives the market backdrop (recent news and earnings-season themes), its current holdings with its own prior rationales, and for every candidate its rank, composite score and each factor sub-score. It also receives Flexreport's company snapshot (thesis, bull and bear case, rating) and, depending on the strategy, recent AI analyses of financial statements, earnings calls and price action. It must choose exactly N names from the shortlist and justify each one.
Exhibit 1: The four strategies as they ran from April to September
| Strategy | Live since | Universe beyond the common filters | Mechanical score | Shortlist → held | LLM book weighting |
|---|---|---|---|---|---|
| SMID fundamentals | 28 Apr | Market cap $300M–$10B | Equal-weight average of percentile ranks: revenue and EPS growth (year over year and 3-year CAGR), ROE, ROIC, operating and net margin, margin trend vs 3-year average, profitability vs industry, EPS beats in the last four quarters, valuation nearest the median P/E and P/S, balance-sheet pass/fail | 40 → 20 | Equal; risk-parity from 16 Jun |
| Value & quality | 13 May | Market cap ≥ $500M; 23 cyclical and commodity industries excluded; 11 hard ratio filters (e.g. P/E 1–35, P/FCF 1–35, operating margin ≥ 5%, debt/equity ≤ 2.5, interest coverage ≥ 2.5×) | 40% valuation, 35% quality, 25% balance-sheet strength, each a direction-aware percentile rank within the survivors | 39–50 → 20 | Conviction (HIGH 3×, MED 2×, LOW 1×); risk-parity × conviction from 16 Jun |
| Multi-signal | 28 Apr | Top 100 by a return-forecast model | 30% forecast rank, 25% fundamentals, 20% 3- and 6-month momentum, 15% company-snapshot rating, 10% narrative signals from AI research | 20 → 10 | Conviction; risk-parity × conviction from 16 Jun |
| Momentum | 28 Apr | No size limit | 1/3/6/12-month returns, price and 50-day above 200-day average, MACD, ADX trend strength, RSI (best near 60), on-balance-volume slope, volume vs 20-day average, 21-day return forecast; scaled down for volatility | 50 → 25 | Equal; risk-parity from 16 Jun |
Source: Flexreport strategy registry and scorer code as deployed during the study window. The scorers were rebuilt on 14 Sep 2026, after the data cut-off, and are not what is measured here.
Two design details shaped the LLM's behavior and are part of what is being tested. The prompts carried a soft incumbency rule: keep current holdings unless the thesis has materially weakened or a clearly stronger candidate has appeared, mirroring a portfolio manager who avoids needless turnover. The momentum prompt told the model not to default to keeping names. From June 17, each call also included performance feedback: the book's return versus the S&P 500, its return versus the mechanical book, and per-holding attribution. From June 16 the LLM books were sized by a risk-parity allocation tilted by the model's conviction.
The books ran live, which matters more for language models than for any other strategy. A model asked in 2026 how it would have traded in 2020 already knows what happened in 2020, so a historical backtest of an LLM is contaminated by lookahead bias. Every pick here was timestamped before the returns it is judged on. Returns are close-to-close from each rebalance, measured daily, before trading costs and excluding dividends, against the S&P 500 price index over identical dates. Because both books share the shortlist, the timestamp and every input, the comparison isolates one question: given the same information, does the model's judgment improve on the ranking it was handed?
The screeners consistently beat the LLM
All four rule based strategies beat its LLM peer, and the LLM's underperformance is clearly traced to repeated behavior rather than a handful of unlucky trades
Exhibit 2: Cumulative return since inception: mechanical vs LLM book vs S&P 500
Cumulative returns through 11 Sep 2026. SMID fundamentals: mechanical +15.5%, LLM +5.6%, S&P 500 +6.7%. Value & quality (from 13 May): mechanical +21.0%, LLM +16.5%, S&P 500 +3.5%. Multi-signal: mechanical +11.0%, LLM −2.3%, S&P 500 +6.7%. Momentum: mechanical +2.5%, LLM −2.2%, S&P 500 +6.7%.
Source: Flexreport strategy performance tables (LLM and mechanical books); S&P 500 price index. Daily compounding, before costs, through 11 Sep 2026. Shared vertical scale; hover a panel for daily values.
Exhibit 3: Performance summary
| Strategy | Mechanical | LLM | S&P 500 | Mech − LLM | Sharpe mech / LLM | Max drawdown mech / LLM | Avg weekly gap (p-value) | Weeks LLM beat mech |
|---|---|---|---|---|---|---|---|---|
| SMID fundamentals | +15.5% | +5.6% | +6.7% | +9.9 pts | 2.39 / 1.05 | −5.9% / −6.5% | −46 bp (0.29) | 10 of 20 |
| Value & quality | +21.0% | +16.5% | +3.5% | +4.5 pts | 3.30 / 2.99 | −4.2% / −3.7% | −22 bp (0.54) | 8 of 18 |
| Multi-signal | +11.0% | −2.3% | +6.7% | +13.3 pts | 0.94 / 0.06 | −12.7% / −13.3% | −69 bp (0.39) | 10 of 20 |
| Momentum | +2.5% | −2.2% | +6.7% | +4.7 pts | 0.58 / −0.39 | −4.6% / −9.5% | −23 bp (0.33) | 11 of 20 |
Windows: 28 Apr – 11 Sep 2026 (value & quality from 13 May). Sharpe ratios annualized from roughly four months of daily returns (risk-free rate 0). Weekly gap = mean of weekly LLM-minus-mechanical returns; p-value from a one-sample t-test. Pooled across strategies: −44 bp a week, t = −1.46, p = 0.16 over 20 weeks.
The gap mostly boils down to stock selection
The LLM book differs from the mechanical book in two ways: which stocks it buys, and how much it puts in each. To measure them separately, we held the LLM's stocks in equal amounts. Comparing that with the mechanical book shows the effect of its stock picks. Comparing it with the LLM's actual book shows the effect of its position sizes.
Exhibit 4: Where the LLM book lost ground: stock selection vs position sizing (percentage points, since inception)
Selection effect / weighting effect in percentage points: SMID fundamentals −10.3 / +2.3; value & quality +0.8 / −4.7; multi-signal −22.1 / +1.1; momentum −2.4 / −3.5.
Source: Flexreport weekly holdings snapshots and daily prices. Weekly close-to-close reconstruction (21 weeks; 18 for value & quality), so totals differ slightly from Exhibit 3's daily series.
Selection explains almost all of the damage in SMID (−10.3 points) and multi-signal (−22.1). Weighting was mixed: it helped slightly in those two strategies and cost 3.5 points in momentum and 4.7 in value & quality. Value & quality is the one case where the model's selection did no harm (+0.8 points). There, eleven hard ratio filters had already narrowed the field to comparable companies, leaving little room for the drift described below.
Exhibit 5: What the model changed each week, and how the changes performed
| Strategy | Mechanical book kept by LLM | Names swapped per week | Swapped in: avg weekly return | Swapped out: avg weekly return | Weeks the swaps helped | Top N vs next N in shortlist |
|---|---|---|---|---|---|---|
| SMID fundamentals | 56% | 8.9 | +36 bp | +118 bp | 37% | +90 bp/wk |
| Value & quality | 51% | 9.9 | +70 bp | +49 bp | 38% | +44 bp/wk |
| Multi-signal | 60% | 4.0 | −67 bp | +213 bp | 32% | +129 bp/wk |
| Momentum | 52% | 11.9 | +8 bp | +23 bp | 53% | +16 bp/wk |
“Swapped in” = LLM picks ranked outside the mechanical top N; “swapped out” = mechanical top-N names the LLM left out. “Weeks the swaps helped” counts weeks in which swapped-in names beat swapped-out names. The last column is the average weekly return of shortlist ranks 1…N minus ranks N+1…2N.
The model is not making marginal adjustments. It replaces 40–50% of the mechanical book every week, which amounts to a different portfolio. And the ranking it overrides carries real, if modest, information: inside each shortlist, the top half out-returned the bottom half by 16 to 129 basis points a week. Overriding half of that ranking every week, with swaps that lose to the names they replace in most weeks, is the whole story in arithmetic form.
Why the model underperforms
1. It evaluates companies one at a time, ignoring the ranking
A factor screen compares every candidate against every other on the same yardsticks and turns that comparison into an ordering. The model received that ordering, with every sub-score beside every name, and largely set it aside. In SMID and momentum, its probability of picking a stock was essentially flat across the shortlist: a bottom-quartile name was as likely to make the book as a top-quartile one. In value & quality the model mainly steered clear of the bottom quartile; only in multi-signal did rank clearly shape its choices.
Exhibit 6: Probability a shortlist name is selected, by its screen rank quartile
LLM pick rate by screen rank quartile (top, 2nd, 3rd, bottom): SMID fundamentals 50%, 64%, 43%, 51%; value & quality 49%, 48%, 53%, 34%; multi-signal 70%, 52%, 45%, 34%; momentum 50%, 59%, 52%, 48%. The mechanical book picks the top two quartiles by construction.
Source: Flexreport weekly shortlists, 28 Apr – 7 Sep 2026. Quartiles of each week's shortlist by mechanical rank.
The model's own words point the same way. Across 1,515 written rationales, 70% lean on narrative vocabulary: thesis, catalyst, demand, guidance, moat, execution, backlog. Only 32% contain a single number, and 2% refer to the screen's rank or score. The model reasons the way a well-read analyst writes, name by name, in prose. It does not work from a cross-sectional comparison it can calibrate. A qualitative “this looks strong” cannot tell the 8th-best name from the 28th.
2. It anchors on what it already owns
The strongest predictor of an LLM pick is not rank, size or news flow. It is whether the model held the stock the week before. When the screen demoted an existing holding out of its top N, the model kept it 91% of the time in SMID, 84% in multi-signal and 96% in value & quality. Most of what the model “adds” to the mechanical book is last week's portfolio: 62–82% of its swaps in those three strategies were existing holdings.
The prompt asked for soft incumbency, so some stickiness is by design. But the model applied it far more literally than the rule implies. “Default to keeping” became “almost always keep,” even as the numbers underneath those names deteriorated. Momentum, whose prompt discouraged stickiness, shows the pattern at a much lower level (75% vs 45% for new names), which suggests the instruction amplified an anchoring tendency rather than creating it.
AGX — SMID fundamentals, 17 Aug, screen rank 36 of 40, next week −20.4%
“Argan remains a direct beneficiary of power and data center construction with record revenue, strong margins, no debt, and a large backlog. We are keeping it because the end market is still favorable and the balance sheet gives management room to keep returning capital.”
IESC — SMID fundamentals, 17 Aug, screen rank 29 of 40, next week −18.4%
“IESC remains a clean power and data center execution story with strong multi year revenue and EPS growth, repeat beats, and solid returns on capital. The AI and electrification buildout still supports demand, and we have not seen evidence that the operating story is fading.”
CBZ — Multi-signal, 16 Jun, screen rank 18 of 20, next week −15.9%
“CBIZ stays because guidance was reaffirmed, EPS guidance moved higher, leverage is falling, and advisory mix plus AI and offshore productivity support steadier earnings quality.”
3. It gravitates to the companies with the best-told stories
The names the model brings in are systematically bigger and better covered than the names they displace. Median daily dollar volume of swapped-in stocks runs 1.7× (value & quality) to 25× (multi-signal) that of the swapped-out stocks, and they carry more news in the preceding month. Bigger, more-covered companies come with richer inputs: fuller company snapshots, detailed earnings-call analyses, more headlines. Those inputs make a more persuasive case for a model that reasons in prose. A $300 million company with excellent numbers and thin coverage gives the model less to write about, and loses its slot to a familiar name with a good story.
Exhibit 7: Anchoring and attention: how swapped-in names differ from the names they replace
| Strategy | Off-screen name already held: kept | Off-screen name not held: picked | Swaps that were existing holdings | Median $ volume in / out | Median news (30d) in / out |
|---|---|---|---|---|---|
| SMID fundamentals | 91% | 26% | 62% | $58M / $26M | 2 / 0 |
| Value & quality | 96% | 12% | 82% | $73M / $44M | 2 / 1 |
| Multi-signal | 84% | 21% | 64% | $78M / $3M | 3 / 1.5 |
| Momentum | 75% | 45% | 19% | $80M / $19M | 1 / 0 |
“Off-screen” = on the shortlist but ranked outside the mechanical top N. Dollar volume = 30-day average close × volume before each rebalance; news = significant company news items in the 30 days before each rebalance.
Exhibit 8: What predicts an LLM pick: odds ratio per standard deviation, with 90% interval
Odds ratios per standard deviation (SMID / value & quality / multi-signal / momentum): already held last week 5.81 / 18.23 / 4.14 / 1.87; screen rank 1.19 / 1.28 / 1.46 / 1.10; dollar volume 1.13 / 1.03 / 2.05 / 1.45; news items 1.37 / 1.42 / 0.91 / 0.91.
Logistic regression of LLM selection on standardized predictors within each strategy's shortlists (418 to 1,018 candidate-weeks per strategy); 90% intervals from a 200-draw bootstrap over rebalance weeks. Log scale. An odds ratio of 2 means one standard deviation more of that predictor doubles the odds of selection.
Read together, Exhibits 6–8 describe a consistent decision process. Prior ownership dominates in every strategy. Rank matters a little, significantly so only in multi-signal and value & quality. Size predicts selection in momentum and multi-signal, and news coverage in SMID and value & quality. That is the profile of a thoughtful reader of company research, not of a disciplined allocator working from a scorecard.
4. More reasoning is not better forecasting
These books ran at the model's highest reasoning setting, producing 18,000 to 22,000 output tokens per decision. The rationales are articulate, accurate about the businesses, and frequently right about fundamentals, yet the book still lost to a sort. Weekly stock returns are dominated by noise, and the edge in these factors is small: a rank correlation with next-week returns of about 0.05. That kind of edge pays only when applied consistently across many names, week after week. Discretion that is not itself informative dilutes it, and discretion with a systematic tilt toward large, familiar, already-owned names erodes it.
Momentum failed for both books; the model made it slightly worse
Momentum is the weakest result in the study, and the reason is mostly not the model. The mechanical momentum book itself trailed the S&P 500 by 4.2 points: the signals it is built on simply did not work this summer. Inside the momentum shortlist, the classic ingredients were negatively related to the following week's return. Higher RSI, stronger one-month performance, a bullish MACD and a surge in volume all tended to precede weaker weeks. In a separate test across the whole liquid U.S. market, buying the week's 25 best performers by one-day or five-day return lagged an equal-weight universe by roughly 1% a week from late April onward (not statistically significant on its own).
Exhibit 9: Momentum factors inside the shortlist: rank correlation with next-week return
Mean weekly rank correlation with next-week return (t-stat): RSI level −0.094 (−2.55); 1-month return −0.090 (−1.92); volatility −0.082 (−1.78); volume vs 20-day average −0.081 (−2.32); strong trend flag −0.077 (−1.76); 3-month return −0.071 (−1.56); MACD bullish −0.069 (−2.28); 6-month return −0.043; 12-month return −0.029; ADX trend strength −0.021; price and 50-day above 200-day +0.007; 21-day return forecast +0.079 (1.66); RSI closeness to 60 +0.084 (2.04).
Mean weekly Spearman correlation between each factor and the following week's return across the 50-name momentum shortlist, 21 rebalances. Dark bars: |t| ≥ 2. The return forecast had not been updated since 27 March, so it is a static ranking.
When the underlying signal carries no information, there is no edge for the model to add. The measured LLM gap in momentum is therefore modest (−4.7 points), and most of it came from weighting (−3.5) rather than selection (−2.4). The model's swaps were close to a coin flip: they helped in 53% of weeks. The same tilts still showed up. Swapped-in momentum names had 4× the dollar volume of the names they replaced, and were slightly less stretched (RSI 66 vs 70) but more volatile (average true range 2.1% of price vs 1.7%). The momentum rationales were technical in vocabulary (99% cite trend, RSI, MACD or volume), yet 60% also leaned on fundamentals or demand stories outside the strategy's mandate.
A change in methodology might not fix things
Of course, these results might be due to a flawed methodology, but they currently align with other findings in the field, most notably the public trading contests Bloomberg surveyed in May, which highlighted a variety of LLM trading strategies that performed poorly.
The verdict was unflattering: most of the systems lose money, they trade too much, and they make wildly different decisions when given identical instructions.
The best known, Nof1's Alpha Arena, gave eight frontier models, among them Anthropic's Claude, Google's Gemini, OpenAI's ChatGPT and xAI's Grok, $10,000 apiece in four two-week contests trading U.S. tech stocks. Together they lost about a third of their capital, and only six of the 32 results finished in profit. The Flat Circle blog, tracking 11 such arenas, found the median model profitable in just two of them.
“LLMs can’t really make money by themselves. You need basically a very sophisticated harness and scaffolding and data platform in order to even give them a chance.”
— Jay Azhang, founder of Nof1, speaking to Bloomberg, May 6, 2026
Our study is, in effect, the experiment Azhang describes. The model was not trading alone. It worked inside a curated data platform, received a pre-ranked shortlist with every factor score beside every name, and had company research, earnings-call analysis and market context for each candidate. His diagnosis of what goes wrong also matches what we measured. Models, he said, are good at research and at reaching for the right tools, but do not yet know how much each of the many variables that move a stock actually matters. In our books that weighing is precisely the ranking the model set aside (Exhibit 6). The platform gave the model a chance. It did not make the model a better allocator than the ranking it was handed.
Strong researchers, weak stock pickers
The model's research was often excellent. Its rationales were fluent, specific and mostly correct about the companies involved, and the analyses it drew on (company snapshots, earnings-call and financial-statement reviews) are exactly the work frontier models do well. What failed was the step from understanding companies to ranking them against each other and committing capital. There the model behaved like a reader rather than an allocator. It weighed stories over scorecards, favored the familiar and the already owned, and replaced half of a modestly informative ranking with judgment that was not.
This does not rule out capable AI stock pickers. However, it does show that an out-of-the-box frontier model, given a strong shortlist and every piece of research on it, can make a portfolio considerably worse.
Flexreport still believes that a strong data platform is a necessary ingredient for LLMs to successfully pick stocks. However, our results suggest it does not guarantee success, and that the model should not be the final allocator. As a result, we are changing our own process, with the mechanical scorers becoming the default approach for every strategy. The LLM books will still be available alongside them.
We will continue to experiment with the LLM approach, as it offers exciting possibilities, and our view would of course change if LLMs start to consistently outperform rule-based strategies. We will keep publishing the results of both approaches on our Stock Picks page.
Methods: live books from Flexreport's holdings snapshots and daily performance tables; LLM calls verified in the platform's model cost log (gpt-5.4-2026-03-05, reasoning effort xhigh). Weekly t-tests on LLM-minus-mechanical returns; logistic regressions on standardized predictors with week-block bootstrap intervals; rationale language classified by keyword dictionaries over all 1,515 LLM rationales. Not investment advice. Past performance does not guarantee future results.
Read more from the Flexreport Finance blog · Get started with Flexreport Finance