Do LLM “Crowds” Produce Investment Signals? An Empirical Test


Steven Edwards

This paper tests whether aggregating stock selections from a large, philosophically diverse ensemble of LLM personas can produce genuine investment signals beyond passive benchmark exposure. The author built 100 distinct investor pe

rsonas spanning value, momentum, growth, ESG, quantitative, and contrarian philosophies, each independently selecting fifty US-listed equities. The picks were backtested across three weighting schemes over a 561-trading-day period beginning January 2024, strictly after the model’s training data cutoff to eliminate look-ahead bias.

The market-cap-weighted consensus portfolio delivered a striking 29.5% annualized return and a statistically significant 9.65% annualized alpha that survives a six-factor decomposition. But the diversification looks illusory under closer inspection: an equal-weighted version of the same ensemble converges to just 88% of the full universe’s composition with as few as 25 personas, and constituent analysis shows 92% of the fifty stocks are S&P 500 members that overlap heavily with a single neutral LLM call. Bootstrap sampling confirms the “crowd” reproduces the same consensus even with only 25 of the 100 personas.

The paper’s conclusion is a caution for quant allocators: the apparent alpha likely reflects concentrated exposure to large-cap, mega-cap technology names during an AI-driven re-rating window rather than genuine independent stock-selection skill from the LLM crowd. Free-form prompting fails to generate novel investment signals and instead maps onto existing financial media prominence.

Link to full article: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6543259


Other news

  • Boundaries of Time Series Momentum

    Matti Suominen and Erik Hjalmarsson Time-series momentum is one of the most reliable and heavily backtested anomalies in quantitative finance, serving as a foundational alpha source for managed futures and trend-following strategies. This paper uncovers a structural vulnerability that every practitioner must account for: momentum returns persist reliably during normal business cycles, but break down…

  • Do LLM “Crowds” Produce Investment Signals? An Empirical Test

    Steven Edwards This paper tests whether aggregating stock selections from a large, philosophically diverse ensemble of LLM personas can produce genuine investment signals beyond passive benchmark exposure. The author built 100 distinct investor pe rsonas spanning value, momentum, growth, ESG, quantitative, and contrarian philosophies, each independently selecting fifty US-listed equities. The picks were backtested across…

  • Testing an AI-Assisted Research Workflow for Multi-Asset Pullback Strategy Discovery

    This study builds a short-term mean-reversion strategy across six liquid ETFs spanning equities, fixed income, currencies, gold, and commodities (2006–2025), using a 200-day trend filter and a multi-day pullback trigger. Beyond the strategy itself — which delivers strong risk-adjusted returns while invested only ~21% of the time — the paper tests ChatGPT and Claude as…