Testing an AI-Assisted Research Workflow for Multi-Asset Pullback Strategy Discovery


Soňa Beluská, Quantpedia

This study develops a systematic short-term reversal strategy, capturing the tendency of an asset to partially retrace after a brief run of down days, across six liquid ETFs spanning equities, fixed income, currencies, gold, and commodities from 2006 to 2025. It pairs a 200-day moving-average trend filter with a multi-day pullback trigger, volatility-adjusted sizing, and equal-weighted allocation. Uniquely, nearly all of the analysis was performed using ChatGPT and Claude, so the paper doubles as a test of how useful those tools are for real quantitative research.

The best specification (200-day MA trend filter, 2-day pullback, 1-day hold) delivered a Sharpe of ~0.95, nearly double the best passive benchmark, despite being invested less than half the time. The mean-reversion edge is concentrated almost entirely in the first day after entry.

A more conservative 3-day pullback version is recommended for institutional use: comparable returns to buy-and-hold but with a drawdown roughly 2.4x smaller (-13.69% vs -33.28%) and a best-in-class Calmar ratio, while deploying only ~21% of capital.

Sub-period tests show no evidence of alpha decay, the most recent window (2021–2025) was the strongest on record (Sharpe 1.25), and win rates stayed above 56% throughout.

On the AI comparison: Claude better understood the research intent and produced more coherent written analysis and article drafts, while ChatGPT over-optimized and produced nicer-looking graphs. Both reached the same strategy results, but the authors stress that precise instructions and a final independent human check remain essential.

Link to full article: https://quantpedia.com/testing-an-ai-assisted-research-workflow-for-multi-asset-pullback-strategy-discovery/


Other news

  • Testing an AI-Assisted Research Workflow for Multi-Asset Pullback Strategy Discovery

    This study builds a short-term mean-reversion strategy across six liquid ETFs spanning equities, fixed income, currencies, gold, and commodities (2006–2025), using a 200-day trend filter and a multi-day pullback trigger. Beyond the strategy itself — which delivers strong risk-adjusted returns while invested only ~21% of the time — the paper tests ChatGPT and Claude as…

  • Guardrails Make the Researcher: What an AI Agent Got Right (And Wrong) Replicating Nine Equity Anomalies

    An autonomous AI research agent was tasked with replicating nine published U.S. equity anomalies on clean, survivorship-free data. On a faithful build, none survive out-of-sample — and the lone apparent survivor turned out to be the agent’s own construction error. The real lesson is that an AI researcher is only as trustworthy as the guardrails…

  • How Wise is the Crowd? Bias and Edge in Prediction Markets

    Using tick-level data from Polymarket and Kalshi, this study finds the classic favorite-longshot bias is largely a ‘Yes Bias,’ that the most capitalized traders systematically underperform smaller ones, and that the most vocal participants add no informational edge. It offers a reproducible method for denoising prediction-market probabilities.