Project · Failed

AI Trading Arena

An experiment: local and paid language models against a rules engine, a coin-flip control, and buy-and-hold SPY, on US equities, with a real cost model. Built on the wreckage of an earlier intraday bot.

Original bot: outcome and lesson

Independent audit of the original bot (2-year window, Jul 2023 – Jun 2025, 1,520 trades): net P/L −$835.86, gross +$10.62, profit factor 0.771 net / 1.003 gross. Slippage was −$846.48 — 101% of the reported loss. The system was not “losing on a weak edge.” It had no signal, and paid to trade.

What the Arena showed

In the later arena, buy-and-hold beat the mechanical agents over 180-day windows. Local LLMs did not beat them in a trustworthy test: the models were clamped to ~4k tokens and were not reading the full market view, so those rows are omitted.

This is not a signal service. There is no live trading from this page.

Across all 316 windows Archived resultsRun: July 25, 2026 · Controls only · 316 windows · 180-day horizon · $100 starting balance
Strategy Median final balance ($) Profitable windows vs. random · reported
Buy & hold SPY 105.93 74% +27.27 p=0.000
Mean reversion 81.93 35% +5.74 p=0.004
Random 77.37 11% baseline
Legacy momentum 32.01 0% −43.65 p=0.000

Historical experiment, not live performance. Median balance is final wealth from a $100 start; profitable windows finish above $100. Comparisons and p-values are reproduced from the original summary.

One example window

Mar 31–Sep 27, 2016 · first chronological window; not selected for performance
$31$-69Mar 31, 2016Sep 27, 2016
Selected date: cumulative recorded P/L and reported costs, in dollars
StrategyP/L ($)Costs ($)
Buy & hold SPY+5.740.48
Mean reversion-36.3924.74
Random+16.6122.83
Legacy momentum-60.9049.66

Recorded realized P/L from closed trades; reported costs shown separately. This is not final portfolio balance or net return. Lines step when trades close, not with daily market prices. This single window does not represent the median across all 316 windows.

Before a rerun

I’m considering reopening the experiment with current local and paid models. A rerun needs a passing context check so every model can read the full market view.

← Projects