The AI Strategy Report Card · July 2026

Eight entries. One prompt. Real backtests.
Two passing grades.

We asked ChatGPT, Gemini, Grok, DeepSeek, Perplexity and Manus (Meta) for their best fully-mechanical BTC/ETH strategy — the identical prompt, first answer only, no retries — then ported each one faithfully and graded it on 6 years of data with our five hard gates. On 2026-07-11 we added the hype of the moment, Hermes Agent, plus the control it demands: the same Claude model bare, same prompt. Here's the honest scoreboard.

ModelGrade (taker costs)Grade (maker costs)OOS tradesNet bp/tradeWin rateR:R
Hermes Agent (Claude backend)Nous Research / AnthropicBB116+120.335.3%2.92
GeminiGoogleBB217+63.434.6%2.55
ChatGPTOpenAIFB42+57.628.6%4.13
ManusMetaFF67−28.937.3%1.25
GrokxAIFF1,132−21.036.2%1.24
DeepSeekDeepSeekFF4,279−20.926.2%0.60
Claude (control)AnthropicFD152−1.136.8%1.63
PerplexityPerplexityFF30−53.430.0%1.16

Net bp/trade and win rate shown at taker costs (20bp round trip incl. slippage — every spec specified market orders). Maker sensitivity = our 4bp archetype-sweep terms. OOS = out-of-sample: the final 40% of six years, untouched by the strategy's design.

Hermes Agent (Claude backend) (Nous Research / Anthropic)

B

New top of the class — added 2026-07-11 after the "best AI for trading" hype cycle. The most-starred agent framework on GitHub, run in safe mode with no tools and Claude Opus 4.8 inside: a 4h Donchian breakout gated by a daily EMA200 regime and ADX, chandelier trail, nothing clever. Every gate passes at both cost models, and its self-predicted win rate (35–45%) landed within a point of reality (35.3%). Read the control row below before crediting the harness.

Gemini (Google)

B

Top of the class — and the simplest spec of the six: a 4h EMA50/200 regime with an EMA50-cross entry, ATR stop, and a ratcheting trail. Positive on both assets, train and test agree, survives its own taker costs.

ChatGPT (OpenAI)

F / B

One gate from passing. Strong positive expectancy and a 4:1 realized R:R, but ETH does all the work — BTC is net negative out-of-sample, so the cross-asset robustness gate fails at taker costs. Its own predictions: R:R ✓, win rate ✗ (predicted 40–50%, got 29%), trade frequency ✗ by 20×.

Manus (Meta)

F

Textbook-clean multi-timeframe trend spec with disciplined risk — and a negative gross edge (−8.9bp before any fees). Not overfit; just edgeless.

Full case study →

Grok (xAI)

F

2,848 trades in six years — the MACD-cross churn it explicitly said it was avoiding. Gross expectancy is a coin flip (−1bp); fees do the rest. Failed four of five gates including survivable drawdown.

DeepSeek (DeepSeek)

F

A 5-minute scalper with 12,842 trades and a gross edge of −0.9bp — pure fee treadmill. Its own writeup predicted a 42% win rate and profit factor 1.35; the 6-year tape says otherwise.

Claude (control) (Anthropic)

F / D

The control for the Hermes row: the SAME model (Claude Opus 4.8), same prompt, bare API, first answer. It chose almost the same skeleton — 4h Donchian breakout, EMA20/50, ADX gate — then buried it under management: a 40% partial at +1.5R, a breakeven move, a 30-bar time stop. 97 time-stop exits later, the edge is gone. One sample each, so the harness delta may be luck — but complexity losing to simplicity is now 8 for 8 on this page.

Perplexity (Perplexity)

F

The most elaborate spec of the six — three timeframes, seven entry conditions, volume confirmation — fired 83 times in six years and lost the most per trade before fees. Complexity isn't edge; here it was just a smaller sample of the same nothing.

What the class of 2026 has in common

  • Everyone brought the same textbook.All six chose trend-following built from EMAs; five of six added RSI. Not one used order-flow, funding, or anything a desk would call alternative data. AI's "best strategy" is the retail canon, formatted beautifully.
  • Complexity anti-correlated with results.The winner was the simplest spec (two EMAs and an ATR). The most elaborate spec finished last per trade. Extra conditions didn't add edge — they subtracted sample.
  • Nobody overfit.Every failing strategy failed honestly — train and test agree. These models write clean, disciplined, professional systems whose expected value is simply not positive. That's the modern trap: AI makes "no edge" look institutional.
  • The two viable entries share a shape: higher timeframe, low frequency, asymmetric exits that let winners run. Costs decide everything at the margin — the same specs graded at maker terms improve by exactly the fee delta, never by more.

Methodology, in full

Each model got the identical prompt asking for its best fully-mechanical BTC/ETH perpetuals strategy with exact parameters, honest costs, and no discretion. We took the first answer — no retries, no cherry-picking. Each spec was hand-ported rule-for-rule with engine-identical conventions: signals on closed bars only, no-lookahead higher-timeframe maps, stop-before-target on same-bar touches, entries exactly as each spec dictates. Non-computable rules (e.g. a reward:risk filter with no defined take-profit) were skipped and disclosed. Every strategy faced the same five gates on the same six years of data: positive out-of-sample expectancy, clears the cost hurdle, robust across assets, survivable drawdown at its own stated sizing, and not overfit.

The 2026-07-11 additions: Hermes Agent v0.18.2 ran in safe mode with all toolsets disabled and Claude Opus 4.8 as its model — one shot, first answer. Because a framework's output is mostly its underlying model, we also graded bare Claude Opus 4.8 over the raw API as a control, same prompt, first answer. The two rows differ by one thing only: the harness. With one sample each, treat the gap as an observation, not a conclusion.

Honest caveats: one sample per model (these are stochastic systems — a different seed writes a different strategy); backtest grades are not live performance, and execution quality is unverified until paper/live fills exist; the two passing-grade entries earn exactly that — a grade, not a deployment recommendation. We'd forward-test before believing anything — so we are.

Forward test — live

Since July 10, 2026, the two gate-passing specs (Gemini and ChatGPT) run through a nightly walk-forward replay on fresh exchange candles. Only trades entered after publication count — no backtest can sneak in, and open positions are never force-closed to flatter the numbers. Both are low-frequency systems, so the sample builds slowly — the live record publishes on our forward-test scoreboard, win or lose. If the grades were luck, that's where it shows.

Grade your own — or your AI's

Free at tessen.ai/grade. If your AI of choice writes you a strategy, make it check its own homework: Tessen is an MCP connector any agent can call — here's the 5-minute setup. 93.5% of everything we grade gets an F. Now you know the frontier models mostly do too.

Want the raw evidence?

The complete raw pack — all 8 faithful strategy ports as readable Python (every ambiguity documented), 16 full grade JSONs with per-gate verdicts at both cost models, and the summary CSV — is free (pay what you want): download the raw pack. Check our work. That's the point.