Back to blog
Spoke · Trust & Safety

Backtest Lies: Overfitting, Survivorship Bias and Cherry-Picked ROI

Image Placeholder · Hero 16:9
A lion staring down a glowing teal chart line that splits into a fake smooth curve and a jagged honest one, amber spark marking the divergence, dark cinematic fintech-noir, near-black background, electric emerald-teal glow (#3ecf8e) as key light, subtle amber-gold spark accent (#f5a70b), atmospheric misty forest and mountains, photorealistic, dramatic rim lighting, 16:9 cinematic wide composition, clean plate, no text
Quick Answer

Backtest overfitting is tuning a trading strategy so tightly to past data that it fits historical noise instead of a real edge, producing a gorgeous backtest that falls apart on new data.

  • Three tricks fake a backtest: overfitting, survivorship bias, and a cherry-picked window.
  • An honest backtest is out-of-sample, spans bull and bear, and includes fees and slippage.
  • Only live PnL cannot be faked, because you cannot curve-fit a trade that has not happened yet.

A beautiful backtest is the easiest thing in trading to fake.

Start with the biggest trick: overfitting.

Give a strategy enough knobs and an optimiser finds the exact settings that would have printed money on the last three years. It had the answer sheet.

That is a curve-fit to noise, not a prediction. It breaks the moment the market stops repeating itself.

Drag the slider below. Turn the parameters up and watch the backtest climb while the live result caves.

Interactive · The Overfitting Machine
Parameters (knobs)14
Backtest (in-sample) Backtest promise Live (out-of-sample)
In-sample backtest
Out-of-sample (live)
Backtest fit (R²)
Illustrative model of the bias-variance tradeoff, not TRAPR data. As parameters rise, in-sample fit approaches 100% while the out-of-sample result degrades and eventually turns negative. Not a forecast, and not financial advice.

Any strategy can look genius in hindsight. Hindsight is not a track record.

The tell is in the shape.

An equity curve that is too smooth, a Sharpe ratio that looks impossibly high, and no mention of testing on data the strategy never saw.

When it looks too good, it was fit too hard.

So treat any backtest as a hypothesis. A live record is the only thing that confirms it.

Trick two

What is survivorship bias in a crypto backtest?

Survivorship bias is testing a strategy only on coins that still exist, silently deleting the ones that died. Test a bot on 2026's surviving coins and you baked in a winner's list it could not have known in advance. The result flatters the bot by hiding every coin that went to zero. In crypto that lie is enormous.

0
Of tokens launched since mid-2021 no longer actively traded
CoinGecko, Jan 2026
0
Of all those failures happened in 2025 alone
CoinGecko, Jan 2026

CoinGecko tracked roughly 20.2 million tokens launched between mid-2021 and the end of 2025.

By early 2026, 53.2% were no longer actively traded.

A backtest run on "the coins we trade today" quietly excludes every dead one, so it grades the strategy on a pre-filtered pool of survivors.

In real time it would have held some of the losers. Those losses never reach the pretty chart.

If the backtest only trades coins that survived, it was handed the answers before the test began.

Trick three

What is a cherry-picked backtest window?

Cherry-picking is choosing the test period that makes a strategy shine. Show a DCA bot only the 2023 to 2024 recovery and it looks unstoppable, because averaging down always works when price eventually rips upward. The honest test spans bull, bear, and chop.

Every strategy has a favourite weather.

A trend bot glows in an uptrend. A grid bot glows in chop. A DCA ladder glows in any market that recovers.

Feed it only its best season and it looks like a machine that cannot lose.

Then the regime shifts. The same bot that "never lost" meets a sustained downtrend and bleeds, because it was never tested there.

The same illusion hits your account statement, which is why a "winning" bot can still leave you down. We unpack that in why your bot shows profit but your balance is down.

One favourable window is not a track record. It is a screenshot with good lighting.

The fix

What does an honest backtest look like?

An honest backtest tests out-of-sample on data the strategy was never tuned on, spans the full market including bear phases, and subtracts real fees and slippage. It is framed as a hypothesis to confirm live, not as a guarantee. Those four traits separate research from marketing. Reject anything missing one.

A backtest worth trusting
Out-of-sample by default. Tuned on one slice of history, then tested on a separate slice it never saw. Same data for both is worthless.
Full span, not a flattering window. Bull, bear, and sideways. A strategy that survives only one regime is a bet on that regime continuing.
Fees and slippage included. Every simulated trade pays the real cost of trading. A zero-cost backtest is fiction, especially for anything that trades often.
Framed as a hypothesis. The operator says the backtest suggests an edge and the live record proves it. Anyone presenting a backtest as a promise is selling.

For the mechanics of doing this yourself, see how to backtest a crypto trading strategy.

The proof

Why does live PnL beat any backtest?

Live profit-and-loss is the only proof that cannot be curve-fit, because it happens in real time on real prices with no hindsight. A backtest can be tuned, cherry-picked, and survivorship-washed. A live record, wins and losses included, cannot be.

Nobody can optimise against a candle that has not printed yet.

So a live track record removes every trick in this post at once. No refitting. No deleted losers. No favourite window.

Just what actually happened.

You cannot overfit a trade that has not happened yet. That is why live PnL is the only honest proof.

The catch is that live records are unflattering.

They show the drawdowns, the flat stretches, the losing months.

Which is exactly why they are believable.

Treat every backtest as a claim to be tested. Before you believe a single equity curve, ask three things.

1
Was it tested out-of-sample?
On data the strategy was never tuned on. If tuning and testing used the same history, the result describes the past, not the future.
2
Does it include the dead coins and the bear markets?
Survivorship-bias free, across the full span. If it only trades survivors in a recovery window, it was handed the answers.
3
Is there a live record behind it?
Real trades, wins and losses shown. A backtest is a brochure until a live PnL, drawdowns and all, stands behind it.

If you are still weighing whether automated trading is legitimate at all, start with the pillar, are crypto trading bots a scam?, or the field guide to spotting a crypto trading bot scam.

This article is educational and is not financial advice. Crypto is high-risk and you can lose money, including with any automated strategy. Backtests, historical illustrations, and the interactive model above do not predict future results. Do your own research and consider your own situation before investing.

* Pitch warning
Backtest it yourself, then arm the OX

You just saw the rule: a backtest is a hypothesis, and only a live record proves it.

TRAPR is built around that order of proof. You build a strategy, run a one-click backtest over years of real market data, and arm it live only when the numbers convince you. The shipped presets are shown with their backtests as evidence, never as a promise.

The core long loop is the OX, TRAPR's AUTO TRADE LONG preset, on the Trader tier at $49 a month.

1
Build the strategy
Set the rules in the Lab, no code, everything explicit and fixed up front.
2
Backtest on years of real data
One click runs it over full-span history, fees included, so you see the bad stretches too.
3
Arm it only when convinced
Nothing goes live on your money until the numbers earn it.
4
Judge it on the live record
The backtest was the hypothesis. The live PnL, drawdowns and all, is the verdict.

A backtest here is evidence, never a guarantee. See the loop on the OX or start free.

Illustration only. Not a real backtest, not a return promise, and not financial advice.

FAQ

Common Questions About Backtest Reliability

Is backtesting reliable?+
Backtesting is useful as a hypothesis, not as proof. A backtest is reliable only when it is run out-of-sample on unseen data, spans the full market including bear phases, and includes fees. Even then, live results are what confirm it.
What is overfitting in trading?+
Overfitting is tuning a strategy so tightly to past data that it fits historical noise instead of a real edge. Overfit strategies show near-perfect backtests, then fall apart on new data because they memorised history rather than learning a pattern.
What is survivorship bias in a crypto backtest?+
It is testing a strategy only on coins that still exist, silently excluding the many that collapsed or delisted. CoinGecko found 53.2% of tokens launched since mid-2021 are no longer actively traded, so ignoring the dead ones badly flatters any result.
How can you tell if a bot backtest is overfit?+
Look for three tells: an equity curve that is unusually smooth, a Sharpe ratio that looks too high to be real, and no mention of out-of-sample testing. When results look too good and the strategy was never tested on unseen data, it was almost certainly fit too hard.
Does a high backtest Sharpe ratio mean a good bot?+
Not on its own. A very high Sharpe usually signals a curve fit, not an edge, because real strategies are noisier than that. Experienced traders treat an impossibly high backtest Sharpe as a warning sign rather than proof.
Why does live PnL matter more than a backtest?+
Because a live record happens in real time with no hindsight, so it cannot be curve-fit, cherry-picked, or survivorship-washed. It shows real drawdowns and losing periods, which makes it the only proof of an edge that a backtest cannot fake.
Sources
  1. CoinGecko, "Dead Coins" research, reported by CoinDesk (January 2026). Of roughly 20.2 million tokens launched between mid-2021 and the end of 2025, 53.2% are no longer actively traded, and 86.3% of those failures occurred in 2025.
  2. TRAPR / TAP fact-sheet. Rules-based DCA-patience design: 3 safety orders by default, leverage optional and off by default, an 80% disaster-stop on leveraged positions, with backtests shown full-span and published live PnL.
  3. Reddit r/algotrading and r/algotradingcrypto (community voice-of-customer). The standing skepticism this post answers: "you overfit the shit out of it", "is this survivorship bias free", and "run it live" as the credibility test.
GET THE PHASE ALERT

The market flips. You get the email.

We email you automatically when our algorithm warns of a market regime change, so you can trade accordingly.

Phase alerts are market information, not financial advice.

Judge us on the live record, not a backtest

TRAPR runs a fixed, rules-based playbook and publishes its live PnL, drawdowns included. Run it on your own exchange, on your own money, and hold it to the standard in this post.

Start free