Two people showed me a strategy in the same week. The first had never tested his — he read about the setup, it made sense to him, and he funded the account on Monday. The second had a spreadsheet with 380 backtests in it and he was proud of the one sitting at the top. I didn't believe either result. They'd made the same mistake from opposite ends. Here's the number that changed how I think about this. David Bailey and Marcos López de Prado worked out how much history a backtest actually needs before the winner you picked means anything. With five years of data, you get roughly 45 independent strategy configurations. Go past that, and a strategy showing an in-sample Sharpe ratio of 1.0 has an expected out-of-sample Sharpe of zero. Not lower. Zero. Sharpe ratio is return per unit of volatility, and 1.0 is the level where a fund manager stops apologizing for his year. Bailey's point is that you can manufacture that number out of pure noise if you're willing to click "run" enough times. Think about what one trial actually is. You move the moving average from 50 to 55 and rerun it. You widen the stop by half a percent. You add a volume filter, then take it out, then try the whole thing on the 4-hour instead of the daily. Every one of those is a trial. The counter doesn't reset because you made coffee in between. Forty-five is a Saturday afternoon for most people — I've burned through more than that before lunch. The mechanism is boring once you see it. With enough attempts, randomness produces a winner. Test 45 variations of a rule with no edge whatsoever and one of them will look excellent across five years, purely because something has to come first. You didn't discover an edge. You sorted noise and kept the top of the pile. So why does almost everyone get this backwards? Because effort feels like rigor. The guy with 380 backtests is certain he's been more careful than the guy with none, and in a sense he has been — he's just been thorough at a process that punishes thoroughness unless you account for it. Bailey later published a correction for this called the Deflated Sharpe Ratio, and the idea behind it is uncomfortable. A Sharpe number on its own is meaningless. You need to know how many attempts produced it. The same 1.4 means two entirely different things at 5 trials and at 500. Nobody reports the trial count. Not the person selling you a bot, not the account posting equity curves, and not you in your own notes three months later when you've forgotten how many versions you chewed through to get there. The untested trader isn't safer, by the way. He thinks he skipped a formality. What he actually did was open a backtest with a sample size of one, running live, funded, in a market that charges for lessons. His strategy is going to get tested either way. He just picked the most expensive method available. If I were starting over, I'd change one thing. Before opening the backtester, write down how many variations you're allowed to run — and then log every single one, including the ones you killed after four seconds because the equity curve looked ugly. Those count. Especially those. Then hold something back. Choose your rule using 2019 through 2023 and don't let yourself look at 2024 onward until the parameters are locked. If the held-out stretch falls apart, you didn't find an edge. You found a good fit for one particular slice of history, which is a completely different thing that happens to look identical on a chart. I keep the trial count for our BTC strategy in the same document as the results, because the number is part of the result.
Why More Backtesting Can Make Your Trading Strategy Worse
Full Article
Original Source
Read the full article at Hackernoon →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.