A gold strategy test can look profitable without holding a real edge for five distinct reasons: the test window happened to fall inside one favourable volatility regime, a single outsized winning trade carried an otherwise flat sample, trading costs like rollover spread widening and news-driven slippage were left out of the model, the rules were quietly changed partway through the test, or key parameters were tuned precisely to the noise of one narrow volatility period rather than to a robust relationship.
Each false positive has a fast, mechanical check: splitting the sample into regime thirds exposes regime luck, removing the three largest winning trades exposes outlier dependence, re-applying time-varying spread and slippage exposes unmodelled costs, timestamping rule changes exposes mixed-rule pooling, and sweeping parameters by 10-20% exposes curve fitting.
A gold test that survives all five checks has cleared the most common ways a result misleads its own author, and modest, cost-adjusted expectancy that persists across trend, range and shock conditions is a far more trustworthy signal than a smooth, high-looking equity curve produced by an untested combination of regime luck, outlier dependence and an over-fitted parameter set.