A gold strategy test should end in one of two ways, both defined before the test begins: graduation to live trading when the minimum sample size and macro coverage are met, expectancy is positive after full realistic costs, the result survives outlier-removal and regime-split checks, and the drawdown observed is tolerable at live size, or abandonment when expectancy is negative, when profitability depends on one exact parameter or a handful of event days, or when the rules cannot be executed in real time.
A written time limit alongside the sample-size and expectancy thresholds prevents the most expensive failure mode of all — testing indefinitely because neither a clear pass nor a clear fail feels comfortable to declare, which costs every week an edge could have been traded rather than just risking a decision on incomplete evidence.
An ambiguous result, where the sample size is met but the outcome barely survives an outlier check, is resolved by extending the identical, unchanged rule set at reduced size for roughly half the original sample rather than by adjusting the rules, and any genuine rule change requires a clean restart with the trade count reset to zero and the prior log archived separately.