Sample size is the number of trades a performance statistic is built from, and it governs how much of that statistic is signal rather than noise. Twelve winners out of twenty produces a sixty percent win rate whose confidence interval spans both excellent and unprofitable systems, so the figure supports no conclusion at all.
Strategies with lower win rates require larger samples because their profitability depends on infrequent large winners that a short run may not contain. Segmentation compounds the problem: splitting two hundred gold trades across three sessions and four setups leaves cohorts of roughly a dozen, each individually meaningless, which is the mechanism by which XAUUSD strategies become overfitted to history and fail in forward testing.