GECKO

METHODOLOGY

Is it a real edge, or a lucky coin flip?

A backtest can only ever tell you what happened on one specific run of history. It can't tell you, on its own, whether that result reflects a genuine directional edge or is simply what you'd expect to see anyway if the strategy's LONG/SHORT call were no better than chance. A permutation test is how Gecko checks the difference.

The idea, in plain terms

Take the real trades a strategy actually generated — same instruments, same days, same entry prices, same risk taken on each one. The only thing that changes from trade to trade is whether it went LONG or SHORT. Now shuffle that LONG/SHORT label across the trades at random, as if a coin had been flipped for each one instead of the strategy's actual logic, and re-run the numbers. Do that a thousand times, and you get a thousand "what if this had been chance instead of skill" results — a null distribution.

Then rank the real, unshuffled result against that pile of a thousand random ones. If the real result only beats, say, 500 of them, it's entirely ordinary — exactly what a coin flip would produce half the time. If it beats 968 of them, that's a different story: only 3.2% of random shuffles did as well or better, which is a reasonable working definition of "probably not luck."

Why shuffle direction, specifically

ORB's entry rule decides two separate things: whether to trade a given breakout setup, and which way — LONG or SHORT. The permutation test isolates the second decision. It holds every setup exactly as the strategy found it (same instrument, day, entry price, stop distance) and only randomises the directional call, because that call is the part of the rule most likely to be pattern-matching noise rather than a real signal. A stop-loss and take-profit that get flipped to the opposite direction are mirrored around the entry price — same risk distance, opposite side — so a shuffled trade is still a mechanically valid trade, not a nonsense one.

Where this comes from

Randomisation/permutation testing as a general statistical method goes back to R.A. Fisher in the 1930s — it's a standard tool across statistics, not something specific to trading. Its application to validating trading-rule backtests specifically is most closely associated with David Aronson'sEvidence-Based Technical Analysis (2006), which is the direct inspiration for how it's used here. White's "Reality Check" and the Bootstrap Reality Check are more advanced refinements of the same idea, adding a correction for testing many rules at once — worth knowing about, not yet implemented in Gecko.

What it doesn't do

A permutation test only checks whether the directional call beats chance — it says nothing about whether the strategy is profitable after costs, and nothing about whether the decision to trade at all (as opposed to which side) is meaningful. It also doesn't correct for having tried many different strategy variants or instrument subsets and only reporting the ones that looked good — that correction (multiple-comparisons adjustment) is a separate, harder problem, and any post on this site that leans on a permutation-test result will say explicitly whether that correction has been applied.

It's also worth knowing that Gecko's permutation test necessarily uses both LONG and SHORT trades — shuffling direction on a dataset that's 100% one direction is a no-op. The published backtest figures elsewhere on this site are LONG-only (a scope limitation, not a design choice), so a permutation test's "real average R" figure is not directly comparable to the LONG-only average R quoted for the same strategy config elsewhere. Only the percentile — how the real result ranks against the shuffled ones — carries across.