The Comparison
Two things landed within days of each other this week: TFM’s exit-tuning series found a configuration that nets real money on the full research basket, and ORB’s equivalent full-scale test went the other way, converging to a flat 0.99 profit factor once it was scaled up and genuinely bidirectional. Put side by side, the comparison is the most encouraging result on this site so far — but it’s worth being precise about exactly which numbers are being compared, and how fair that comparison actually is.
| Metric | ORB best-ever (#118, 5-bank subset, LONG-only) | ORB same-scope (#143, 42-stock, bidirectional) | TFM baseline (#150, 42-stock, bidirectional) |
|---|---|---|---|
| Trades entered | 2,633 | 81,402 | 4,433 |
| Profit factor | 1.19 | 0.99 | 1.20 |
| Average R | 0.0697 | −0.0027 | 0.1157 |
| Net result | £1,450.90 | −£509,099.31 | £21,214.65 |
TFM’s £21,214.65 doesn’t just beat ORB’s flat, full-scale number — it beats ORB’s best-ever result, full stop, the trend-filtered banks configuration that took a permutation test and a dedicated post to validate. And it does that on profit factor and average R too, not just net result, where trade count alone could otherwise explain a lot of the gap.
Why This Comparison Is Fair — and Where It Isn’t
The middle column is the important one. ORB’s best-ever number (#118) was found on a hand-picked 5-stock subset, LONG-only, before the SHORT-filtering bug was fixed — exactly the kind of narrow, favourable slice that the scaling post showed doesn’t survive being widened. TFM’s #150, by contrast, was run on the same 42-stock basket, the same bidirectional trading, and a comparable historical window as ORB’s own flat #143 result. That’s the fairest comparison available on this site: same scope, same rules, both strategies tested as hard as Gecko currently tests anything. TFM holds up under those conditions; ORB’s best number never was.
Where it isn’t fair yet: TFM has only ever been tested at this scope, so there’s no equivalent to ORB’s #88 → #143 journey to show whether TFM’s edge also erodes as scope widens further or the test gets harder. It’s entirely possible TFM is exactly where ORB was before the SHORT-fix and the 42-stock scale-up — a result that looks strong on the hardest test it’s faced so far, without yet having faced a harder one.
The Permutation Test
The two biggest caveats above were “no permutation test yet” and “a small, noisier sample” — and a permutation test addresses both at once, since checking a real result against its own null distribution is exactly how you find out whether a small sample’s good number is signal or just what a noisy sample looks like. The same technique that validated ORB’s banks result — shuffle which trades are LONG vs SHORT 1,000 times, mirror each shuffled trade’s price levels to keep its reward:risk profile intact, and see how the real, unshuffled result ranks against that null distribution — was built for TFM and run against the exact #150 configuration and scope.
| ORB broad universe (28 stocks) | ORB banks (5 stocks) | TFM (42 stocks) | |
|---|---|---|---|
| Real average R (bidirectional) | 0.0071 | 0.0289 | 0.1157 |
| Null distribution mean | 0.0016 | 0.0046 | 0.0825 |
| Null distribution stddev | 0.0056 | 0.0131 | 0.0149 |
| Real result beats | 837 / 1,000 shuffles | 968 / 1,000 shuffles | 989 / 1,000 shuffles |
| Percentile | 83.7th (fails) | 96.8th (passes) | 98.9th (passes) |
TFM clears the same bar ORB’s banks result needed to earn “genuine signal” status, and clears it more comfortably — 98.9th percentile against 96.8th, with only 1.1% of random direction-shuffles matching or beating the real result. The higher null-distribution mean here (0.0825, versus 0.0046–0.0016 for ORB) is expected and not a red flag: TFM’s SupResRRExitPlanStrategy targets a fixed 2:1 reward:risk ratio on every trade by construction, so any random direction call on this dataset already carries a structurally favourable payoff shape — the permutation test’s whole point is to check whether the real, unshuffled call still beats that inflated baseline, and it does, decisively.
Building this test surfaced a real bug worth being transparent about: the first version of TFM’s shuffle logic left a flipped trade’s STOP entry price unchanged, which is correct for ORB’s MARKET_AT_CLOSE orders but produces a malformed order for TFM’s high/low-anchored STOP entries — a SELL STOP left sitting above market where a BUY STOP correctly was. That bug inflated the null distribution so badly the real result beat 0 of an initial 5 shuffles, which would have wrongly suggested TFM’s directional call was worse than random. Fixed by having TfmEntryStrategy compute and carry through what the entry price would have been for the opposite direction, so a shuffled trade’s price gets reconstructed correctly instead of mirrored blindly - full detail in the code, not repeated here.
What’s Still Unproven
Two caveats resolved, one remains:
No permutation test yet — resolved above. TFM’s directional call is now shown to be statistically distinguishable from chance, on the same 42-stock, bidirectional scope the net-result comparison uses.
Smaller, noisier sample — also addressed by the permutation test itself: a null distribution built from this exact 4,433-trade sample is what “beats 98.9% of shuffles” is measured against, so the small-sample concern is priced into the result rather than left as an open question.
The config was chosen from a sweep. #150 (NEAREST target, 14-day expiry) was the best of seven configurations tested across two exit-plan levers, and the original post was explicit that nothing rules out a better result at some untested point in between — 10 or 12 days, say, or a zone-selection strategy between NEAREST and SECOND_NEAREST. Picking the best of several tested points and calling it “the finding” carries some of the same multiple-comparisons risk ORB’s own permutation-testing methodology exists to guard against, even though this sweep was principled and pre-specified rather than mined after the fact. This is the one caveat a permutation test can’t address by itself — it would need testing the untested points in between, not re-testing the winner.
What This Opens Up
This is now the first strategy on this site with a clearly positive net result and a permutation test confirming that edge is real, not noise — a materially stronger position than “the biggest number so far.” That’s enough to justify treating TFM as the priority going forward rather than a side experiment — which is already the direction effort has shifted. The natural next steps: narrow in on the exit-plan sweep between the tested points rather than treating NEAREST/14-day as final, and eventually re-run the whole thing on a wider or longer scope the way ORB just was, to see whether this edge is durable or just hasn’t been tested hard enough yet to break.
