Continuing from the reward:risk ladder

The reward:risk sweep settled on RR 3.0 as the point that maximises total profit, permutation-validated at the 100th percentile. That left three other exit-plan levers untested against that baseline: which support/resistance zone gets used as the take-profit target, how long a pending order and open position are allowed to run, and how far beyond the confirmation day’s high/low the entry order sits. Each was swept one at a time, same one-factor-at-a-time discipline as the RR sweep itself, against the identical 110-stock backtest_core basket.

The headline figures

Lever Value tested Net profit Profit factor Average R
(baseline) RR 3.0, NEAREST zone, 14-day expiry, 0.10% buffer £12,018 1.23 0.16
Zone target SECOND_NEAREST £16,001 1.25 0.15
Expiry window 7 days £10,062 1.23 0.16
Expiry window 21 days £9,798 1.21 0.16
Expiry window 28 days £10,460 1.22 0.16
Entry buffer 0.05% (tighter) £12,259 1.23 0.17
Entry buffer 0.50% (wider) −£5,322 1.08 0.05

Every configuration above beats 100% of 1,000 random direction-shuffles - the full permutation-test detail for the standout result follows below.

Zone target is the real find here. Targeting the second-nearest support/resistance zone instead of the nearest one lifts net profit by 33% (£16,001 vs £12,018), with more winning trades (929 vs 867) at a similar win rate and a slightly better profit factor. The likely mechanism: the nearest zone is sometimes a thin, low-conviction level - barely two touches, easily brushed and reversed - while the second-nearest is more often a genuinely established level. Since the stop is arithmetically backed out from the target distance (stop = target distance ÷ reward:risk), a more reliable target also means a more sensibly-placed stop, not just a bigger prize.

Expiry window and entry buffer both turned out to already be close to optimal. Every tested alternative to the 14-day expiry window did worse, not better - 7 days cuts profit by trading out of positions too early; 21 and 28 days do worse still. That’s worth noting against this site’s own earlier speculation (see CLAUDE.md’s design notes) that a shorter window might trade fewer, higher-quality entries - tested now, and it doesn’t hold. The entry buffer shows a similar pattern from the other side: tightening it to 0.05% is a small, plausible improvement (£12,259), but widening it to 0.50% is an outright collapse - profit factor falls to 1.08, average R crashes from 0.16 to 0.05, and the run posts a net loss of £5,322. A wider buffer means chasing price further before the order triggers, which degrades every trade’s entry price and therefore its whole risk:reward shape - not a subtle effect once you look at the number.

The permutation test, for the winning configuration

Baseline (RR 3.0, NEAREST) SECOND_NEAREST
Real average R 0.1613 0.1495
Null distribution mean −0.0165 −0.0025
Null distribution range −0.090 to 0.061 −0.078 to 0.083
Real result beats 1,000 / 1,000 shuffles 1,000 / 1,000 shuffles
Percentile 100.0th 100.0th

Interesting that the real average R is actually slightly lower for SECOND_NEAREST (0.1495 vs 0.1613) despite its higher total net profit - a reminder that average R and total profit answer different questions (quality of a typical trade vs. how much money the whole run made) exactly the way the win/loss breakdown in the RR-ladder post discussed. SECOND_NEAREST produces more trades that clear the entry bar in the first place (2,806 vs 2,862 signals evaluated, but more of them entered and more of those won), which is what actually drives the bigger total profit here.

Alongside the exit-plan sweep, and using the same RR 3.0 baseline, two indicators already being computed by the codebase - but not yet used for anything beyond an optional hard entry gate - were checked against real trade outcomes for the first time. This isn’t a “which config is best” question like the rest of this post; it’s the first real test of whether the trade-grading idea sketched out in recent design conversations actually has something to grade with, using gecko’s own data rather than assuming a finding from elsewhere applies.

ADX (trend strength) shows a hump, not a straight line. Splitting trades into thirds by their confirmation-day ADX reading:

ADX tercile Trades Average R Net profit
Lowest 954 0.13 £2,087
Middle 954 0.23 £8,460
Highest 954 0.13 £1,470

The middle third clearly outperforms both extremes, which underperform each other by almost identical amounts. A plausible read: very low ADX means barely any trend to ride in the first place; very high ADX may mean the move is already extended by the time TFM’s pullback-style entry actually fires. The existing ADX gate (entry.tfm.minimumAdx) only checks a floor - if this pattern holds up, a bounded range would fit the data better than “more trend strength is always better.”

Market cap shows a much larger, cleaner effect. Splitting the same trades by the traded instrument’s market-cap tier:

Market cap Trades Average R Net profit
Large cap 2,334 0.175 £11,774
Mid cap 383 0.149 £1,125
Small cap 145 −0.029 −£881

Large-cap trades carry almost the entire run’s profit. Small-cap trades - on a much thinner sample - are a net loser. This directly contradicts a well-known finding from Bulkowski’s Trading Classic Chart Patterns (his own trendline research found small caps outperforming, in both directions tested) - a useful reminder that a finding from a different strategy, market and era doesn’t automatically transfer, rather than something to take on faith either way.

What’s still unproven

Neither the ADX nor the market-cap result has been permutation-tested. Both are a same-run cohort comparison - genuinely informative, but a weaker form of validation than the direction-shuffle test the rest of this site relies on. Worth doing properly before either is treated as settled, particularly the ADX result, where the effect size is smaller and the mechanism more speculative.

The market-cap sample is thin. 145 small-cap trades is a real number, not a handful, but it’s an order of magnitude smaller than the large-cap sample it’s being compared against - backtest_core itself is already large-cap-heavy (90 of 110 instruments). The direction and the sign flip (profitable to loss-making) look substantial enough to be worth testing as an actual filter rather than dismissed as noise, but “substantial-looking on one run” isn’t the same bar as the permutation-validated results elsewhere on this site.

None of these levers have been tested together yet. RR 3.0, SECOND_NEAREST targeting, and a market-cap exclusion filter have each shown a real effect individually, but never in combination - the next natural test, not yet run as of this post.