The Question
Every ORB result on this site so far — the baseline, every entry/exit variant, the banks-vs-broad-market permutation test — was run against the original 28-stock universe, over roughly seven years of data (2019-01-01 to 2026-08-11), and, as it turns out, LONG trades only. None of that was a deliberate scope choice for this test; it’s just what every prior ORB post inherited from earlier code and earlier data availability.
Gecko has since grown past both of those limits: a wider, 42-stock basket (orb_group_2, detailed on the scope page) with a decade of history (back to 2016-09-01), and — as described below — a real bug fix that lets ORB trade SHORT for the first time. That raises an obvious question: does the edge found on the narrow slice actually hold up once the strategy is allowed to run at the scale Gecko can now support, in both directions?
The Bug: SHORTs Were Silently Dropped
OrbStrategyApp.runBacktest() filtered every trade plan down to TradeDirection.LONG before handing it to TradeExecutionService.execute() — a single filter line, left over from early development and never removed. OrbEntryStrategy genuinely generates SHORT trade plans on a breakdown below the opening range; they were just silently discarded before ever reaching execution. Every backtest_run published through id 130 — the original baseline, every entry/exit tweak, the banks-vs-broad-market comparison — was a LONG-only result, whether or not that was ever stated as a limitation.
The fix was a one-line deletion: trade plans now pass straight through to execution, regardless of direction. TradeOutcomeEvaluator, the P&L calculation, and position sizing were already direction-correct for SHORT and needed no changes — the bug was purely in what got filtered out before those components ever saw a SHORT trade.
Isolating the Fix: Same Stocks, Same Dates, SHORTs Added
Before scaling anything else up, the fix alone was tested in isolation: the same 5-stock subset (STAN, DEBS, HSBA, DGE, BARC), the same 10-year date range, the same baseline configuration — LONG-only (backtest_run #130) versus the identical setup with SHORTs included (#131).
| LONG-only (#130) | Bidirectional (#131) | |
|---|---|---|
| Trade plans / entered | 3,174 | 9,789 |
| Average R | +0.0309 | −0.0100 |
| Profit factor | 1.08 | 0.98 |
| Gross profit | £120,458.92 | £359,892.16 |
| Gross loss | £111,233.28 | £369,095.46 |
| Total commission | £19,044.00 | £58,734.00 |
| Net result | −£9,818.36 | −£67,937.30 |
Trade count roughly triples once SHORTs are included — expected, since the opening-range breakdown side of the setup was never being taken before. But the direction itself flips: average R goes from modestly positive to modestly negative, and the loss more than sextuples in absolute terms. LONG-only was already net-negative on this subset once commission is counted (as profit-factor-only figures on this site have repeatedly obscured elsewhere) — adding genuine SHORT trades didn’t rescue it, it made it worse. A further round of parameter tuning on this same bidirectional subset (backtest_run #132–142, not tabulated here) stayed flat or negative throughout; nothing found in that sweep recovered a real edge.
Scaling Up: The Baseline Across the Full Basket
With the SHORT-fix confirmed and isolated, the strategy’s plain baseline configuration — no trend filter, no size filtering, nothing tuned — was run against the full 42-stock basket over the full available decade (2016-09-01 to 2026-08-31), bidirectionally. This is backtest_run #143, compared below against the original 28-stock, LONG-only baseline (#88) that started this whole series.
| Category | Metric | 28-stock, LONG-only (#88) | 42-stock, bidirectional (#143) |
|---|---|---|---|
| Trade Activity | Trade plans | 21,869 | 81,402 |
| Trades entered | 21,869 (100.00%) | 81,402 (100.00%) | |
| Execution Outcomes | Take profit | 1,813 (8.29%) | 6,848 (8.41%) |
| Stop loss | 6,899 (31.55%) | 25,504 (31.34%) | |
| Timeout | 13,157 (60.16%) | 49,048 (60.26%) | |
| Ambiguous | 0 (0.00%) | 2 (0.00%) | |
| Position Sizes | Total position value | £223,498,978.72 | £809,476,746.77 |
| Average position value | £10,219.90 | £9,944.19 | |
| Largest position value | £46,689.80 | £226,446.00 | |
| Capital turnover | 44,699.80x | 161,895.35x | |
| Profitability Cross-Check | Wins | 9,652 (44.13%) | 34,919 (42.90%) |
| Losses | 12,208 (55.82%) | 46,462 (57.09%) | |
| Breakevens | 9 (0.04%) | 19 (0.02%) | |
| Profit & Loss | Gross profit | £826,140.99 | £2,974,991.00 |
| Gross loss | £795,344.13 | £2,995,678.31 | |
| Profit factor | 1.04 | 0.99 | |
| Total commission | £131,214.00 | £488,412.00 | |
| Average R | 0.0153 | −0.0027 | |
| Net result | −£100,417.14 | −£509,099.31 |
Three things changed at once between these two runs — the SHORT-fix, the basket (28 stocks to 42), and the date range (roughly 7 years to 10) — so this isn’t a single-variable test the way the isolated #130/#131 pair above is. That’s a real caveat, not a minor one. But the direction of every metric that matters is consistent with the isolated comparison, not contradicting it: profit factor drifts from an already-thin 1.04 down to almost exactly break-even at 0.99, average R flips from modestly positive to modestly negative, and the net loss scales up roughly fivefold — considerably faster than the roughly fourfold growth in trade count, because a larger fraction of the additional trades are themselves losing ones.
Worth being explicit about what didn’t change: the shape of the outcome distribution is almost identical between the two runs. Take-profit rate (8.29% → 8.41%), stop-loss rate (31.55% → 31.34%) and timeout rate (60.16% → 60.26%) are all within a fraction of a percentage point of each other. This isn’t a case where a wider basket behaves qualitatively differently — trades still resolve the same way, in the same proportions, at ten times the volume. It’s the same strategy, doing the same thing, just with the LONG-side edge that used to carry it now offset by an equal and opposite drag from the SHORT side.
It’s also worth noting that even the best-tuned ORB configuration found on the narrow slice — the trend-filtered variant from the banks-vs-broad-market post, which reached average R 0.0697 and profit factor 1.19 on the 5-bank subset — was only barely net-positive there: +£1,450.90 on 2,633 trades (backtest_run #118). That same configuration, run on the full 28-stock universe rather than the bank subset (#106), was already net-negative at −£52,803.88 despite a “good-looking” profit factor of 1.06. The edge this strategy family ever demonstrated was thin and narrowly concentrated even before this test — which makes the flat result at full scale unsurprising in hindsight, if not in advance.
What This Means
The LONG-only bug wasn’t a cosmetic issue — it was flattering every ORB result published before it was fixed. Every profit factor, every average R, every net figure attributed to ORB through backtest_run #130 reflected only half of what the strategy actually does. The isolated #130/#131 comparison shows that including the other half doesn’t just add noise around the same conclusion — it reverses the sign.
Scaling up the universe and the history doesn’t rescue the strategy, and if anything makes the problem more visible. Trade-outcome proportions barely move between the 28-stock and 42-stock baselines; what moves is the net figure, and it moves in the wrong direction, faster than trade count alone would predict. There’s no sign here that a wider, more diverse basket smooths out noise into a cleaner edge — the opposite: more stocks, more history, and honest bidirectional trading converge toward profit factor 1.0, which is exactly what “no edge” looks like.
This doesn’t necessarily mean ORB has zero edge anywhere. The banks-vs-broad-market permutation test found a real, statistically-distinguishable signal on the 5-bank subset specifically — but that test predates the SHORT-fix, and #118’s net result on that same subset was already only marginally positive (£1,450.90) even LONG-only. Whether that specific signal survives being re-tested bidirectionally is an open question this post doesn’t answer, since the isolated fix comparison above used a different 5-stock subset (STAN, DEBS, HSBA, DGE, BARC, not the original banks set of BARC, HSBA, LLOY, NWG, STAN).
What This Led To
This result — a flat baseline at the largest scope and history Gecko currently supports, immediately after fixing a bug that had been silently flattering every number published before it — is the point at which continuing to grind ORB variants stopped looking worthwhile. Effort since has shifted to TFM, which already shows a meaningfully different profile at similar scale: a real, positive net result on nearly 4,700 real trades on the same 42-stock, bidirectional basket this post uses. That’s not a claim TFM’s edge is proven — it hasn’t been permutation-tested the way ORB’s banks result was — only that it’s the more promising avenue to keep testing methodically.
