prop-challenge-lab research console

A falsifiable record: what was asked, what was measured, and what the instrument was doing when it measured it. Rendered from the register at 2026-08-31T15:14:19+00:00. A dated artifact, not a live view.

engine_sha
ccc11274ee7264db4a21c298913aa3f9592cc99c
generated
2026-08-31T15:14:21+00:00
archive objects
1,307

1 · Controls

Run first, and shown first. Nothing below is worth reading until the instrument has demonstrated it can find an edge that is there and report nothing where there is none. These recompute on every build; they touch no market data, no key and no network.

Positive control pass

P(pass) = 1.00

A planted edge, in a synthetic world. A harness that cannot find an edge that is there cannot be trusted to report its absence either — which is the claim this whole lab rests on.

Negative control pass

P(pass) = 0.00 · random-entry null 0.00

A dead world, built with the impulse term set to zero. An earlier “no-edge” world still carried a harvestable kick, so it was never a negative control at all — found in 2026-07, and it invalidated a gate.

Calibration gate pass

measured from price +0.0735 0.5σ calibrated
measured from a level +0.7040 5.0σ fails the dead world

Same data, same arithmetic, one different reference price. The dead world has no forward drift whatsoever, so the second reading is the measurement rather than the market: a market that froze solid at the trigger would still print it. That defect produced a real result here at six times its detectable floor, significant on both instruments, both sides, both horizons. No significance test asks the question this gate asks.

Rule profile pass

example-50k · effective 2026-07-02 · account 50,000 · target 3,000 · trailing DD 2,000 · guard 1,000 · min 1 day(s)

The rules are input, and they are dated. assert_fresh refuses a snapshot older than your tolerance rather than warning: a provider retired a plan tier while its own pages still advertised it, and an undated rule set is a silent expiry. These ship as a worked example of a geometry, not as anyone's current terms.

2 · Position

24 hypotheses registered
26 archived runs
24 scored results
$0.00 recorded spend
detectable 10null 11precise but immaterial 3

How each was resolved

24 of 24 hypotheses carry a resolution (25 records; a correction appends rather than edits) record. 14 of 24 rest on a document rather than a computation — their effect sizes live in the cited writeup and were never entered into the register, so this page shows none for them.

measured7read from a scored result in an archived run
audit1a judgement with no estimate and no floor
superseded2replaced before it was ever run
documented14stated in prose, never computed into the register
0.62 alpha allocated across 24 registrations

12.4× the conventional 0.05 single-test allowance. Alpha is a non-renewable resource: a finding at attempt 24 must clear a materially higher bar than the same finding at attempt 3. The gauge is here so that bar is visible rather than remembered.

3 · Findings

Every scored result in the archive, shown against its own detectable floor. The shaded band is the region in which an effect cannot be told apart from nothing; the bar is the 95% interval and the dot is the estimate. “Not significant” is two different findings — an interval that crosses zero but clears the floor is inconclusive and means get more data, while one inside the floor is null and means the question is answered.

cannot be told apart from nothing zero
H-RECLAIM · 2026-08-30-audited

H-RECLAIM COMPARATOR - sealed (unfiltered) expectancy, per trade

null
estimate
-0.03592 R
95% CI
[-0.0835, +0.0135]
floor
0.08996991758658018 -0.40× floor
n
3,161 → 1,663 independent-equivalent

the arm the filter was supposed to improve on.

H-RECLAIM · 2026-08-30-audited

H-RECLAIM PRIMARY - reclaim-filtered expectancy, per trade

detectable
estimate
-0.23931 R
95% CI
[-0.2910, -0.1843]
floor
0.10378428725965629 -2.31× floor
n
1,964 → 1,033 independent-equivalent

the filtered strategy's own expectancy. Compare with the sealed arm below: the intervals do not overlap.

H-RECLAIM · 2026-08-30-audited

H-RECLAIM SECONDARY - paired reclaim minus sealed, shared days only

detectable
estimate
+0.24266 R
95% CI
[+0.1736, +0.3115]
floor
0.1334909642818425 +1.82× floor
n
1,964 → 1,033 independent-equivalent

POSITIVE and it does NOT rescue the filter. This conditions on days a reclaim occurred, which are days price first went through the stop -- exactly where the sealed entry does worst. It answers 'given a reclaim happened, was the later entry better', not 'is the filter worth applying'.

T1-TARGET-LADDER · 2026-08-30-ladder-v2

T1 target_r=1.0 - net R per trade

detectable
estimate
-0.10285 R
95% CI
[-0.1324, -0.0705]
floor
0.05262286279943919 -1.95× floor
n
3,384 → 1,781 independent-equivalent

win rate 0.476; ORB vehicle, stop_range=0.5, risk=$175

T1-TARGET-LADDER · 2026-08-30-ladder-v2

T1 target_r=1.5 - net R per trade

detectable
estimate
-0.09735 R
95% CI
[-0.1340, -0.0577]
floor
0.06357818583823024 -1.53× floor
n
3,384 → 1,781 independent-equivalent

win rate 0.385; ORB vehicle, stop_range=0.5, risk=$175

T1-TARGET-LADDER · 2026-08-30-ladder-v2

T1 target_r=2.0 - net R per trade

detectable
estimate
-0.08511 R
95% CI
[-0.1281, -0.0398]
floor
0.07297645225758571 -1.17× floor
n
3,384 → 1,781 independent-equivalent

win rate 0.330; ORB vehicle, stop_range=0.5, risk=$175

T1-TARGET-LADDER · 2026-08-30-ladder-v2

T1 target_r=3.0 - net R per trade

precise but immaterial
estimate
-0.06694 R
95% CI
[-0.1186, -0.0142]
floor
0.08789775560052443 -0.76× floor
n
3,384 → 1,781 independent-equivalent

win rate 0.268; ORB vehicle, stop_range=0.5, risk=$175

T1-TARGET-LADDER · 2026-08-30-ladder-v2

T1 target_r=4.0 - net R per trade

null
estimate
-0.03717 R
95% CI
[-0.0956, +0.0236]
floor
0.1000020079817115 -0.37× floor
n
3,384 → 1,781 independent-equivalent

win rate 0.242; ORB vehicle, stop_range=0.5, risk=$175

T1-TARGET-LADDER · 2026-08-30-ladder-v2

T1 target_r=6.0 - net R per trade

null
estimate
-0.02128 R
95% CI
[-0.0893, +0.0523]
floor
0.11361893769086785 -0.19× floor
n
3,384 → 1,781 independent-equivalent

win rate 0.223; ORB vehicle, stop_range=0.5, risk=$175

T2-COST-CONDITIONED · 2026-08-30-cost-ladder

T2 cost_R<=0.05 - net R per trade

null
estimate
+0.02296 R
95% CI
[-0.0361, +0.0840]
floor
0.11543800752505227 +0.20× floor
n
1,226 → 645 independent-equivalent

kept 36.2% of trades

T2-COST-CONDITIONED · 2026-08-30-cost-ladder

T2 cost_R<=0.07 - net R per trade

null
estimate
-0.01152 R
95% CI
[-0.0639, +0.0406]
floor
0.09657798063502403 -0.12× floor
n
1,836 → 966 independent-equivalent

kept 54.3% of trades

T2-COST-CONDITIONED · 2026-08-30-cost-ladder

T2 cost_R<=0.1 - net R per trade

null
estimate
-0.03211 R
95% CI
[-0.0804, +0.0179]
floor
0.086290625752713 -0.37× floor
n
2,377 → 1,251 independent-equivalent

kept 70.2% of trades

T2-COST-CONDITIONED · 2026-08-30-cost-ladder

T2 cost_R<=None - net R per trade

detectable
estimate
-0.08511 R
95% CI
[-0.1281, -0.0398]
floor
0.07297645225758571 -1.17× floor
n
3,384 → 1,781 independent-equivalent

kept 100.0% of trades

T4-COST-FILTER-AT-LOW-TARGET · 2026-08-30-low-target

T4 cost_R<=0.05 - net R per trade

null
estimate
-0.02179 R
95% CI
[-0.0635, +0.0195]
floor
0.06945045557523696 -0.31× floor
n
1,226 → 887 independent-equivalent

kept 36.2% of trades

T4-COST-FILTER-AT-LOW-TARGET · 2026-08-30-low-target

T4 cost_R<=0.07 - net R per trade

precise but immaterial
estimate
-0.03865 R
95% CI
[-0.0758, -0.0013]
floor
0.05859399141997629 -0.66× floor
n
1,836 → 1,328 independent-equivalent

kept 54.3% of trades

T4-COST-FILTER-AT-LOW-TARGET · 2026-08-30-low-target

T4 cost_R<=0.1 - net R per trade

precise but immaterial
estimate
-0.04939 R
95% CI
[-0.0847, -0.0133]
floor
0.052568626722908414 -0.94× floor
n
2,377 → 1,720 independent-equivalent

kept 70.2% of trades

T4-COST-FILTER-AT-LOW-TARGET · 2026-08-30-low-target

T4 cost_R<=None - net R per trade

detectable
estimate
-0.10285 R
95% CI
[-0.1324, -0.0705]
floor
0.04488494774988622 -2.29× floor
n
3,384 → 2,448 independent-equivalent

kept 100.0% of trades

X1-COMPRESSION · r1

X1 PRIMARY — contraction run (prior days) -> next-day range / ATR

detectable
estimate
-0.08900 Spearman r
95% CI
[-0.1345, -0.0431]
floor
0.07 -1.27× floor
n
3,442 → 1,811 independent-equivalent

positive = the mechanism holds; negative = volatility persists across days instead

X1-COMPRESSION · r1

X1 SECONDARY (CONFOUND CHECK) — opening-range ratio -> same-day post-range expansion

detectable
estimate
+0.35033 Spearman r
95% CI
[+0.3093, +0.3901]
floor
0.07 +5.00× floor
n
3,442 → 1,811 independent-equivalent

a POSITIVE reading is expected even if the mechanism is false: intraday volatility persists, so a quiet open means a quiet day

X2-OVERNIGHT-GAP · r1

X2 CONTROL (CONFOUND) — |gap| -> |forward move|

detectable
estimate
+0.16305 Spearman r
95% CI
[+0.1181, +0.2073]
floor
0.07 +2.33× floor
n
3,480 → 1,831 independent-equivalent

a large POSITIVE is expected even if the primary is zero: a big gap means a busy day. Volatility persistence, NOT directional information

X2-OVERNIGHT-GAP · r1

X2 PRIMARY — signed gap -> signed forward return, first session hour

null
estimate
+0.00249 Spearman r
95% CI
[-0.0433, +0.0483]
floor
0.07 +0.04× floor
n
3,480 → 1,831 independent-equivalent

positive = incomplete absorption (continuation); negative = overreaction (gaps fade)

X2-OVERNIGHT-GAP · r1

X2 SECONDARY — same predictor, held to the session close

null
estimate
-0.00225 Spearman r
95% CI
[-0.0481, +0.0436]
floor
0.07 -0.03× floor
n
3,480 → 1,831 independent-equivalent

pre-specified secondary horizon

X3-PRINT-SIZE · r1

X3 CONTROL (CONFOUND) — large-print share -> |forward move|

null
estimate
-0.05252 Spearman r
95% CI
[-0.1387, +0.0345]
floor
0.13 -0.40× floor
n
970 → 510 independent-equivalent

a large POSITIVE is expected even if the primary is zero: a chunky tape is a busy tape. Third experiment running, third such control

X3-PRINT-SIZE · r1

X3 PRIMARY — large-print share -> signed forward return, next hour

null
estimate
+0.02596 Spearman r
95% CI
[-0.0610, +0.1125]
floor
0.13 +0.20× floor
n
970 → 510 independent-equivalent

positive = large prints mark flow that extends; negative = they mark exhaustion/absorption

Audit verdicts

Not scored effects — these return a judgement, not a number with an interval, so they have no floor to be measured against. They are here because an audit can clear or void a sealed verdict, which makes them load-bearing in a way their plainness hides.

hypothesisrun nverdict
Z04-ORB-OBTAINABLE2026-08-01-audit6,774CLEAN - ORB does not share the fade's defect

4 · Figures

Rendered by make plots from the same register. Embedded rather than linked, so the page stays one file.

What edge the rules actually require

2026-08-30T13:59:23.239624 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
Win rate against payoff ratio, with the pass gate marked. Asymmetry is the lever, not hit rate: at 2.0R the rules clear at a 45% win rate; at 1.0R they need 60-65%, which is exactly where most retail strategies live.

Why a rising curve is a warning, not a result

2026-08-30T13:59:25.439398 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
P(pass) against risk per trade. A monotonic rise means variance harvesting -- bigger bets clearing a fixed target more often. Only an interior peak is consistent with a real edge.

Only the assumption is positive

2026-08-30T13:59:25.865444 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
The same signal priced against every order a human could actually place. The engine's assumed fill makes money; the limit, the market and the stop all lose it. The edge was in the gap between where the simulator booked the trade and where a trade could have happened.

The number that contained its own answer

2026-08-30T13:59:26.378543 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
total = the distance price had already travelled + what happened next. An exact identity, so it cannot manufacture a better number. 94.8% of the headline was the reference price.

5 · The register

Every hypothesis, in the state the record actually holds it, with its mechanism written before the number existed.

C2-MAE-PREDICTABLErange/ATR r = -0.131, partly algebraic, 1.7% of variance2 runs
axis
conditioning
search space
4
alpha
0.02
outcome
range/ATR r = -0.131, partly algebraic, 1.7% of variance
effect
—
interval
—

Statement

Adverse excursion is FORESEEABLE before entry: pre-trade volatility measures predict MAE/stop-distance, so days on which the stop is likely to be run can be identified in advance.

Mechanism written before the number existed

H-STOPWIDTH established MAE is LARGE -- 77% of trades exceed the stop, median 3x -- but not whether it is foreseeable. The mechanism here is the strongest available to us and is not a market-edge claim at all: MAE is essentially a realised-volatility quantity, and volatility CLUSTERS. That is one of the most robust facts in financial time series. If it holds intraday on this setup, a day's likely excursion is partly knowable from the volatility that preceded it.
Note what this would and would not buy. It is NOT a directional edge and cannot become one. At most it is a SIZING or STAND-ASIDE filter: skip, or size down, on days where the stop sits well inside the expected noise. That is worth having only if a strategy with a real entry ever exists -- so this is groundwork, honestly labelled.

Power plan

{
 "cluster_size": 2,
 "effect": 0.15,
 "intra_cluster_r": 0.9,
 "test": "correlation"
}

Resolved range/ATR r = -0.131, partly algebraic, 1.7% of variance documented

stated in prose, never computed into the register

stated in
docs/PROGRAMME-CONCLUSION.md
resolved
2026-08-30T16:52:11+00:00

No effect size, interval or floor is shown because none was ever computed into the register. The source document’s numbers are authoritative; this record is not.

Decision the part no run can produce

Precise but immaterial -- the corner of the 2x2 a significance test alone would have reported as a finding. Detectable, partly algebraic, and worth 1.7% of variance. Conditioning on adverse excursion is closed as an avenue.

source pinned at 4580cf2667b92398…

2026-08-03-m5-reproduction controls passed

recorded
2026-08-09T17:46:37+00:00
engine_sha
716c6ac4115191df3ccf7ee2bbeea8779b3f1a18-dirty
spend
$0.00
metricvalue
MES.mins to breakout-0.0501
MNQ.mins to breakout-0.0388
n3260
pooled.5d mean range/ATR0.0075
pooled.mins to breakout-0.044
pooled.overnight gap/ATR-0.0387
pooled.range/ATR-0.1312
Config
{
 "estimator": "occams.stats.spearman",
 "holds_constant": "data + setup logic (frozen script imported)",
 "tool": "scripts/backfill_m5.py"
}
Note
M5 REPRODUCTION — supersedes the `mins to breakout` values in 2026-08-01-four-predictors. NO ALPHA SPENT: same estimand, same data, same setup; only the rank function changed.

The original used argsort(argsort(x)), which gives tied values arbitrary DISTINCT ranks in row order. `mins to breakout` is the only integer-valued predictor in the set -- 94 distinct values across 3,260 observations, 97.1% tied -- and it is the only one that moved. MNQ -0.0304 -> -0.0388, a 28% change.

THE CONCLUSION IS UNCHANGED. range/ATR, the one predictor C2 reported as surviving, reproduces EXACTLY at -0.1312. Minutes to breakout was noise before and is noise now. But a recorded number was wrong, and this appends rather than edits it.

Detail: docs/M5-BACKFILL.md

2026-08-01-four-predictors controls passed

recorded
2026-08-03T17:57:24+00:00
engine_sha
7bb2312f4806b80b62e657d09489359aaef4ee48-dirty
spend
$0.00
metricvalue
MES.5d mean range/ATR0.0146
MES.mins to breakout-0.0522
MES.n1631
MES.overnight gap/ATR-0.0443
MES.range/ATR-0.1293
MNQ.5d mean range/ATR0.0053
MNQ.mins to breakout-0.0304
MNQ.n1629
MNQ.overnight gap/ATR-0.0335
MNQ.range/ATR-0.1196
n3260
pooled.5d mean range/ATR0.0076
pooled.mins to breakout-0.0459
pooled.overnight gap/ATR-0.0387
pooled.range/ATR-0.1312
Provenance the frozen script and its output, archived before the run
exp_mae_predictable.py4,297 Bed8924f26a91b6a4…
output.txt1,063 B8bf47c1d2f9d3607…
Config
{
 "instruments": [
  "MES",
  "MNQ"
 ],
 "outcome": "log(MAE / stop_distance)",
 "predictors": [
  "range/ATR",
  "overnight gap/ATR",
  "5d mean range/ATR",
  "mins to breakout"
 ],
 "search_size": 4,
 "span": "2019-2026",
 "test": "Spearman"
}
Note
OUTCOME: ONE of four predictors survives, and it is weaker than it looks.
range/ATR: r = -0.131 pooled, -0.129 MES, -0.120 MNQ. Consistent across instruments and above the detectable floor (r=0.067 at effective n=1,736). The other three are noise: overnight gap -0.039, 5-day mean range +0.008, minutes-to-breakout -0.046.
THE CAVEAT, and it matters more than the coefficient. The stop is defined as extreme +/- 0.2 x range HEIGHT, so stop distance scales with height by construction, while MAE scales roughly with daily volatility. MAE/stop is therefore ALGEBRAICALLY close to ATR/height -- the inverse of the predictor. Some unknown share of r=-0.131 is definitional rather than informative.
What argues against it being PURELY definitional is the magnitude: a clean inverse relationship would show rank correlation near -1, not -0.13. So MAE is dominated by day-specific noise rather than by the scaling. But that cuts both ways -- it also means the predictable component is small.
PRACTICAL VERDICT: r=-0.131 explains ~1.7% of variance. As a stand-aside filter that is negligible, and it would be filtering a strategy that has no obtainable entry anyway. NOT worth acting on.
Disentangling the definitional component would be a NEW hypothesis with its own alpha -- deliberately NOT done here, because the registration said one look and no fifth predictor, and the practical verdict does not change either way.

SCRIPT: experiments/C2-MAE-PREDICTABLE/2026-08-01-four-predictors/exp_mae_predictable.py
OUTPUT: experiments/C2-MAE-PREDICTABLE/2026-08-01-four-predictors/output.txt
E2-OBJECTIVEP(pass) 0.042 against a 0.55 gate2 runs
axis
objective
search space
5
alpha
0.01
outcome
P(pass) 0.042 against a 0.55 gate
effect
—
interval
—

Statement

Scored against the REAL objective -- P(pass) on the verified Zero tier -- the obtainable fade clears the G1 gate of 0.55, and P(pass) is non-monotonic in risk size, peaking BELOW the expectancy-maximising size.

Mechanism written before the number existed

The challenge is a path-dependent survival problem: trailing drawdown floor, profit target, minimum days. Expectancy ignores path. A lower-variance configuration should therefore reach the target more often before touching the floor, which would put the P(pass) maximum at a smaller size than the one maximising expectancy.

Resolved P(pass) 0.042 against a 0.55 gate documented

stated in prose, never computed into the register

stated in
docs/PROGRAMME-CONCLUSION.md
resolved
2026-08-30T16:52:06+00:00

No effect size, interval or floor is shown because none was ever computed into the register. The source document’s numbers are authoritative; this record is not.

Decision the part no run can produce

Fails twice over, and independently of the entry defect: even the artifact variant of the strategy reaches only 0.258 against a 0.55 gate. Confirms the objective was never close rather than narrowly missed.

source pinned at 4580cf2667b92398…

2026-08-01-zero-tier-INTERPRETATION controls passed

recorded
2026-08-03T11:23:54+00:00
engine_sha
167a4aa5a953a2f23b681674bef92f02826023a2-dirty
spend
$0.00
metricvalue
seeexperiments/E2-OBJECTIVE/2026-08-01-zero-tier.json
Config
{
 "supersedes_note_of": "2026-08-01-zero-tier"
}
Note
OUTCOME: hypothesis REJECTED on BOTH claims, and the second rejection is the useful one.
(1) The obtainable fade does not clear G1. P(pass) 0.042 at 90 days against a 0.55 gate, with P(breach) 0.440. The SEALED variant -- the one whose entry cannot be obtained -- reaches only 0.258, so even the artifact would have failed. Verdict v3's NO-GO stands twice over.
(2) MY REGISTERED PREDICTION WAS WRONG. P(pass) is MONOTONICALLY INCREASING in risk size, not non-monotonic: 0.000 at $75, 0.009 at $125, 0.042 at $175, 0.088 at $250, 0.174 at $350. There is no interior peak.
WHY, and this is worth keeping: with NEGATIVE expectancy the only route to a +$3,000 target is variance, so P(pass) rises with size -- and P(breach) rises faster (0.811 at $350). At $75 the strategy neither passes nor breaches; it simply drifts. The interior peak I predicted REQUIRES a positive edge, because only then does reducing variance preserve the target while moving the floor further away.
NEW LAB CAPABILITY implied: the SHAPE of P(pass) versus size is a diagnostic for whether an edge exists at all, independent of any expectancy estimate. Monotonic-increasing means variance harvesting and no edge; an interior peak means a real edge. Worth adding to the harness as a standard readout.

2026-08-01-zero-tier controls passed

recorded
2026-08-03T11:20:58+00:00
engine_sha
167a4aa5a953a2f23b681674bef92f02826023a2-dirty
spend
$0.00
metricvalue
obtainable_limit.60.median_days41
obtainable_limit.60.n1682
obtainable_limit.60.p_breach0.2913
obtainable_limit.60.p_pass0.0178
obtainable_limit.90.median_days77
obtainable_limit.90.n1652
obtainable_limit.90.p_breach0.4395
obtainable_limit.90.p_pass0.0424
risk_ladder.125.median_days71
risk_ladder.125.n1652
risk_ladder.125.p_breach0.2452
risk_ladder.125.p_pass0.0085
risk_ladder.175.median_days77
risk_ladder.175.n1652
risk_ladder.175.p_breach0.4395
risk_ladder.175.p_pass0.0424
risk_ladder.250.median_days52
risk_ladder.250.n1652
risk_ladder.250.p_breach0.7197
risk_ladder.250.p_pass0.0878
risk_ladder.350.median_days31
risk_ladder.350.n1652
risk_ladder.350.p_breach0.8111
risk_ladder.350.p_pass0.1743
risk_ladder.75.median_days
risk_ladder.75.n1652
risk_ladder.75.p_breach0.0
risk_ladder.75.p_pass0.0
sealed_unobtainable.60.median_days52
sealed_unobtainable.60.n1682
sealed_unobtainable.60.p_breach0.1849
sealed_unobtainable.60.p_pass0.0886
sealed_unobtainable.90.median_days70
sealed_unobtainable.90.n1652
sealed_unobtainable.90.p_breach0.2748
sealed_unobtainable.90.p_pass0.2579
Provenance the frozen script and its output, archived before the run
exp_objective.py5,376 Be8679daaf6fa8429…
output.txt1,580 Bc627fefc0c5a1556…
Config
{
 "account": 50000,
 "calendar_filter": true,
 "horizons": [
  60,
  90
 ],
 "instrument": "MES",
 "min_days": 1,
 "risk_ladder": [
  75,
  125,
  175,
  250,
  350
 ],
 "target": 3000,
 "tier": "verified Zero 50k",
 "trailing_dd": 2000
}
Note
PLACEHOLDER - replaced below once the outcome is written up.

SCRIPT: experiments/E2-OBJECTIVE/2026-08-01-zero-tier/exp_objective.py
OUTPUT: experiments/E2-OBJECTIVE/2026-08-01-zero-tier/output.txt
E3-COSTSMES spread 1.0 tick, MNQ 2.0 -- assumption optimistic1 run
axis
order flow
search space
1
alpha
0.0
outcome
MES spread 1.0 tick, MNQ 2.0 -- assumption optimistic
effect
—
interval
—

Statement

The sealed cost model -- $1.25 per side commission and 1 tick of slippage -- is accurate for our order sizes on MES and MNQ. Specifically: the effective spread is 1 tick, and a typical order of 1-6 contracts does not move price.

Mechanism written before the number existed

Commission was measured at 0.05-0.09R per trade, which is decisive at an edge of ~0.03R gross. Slippage enters the same arithmetic and is currently ASSUMED. If real slippage is 2 ticks rather than 1, every net figure in the register shifts by roughly the same magnitude as the edge itself. The trades tape carries aggressor side and size, so both the effective spread and our size relative to typical prints can be measured rather than assumed.

Resolved MES spread 1.0 tick, MNQ 2.0 -- assumption optimistic documented

stated in prose, never computed into the register

stated in
docs/PROGRAMME-CONCLUSION.md
resolved
2026-08-30T16:52:08+00:00

No effect size, interval or floor is shown because none was ever computed into the register. The source document’s numbers are authoritative; this record is not.

Decision the part no run can produce

Cost assumption tightened, not overturned. The sealed one-tick assumption is correct for MES and optimistic for MNQ, where the effective spread is double it. Changes no verdict -- every family failed by far more than the correction -- but it removes an optimism the numbers had been carrying silently.

source pinned at 4580cf2667b92398…

2026-08-01-tape-measured controls passed

recorded
2026-08-03T11:37:20+00:00
engine_sha
3f9bfd3af6196b166061da441e767d01190451aa-dirty
spend
$0.00
metricvalue
MES.print_size_median1
MES.print_size_p906
MES.share_prints_ge_11.0
MES.share_prints_ge_30.2645
MES.share_prints_ge_60.1079
MES.spread_ticks_mean0.642
MES.spread_ticks_median1.0
MES.walk_ticks_median_10.0
MES.walk_ticks_median_30.0
MES.walk_ticks_median_60.0
MES.walk_ticks_p90_10.0
MES.walk_ticks_p90_31.0
MES.walk_ticks_p90_61.0
MNQ.print_size_median1
MNQ.print_size_p904
MNQ.share_prints_ge_11.0
MNQ.share_prints_ge_30.1798
MNQ.share_prints_ge_60.0358
MNQ.spread_ticks_mean1.597
MNQ.spread_ticks_median2.0
MNQ.walk_ticks_median_10.0
MNQ.walk_ticks_median_31.0
MNQ.walk_ticks_median_61.0
MNQ.walk_ticks_p90_10.0
MNQ.walk_ticks_p90_33.0
MNQ.walk_ticks_p90_64.0
Provenance the frozen script and its output, archived before the run
exp_costs.py4,299 B9de10bb389d84fd5…
output.txt1,634 Ba41f517c921a032f…
Config
{
 "assumed_model": "1 tick slippage, $1.25/side",
 "instruments": [
  "MES",
  "MNQ"
 ],
 "sessions_per_inst": 60,
 "source": "R2.1 trades tape, 09:45-10:15 ET"
}
Note
OUTCOME: SPLIT. The assumption holds for MES and is OPTIMISTIC for MNQ.
MES: effective spread 1.00 ticks median, and even a 6-lot fills with 0 ticks of travel (p90 1). 6 contracts sits at the p90 of print size. The sealed 1-tick assumption is CORRECT for MES.
MNQ: effective spread 2.00 ticks median (mean 1.60) -- DOUBLE the assumption. A 3-lot travels 1 tick median / 3 at p90; a 6-lot 1 / 4. Only 3.6% of MNQ prints are 6 contracts or larger, so a 6-lot is a large order there, unlike MES where it is ordinary.
MAGNITUDE, stated so it is not over-read: one extra tick per side on MNQ is $1.00 per contract, roughly 0.006R at 1 contract and 0.017R at 3. That is real and it makes every MNQ figure in the register slightly optimistic, but it is an order of magnitude below the gap that matters (0.03R gross against 0.25R required). It changes no conclusion; it tightens the error bar.
ACTION: raise slippage_ticks to 2 for MNQ in the instrument costs, leave MES at 1, and re-state that MNQ results predating this were measured on an optimistic model.

SCRIPT: experiments/E3-COSTS/2026-08-01-tape-measured/exp_costs.py
OUTPUT: experiments/E3-COSTS/2026-08-01-tape-measured/output.txt
H-RECLAIMdetectable2 runs
axis
price
search space
1
alpha
0.01
outcome
detectable
effect
-0.2393
interval
[-0.2910, -0.1843]

Statement

Requiring price to trade through the stop, back through it, and then reclaim the original entry before entering will improve the fade's per-trade expectancy.

Mechanism written before the number existed

User observation from chart inspection: on several 2026-07 setups the eventual direction was correct but the stop was hit first. Stated BEFORE the test was run. The proposed mechanism is that a reclaim confirms the fade thesis and filters days that simply keep going.

Resolved detectable measured

read from a scored result in an archived run

from run
2026-08-30-audited primary
effect
-0.23931
95% CI
[-0.2910, -0.1843]
floor
0.10378428725965629
resolved
2026-08-30T18:19:13+00:00
supersedes
resolutions/H-RECLAIM/documented.json
Decision the part no run can produce

Entry filter rejected, now on measured grounds. The reclaim-filtered strategy returns -0.239R per trade with a 95% interval of [-0.291, -0.184] -- detectably negative at 2.3 times its own floor. The sealed arm it was meant to improve on is -0.036R with an interval crossing zero: NULL, indistinguishable from nothing. So the filter does not fail to help, it actively converts an unprofitable-but-indistinguishable strategy into a reliably losing one. Confirmation costs the move being confirmed.

source pinned at ee712d854cbf82d2…

2026-08-30-audited controls passed

recorded
2026-08-30T18:19:00+00:00
engine_sha
b9a735e8289e5f9f85e747eea2abe1cf2c6309ac-dirty
spend
$0.00
metricvalue
comparator.ci[-0.08349605615389351, 0.013471034229541627]
comparator.crosses_zeroTrue
comparator.d-0.035917566773626265
comparator.estimate-0.035917566773626265
comparator.floor0.08996991758658018
comparator.floor_multiples-0.3992175133322973
comparator.n3161
comparator.n_eff1663
comparator.nameH-RECLAIM COMPARATOR - sealed (unfiltered) expectancy, per trade
comparator.notethe arm the filter was supposed to improve on.
comparator.unitR
comparator.verdictnull
n_paired1964
n_reclaim1964
n_sealed3161
primary.ci[-0.29097243965739494, -0.18426419181005807]
primary.crosses_zeroFalse
primary.d-0.2393066627873168
primary.estimate-0.2393066627873168
primary.floor0.10378428725965629
primary.floor_multiples-2.3058082211288804
primary.n1964
primary.n_eff1033
primary.nameH-RECLAIM PRIMARY - reclaim-filtered expectancy, per trade
primary.notethe filtered strategy's own expectancy. Compare with the sealed arm below: the intervals do not overlap.
primary.unitR
primary.verdictdetectable
secondary_paired.ci[0.17358297226237054, 0.31150096492367946]
secondary_paired.crosses_zeroFalse
secondary_paired.d0.24265522257782926
secondary_paired.estimate0.24265522257782926
secondary_paired.floor0.1334909642818425
secondary_paired.floor_multiples1.817765148999192
secondary_paired.n1964
secondary_paired.n_eff1033
secondary_paired.nameH-RECLAIM SECONDARY - paired reclaim minus sealed, shared days only
secondary_paired.notePOSITIVE and it does NOT rescue the filter. This conditions on days a reclaim occurred, which are days price first went through the stop -- exactly where the sealed entry does worst. It answers 'given a reclaim happened, was the later entry better', not 'is the filter worth applying'.
secondary_paired.unitR
secondary_paired.verdictdetectable
Provenance the frozen script and its output, archived before the run
calcs.json1,141 B1eea15ed12c7cb9b…
f5_reclaim.py854 B90d54d48bdbda147…
output.txt4,038 Bd542a9ee98a7f943…
Config
{
 "cluster_size": 2,
 "engine": "M1-M7 audited",
 "intra_cluster_r": 0.9,
 "reproduces": "2026-08-01-full-history (10/10)",
 "seed": 20260830
}
Note
F5 recompute. The archived run recorded bare means with no interval, no floor and no effective n, so it could not be resolved as a measured result. This reproduces it 10/10 and then states the estimand: the FILTERED strategy expectancy against the sealed arm. Strategy level, not trade level -- the filter changes which days are traded.

SCRIPT: experiments/H-RECLAIM/2026-08-30-audited/f5_reclaim.py
OUTPUT: experiments/H-RECLAIM/2026-08-30-audited/output.txt
CALCS: experiments/H-RECLAIM/2026-08-30-audited/calcs.json

2026-08-01-full-history CONTROLS NOT PASSED

recorded
2026-08-03T06:51:13+00:00
engine_sha
e22df70f08b135289de5e57dfcd9d64b671a1c05-dirty
spend
$0.00
metricvalue
MES.setups1706
MES.v0_r_per_trade-0.046
MES.v0_trades1623
MES.v5_r_per_trade-0.262
MES.v5_trades1018
MNQ.setups1622
MNQ.v0_r_per_trade-0.025
MNQ.v0_trades1538
MNQ.v5_r_per_trade-0.215
MNQ.v5_trades946
years_v5_worse_than_v07 of 8 on each instrument
Config
{
 "calendar_filter": false,
 "folds": "pooled (NOT the sealed OOS/lockbox split)",
 "grid_size": 1,
 "instruments": [
  "MES",
  "MNQ"
 ],
 "k_stop": 0.2,
 "risk_usd": 175.0,
 "rule": "reclaim-before-entry",
 "years": "2019-2026"
}
Note
OUTCOME: hypothesis REJECTED. The reclaim is anti-confirmation, not confirmation: you enter at the same price with the same stop but only after the market has demonstrated it can punch through that stop, so the filter selects for days where the stop is weakest. It removes ~40% of trades and removes the wrong 40%. CAVEAT, recorded rather than buried: controls_passed=False because this run used a re-implementation, pooled years and no calendar filter -- it is NOT the sealed engine. The V0-vs-V5 COMPARISON is valid because both share those simplifications; the ABSOLUTE levels are not, and reconciling them is task X1.
H-SECONDPUSHgross +0.011 -> net -0.082 R; costs ate the signal1 run
axis
price
search space
1
alpha
0.02
outcome
gross +0.011 -> net -0.082 R; costs ate the signal
effect
—
interval
—

Statement

Use the failed-breakout as a SIGNAL only, take no fade, and instead enter WITH the original breakout direction when price returns through extreme +/- 0.2h (the level where the fade would have been stopped). Stop at the range boundary; target one range height beyond entry.

Mechanism written before the number existed

H-STOPWIDTH established that 77% of fade setups see price travel through the stop, median 3x the stop distance. If the fade is systematically on the wrong side, the second push through the extreme is the tradeable event: the pullback has shaken out the faders and the original breakout resumes. Distinct from the engine's 'follow' null, which enters at the BOUNDARY rather than at the stop level.

Resolved gross +0.011 -> net -0.082 R; costs ate the signal documented

stated in prose, never computed into the register

stated in
docs/PROGRAMME-CONCLUSION.md
resolved
2026-08-30T16:52:01+00:00

No effect size, interval or floor is shown because none was ever computed into the register. The source document’s numbers are authoritative; this record is not.

Decision the part no run can produce

Closed. The only family here with a positive gross signal, and the costs consumed all of it and more. Kept as the clearest single demonstration that a gross number is worthless on these instruments: the strategy was right about the market and wrong about the arithmetic.

source pinned at 4580cf2667b92398…

2026-08-01-full-history CONTROLS NOT PASSED

recorded
2026-08-03T07:38:10+00:00
engine_sha
93029884fbc39c4bf5cb62aa0f59fd350e0e6282-dirty
spend
$0.00
metricvalue
MES.cost_r-0.093
MES.gross_r0.011
MES.median_contracts6
MES.net_r-0.082
MES.setups1734
MES.trigger_rate0.734
MES.triggered1272
MES.win0.281
MNQ.cost_r-0.053
MNQ.gross_r0.039
MNQ.median_contracts3
MNQ.net_r-0.015
MNQ.setups1737
MNQ.trigger_rate0.699
MNQ.triggered1214
MNQ.win0.285
fade_for_comparison.MES.cost_r-0.088
fade_for_comparison.MES.gross_r0.042
fade_for_comparison.MES.net_r-0.046
fade_for_comparison.MNQ.cost_r-0.052
fade_for_comparison.MNQ.gross_r0.026
fade_for_comparison.MNQ.net_r-0.026
Config
{
 "calendar_filter": false,
 "entry": "stop order at extreme +/- 0.2h, WITH the breakout",
 "folds": "pooled",
 "instruments": [
  "MES",
  "MNQ"
 ],
 "risk_usd": 175.0,
 "span": "2019-05-06..2026-07-03",
 "stop": "range boundary",
 "target": "one range height beyond entry"
}
Note
OUTCOME: REJECTED as a standalone strategy -- net -0.082R (MES) / -0.015R (MNQ), negative in 6 of 8 years on MES. BUT the decomposition is the real finding, and it is bigger than this hypothesis. BOTH directions are GROSS POSITIVE and both are made negative by COMMISSION: fade +0.042/-0.088, second push +0.011/-0.093 on MES. Cost drag is 0.05-0.09R PER TRADE because true-risk sizing buys many contracts when the stop is tight -- MES median 6 contracts x $2.50 round turn = $15 against a $175 risk budget = 0.086R of pure commission. This reconciles H-STOPWIDTH's ladder, where net improved monotonically toward zero (x1 -0.046 -> x8 -0.013 on MES) precisely as widening stops shrank position size and therefore commission. THE STRATEGIC IMPLICATION: gross signal on this setup family is roughly +0.02 to +0.04R. The challenge needs +0.25R. Even at ZERO cost the family is an order of magnitude short, so no further price-based variant of it can close the gap. This is the strongest available argument for R2 (a new information axis) over more search on price.
H-STOPWIDTHladder flat at every rung, 77.5% exceed stop1 run
axis
price
search space
5
alpha
0.02
outcome
ladder flat at every rung, 77.5% exceed stop
effect
—
interval
—

Statement

The fade stops out too early: widening the stop, or removing it, will improve per-trade expectancy because the direction is usually right over a longer horizon.

Mechanism written before the number existed

User observation from chart inspection, corroborated on 7 live setups where adverse excursion ran 5-12x the stop distance. If the stop sits inside ordinary intraday noise, it should be converting winners into losers.

Resolved ladder flat at every rung, 77.5% exceed stop documented

stated in prose, never computed into the register

stated in
docs/PROGRAMME-CONCLUSION.md
resolved
2026-08-30T16:52:00+00:00

No effect size, interval or floor is shown because none was ever computed into the register. The source document’s numbers are authoritative; this record is not.

Decision the part no run can produce

Stop width closed as an avenue. Expectancy is unchanged across a four-rung ladder from tight to wide, while 77.5% of trades exceed their stop at some point at a median of three times the stop distance. The adverse excursion is enormous and none of the answer lives in tuning it.

source pinned at 4580cf2667b92398…

2026-08-01-full-history CONTROLS NOT PASSED

recorded
2026-08-03T07:20:15+00:00
engine_sha
d5837b661e6fee02a54ca7fc32f692a45cf09757-dirty
spend
$0.00
metricvalue
MES.ladder.nostop[1622, 0.647, -0.084]
MES.ladder.x1[1622, 0.309, -0.046]
MES.ladder.x2[1523, 0.429, -0.05]
MES.ladder.x4[1247, 0.543, -0.028]
MES.ladder.x8[748, 0.642, -0.013]
MES.mae_over_stop_median3.0
MES.mae_over_stop_p9010.9
MES.setups1650
MES.share_mae_exceeds_stop0.775
MES.worst_trade_nostop_usd-4444
MNQ.ladder.nostop[1537, 0.591, -0.158]
MNQ.ladder.x1[1537, 0.274, -0.026]
MNQ.ladder.x2[1268, 0.38, -0.059]
MNQ.ladder.x4[665, 0.487, -0.061]
MNQ.ladder.x8[158, 0.595, -0.042]
MNQ.mae_over_stop_median2.8
MNQ.mae_over_stop_p908.8
MNQ.setups1645
MNQ.share_mae_exceeds_stop0.763
MNQ.worst_trade_nostop_usd-6349
legend[trades, win_rate, R_per_trade]
Config
{
 "calendar_filter": false,
 "folds": "pooled (NOT the sealed OOS/lockbox split)",
 "instruments": [
  "MES",
  "MNQ"
 ],
 "risk_usd": 175.0,
 "sizing": "fixed risk, so a wider stop buys fewer contracts",
 "span": "2019-05-06..2026-07-03"
}
Note
OUTCOME: the OBSERVATION is confirmed, the HYPOTHESIS is rejected. 77% of trades do have an adverse excursion exceeding the stop, median 3x -- the stop genuinely sits inside ordinary noise. But widening it does not help: win rate climbs steeply (31%->64% on MES) while expectancy stays flat and negative. That is the signature of NO EDGE -- moving along an iso-expectancy curve, trading win rate against loss size. Removing the stop is worst on both instruments and produces single-trade losses of -$4,444 and -$6,349 against a $175 risk budget (25x and 36x), which disqualifies it structurally regardless of expectancy. CAVEAT: controls_passed=False. This V0 baseline reads -0.046R / -0.026R where the sealed verdict read +0.1R. The ladder COMPARISON is valid (all rungs share the simplifications); the ABSOLUTE levels are not. Two implementations disagreeing by 0.15R across 3,295 setups makes X1 MORE urgent, not less.
H4-ORDERFLOWsuperseded0 runs
axis
order flow
search space
1
alpha
0.03
outcome
superseded
effect
—
interval
—

Statement

Signed aggressor imbalance during the breakout leg predicts the outcome of the failed-breakout setup. Specifically: a breakout pushed by weak aggression (low or opposing signed imbalance) fades more reliably than one pushed by strong aggression.

Mechanism written before the number existed

The fade's own story is an order-flow story -- 'the breakout had no participation behind it' -- yet every test to date has used price alone, and price alone is public information that H-SECONDPUSH showed carries only ~+0.03R gross against a +0.25R requirement. Aggressor side is information not visible on a chart. If the fade thesis is mechanically true, the trades tape at the breakout should separate the setups that revert from the ones that resume.

Resolved superseded superseded

replaced before it was ever run

superseded by
H4-ORDERFLOW-v2
resolved
2026-08-30T16:50:06+00:00
Decision the part no run can produce

Never run. Superseded before any data existed by H4-ORDERFLOW-v2, which changed only the SAMPLE and for budget reasons -- the analysis plan is identical. Costs no alpha beyond the successor's, because it was replaced rather than tested and abandoned.

source pinned at cc633b4319d480fd…

H4-ORDERFLOW-v2superseded0 runs
axis
order flow
search space
1
alpha
0.03
outcome
superseded
effect
—
interval
—
supersedes
H4-ORDERFLOW

Statement

Signed aggressor imbalance during the breakout leg predicts the outcome of the failed-breakout setup: a breakout pushed by weak aggression fades more reliably than one pushed by strong aggression.

Mechanism written before the number existed

Unchanged from H4-ORDERFLOW. The fade's own story is an order-flow story that has only ever been tested through price, and H-SECONDPUSH showed price carries ~+0.03R gross against a +0.25R requirement.

Resolved superseded superseded

replaced before it was ever run

superseded by
H5-FLOW-PREDICTS
resolved
2026-08-30T16:50:07+00:00
Decision the part no run can produce

Never run. Superseded before any data existed by H5-FLOW-PREDICTS, which restated the same order-flow axis with a different outcome variable (signed return to the session close rather than the fade outcome). The order-flow question survives in H5; this framing of it does not.

source pinned at 5d5113f74dd29374…

H5-FLOW-PREDICTSr = -0.072, UNDERPOWERED (n_eff 474, floor 0.131)1 run
axis
order flow
search space
1
alpha
0.03
outcome
r = -0.072, UNDERPOWERED (n_eff 474, floor 0.131)
effect
—
interval
—
supersedes
H4-ORDERFLOW-v2

Statement

Signed aggressor imbalance during the opening-range breakout predicts the DIRECTION of the subsequent move, measured as the signed return from the failure close to the 16:00 ET close, normalised by the range height.

Mechanism written before the number existed

SUPERSEDES H4, which asked whether imbalance improves the FADE. The fade is dead -- its entry cannot be obtained (#3a) -- so that question is moot and its answer would be unusable either way.
The fresh question is prior to any strategy: does the tape carry directional information at the breakout AT ALL? A breakout pushed by heavy one-sided aggression is a different event from one that drifts through on thin flow, and that distinction is invisible in OHLC. If no such information exists, no strategy built on this setup can work and the programme should say so. If it does, the strategy is designed AFTERWARDS, around an entry that is obtainable by construction -- which is the inversion of how the fade was built.

Power plan

{
 "alpha": 0.05,
 "effect": 0.1,
 "power": 0.8,
 "test": "correlation"
}

Resolved r = -0.072, UNDERPOWERED (n_eff 474, floor 0.131) documented

stated in prose, never computed into the register

stated in
docs/PROGRAMME-CONCLUSION.md
resolved
2026-08-30T16:52:09+00:00

No effect size, interval or floor is shown because none was ever computed into the register. The source document’s numbers are authoritative; this record is not.

Decision the part no run can produce

Inconclusive, NOT null, and the distinction is the point: the sample cannot see an effect the size of the one measured, so this records a question probed and not answered. The same data shortage later left X3 unresolved on the same axis. Ambiguous nulls cost alpha and buy nothing, which is why the power gate now refuses them in advance.

source pinned at 4580cf2667b92398…

2026-08-01-first-look controls passed

recorded
2026-08-03T15:32:58+00:00
engine_sha
48304bb211e1289d97974ea3cb2cb6b2db481a86-dirty
spend
$0.00
metricvalue
MES.fwd_mean-0.0333
MES.imbalance_mean-0.0774
MES.n453
MES.spearman-0.0601
MNQ.fwd_mean0.0254
MNQ.imbalance_mean-0.0595
MNQ.n448
MNQ.spearman-0.0885
n901
pooled.spearman-0.0723
Provenance the frozen script and its output, archived before the run
exp_flow_predicts.py5,111 Bba3dc17bf4370657…
output.txt531 Bb943eada0839ee0c…
Config
{
 "outcome": "signed forward return to 16:00 ET, in range-heights",
 "predictor": "signed aggressor imbalance, breakout bar through failure close",
 "sessions": "525 x2 purchased, 453/448 overlapping",
 "test": "Spearman, pooled and per instrument",
 "window": "09:45-10:15 ET"
}
Note
OUTCOME: NOT ESTABLISHED, but the first result this session that was not killed outright. Handle carefully.
Pooled Spearman r = -0.0723 (MES -0.060, MNQ -0.089) -- consistent in SIGN and rough magnitude across both instruments. The pre-specified tercile secondary is a clean monotonic gradient: low imbalance +0.216 range-heights forward, mid -0.030, high -0.198. A spread of 0.41 range-heights between terciles is economically meaningful IF it is real.
DIRECTION, and it is counterintuitive: NEGATIVE r means heavy aggression WITH the breakout predicts REVERSAL, while breakouts that drift through on weak or opposing flow tend to CONTINUE. That is the classic exhaustion shape -- aggressive buyers lifting into a breakout are late, and their flow is the fuel for the move back.
WHY IT IS NOT ESTABLISHED, stated plainly:
(a) UNDERPOWERED FOR WHAT IT FOUND. The sample was powered to detect r=0.10 (n>=783) and passed that gate at n=901. The observed r=0.072 is SMALLER than the effect we powered for; detecting it would need n~1,500.
(b) THE POOLED n IS NOT 901. MES and MNQ are ~90% correlated intraday, so 453+448 setups are not 901 independent observations. Effective n is nearer 453, where the detectable effect is r=0.13 -- almost double what we saw. My power plan treated pooling as free and it is not. That is a defect in the plan, not in the result.
WHAT WOULD SETTLE IT: ~780 INDEPENDENT setups, meaning ~3 years per instrument rather than the ~2 we bought, and the correlation between instruments handled explicitly rather than by pooling. NOT bought on the strength of this -- a suggestive gradient is exactly what a second look is most likely to flatter.

SCRIPT: experiments/H5-FLOW-PREDICTS/2026-08-01-first-look/exp_flow_predicts.py
OUTPUT: experiments/H5-FLOW-PREDICTS/2026-08-01-first-look/output.txt
ICT-P1-FVG-0.00036 ATR, d -0.0015; no forward information1 run
axis
price (1-minute bars, owned)
search space
3
alpha
0.05
outcome
-0.00036 ATR, d -0.0015; no forward information
effect
—
interval
—

Statement

A fair value gap (3-bar imbalance) carries directional forward information. Specifically: after price returns to the gap's consequent encroachment (50%), the signed forward return in the direction of the displacement that created the gap has mean > 0 over a 30-minute horizon, in ATR units.

Mechanism written before the number existed

ICT doctrine holds that a displacement leg leaves an imbalance -- one side of the book was not traded -- and that price returns to 'rebalance' it before continuing. Stated as a market claim rather than a metaphor: a fast directional move leaves resting interest unfilled; that interest is still there when price returns, so the return is met by the same direction of pressure that caused the move. If true, a touch of the gap midpoint marks a point where the original direction resumes.

This is TESTED AS A PRIMITIVE, not as a strategy, and that is the whole point of registering it this way. All four strategies in the source playbook -- 2022 Model, Silver Bullet, Judas Swing, Turtle Soup -- enter with a limit at the FVG or its CE. The entry is their only shared component. If the gap carries no forward information, all four die on this one register entry rather than four.

The outcome is a strategy-free signed forward return: no entry assumption, no stop, no target. This is deliberate. The fade's defect entered through a P&L outcome that smuggled in an unobtainable fill, and a signed return from a touched price to a later price cannot carry one.

PRIOR, stated before the number: I expect this to be at or near zero. FVG detectors ship with every charting platform, which makes this the most crowded thing on our map. The one argument the other way is that crowding at a LIMIT level can be self-fulfilling -- resting orders are the liquidity -- where crowding at a breakout is self-defeating.

Power plan

{
 "alpha": 0.05,
 "cluster_size": 2,
 "effect": 0.07,
 "intra_cluster_r": 0.9,
 "power": 0.8,
 "test": "mean_shift"
}

Resolved -0.00036 ATR, d -0.0015; no forward information documented

stated in prose, never computed into the register

stated in
docs/PROGRAMME-CONCLUSION.md
resolved
2026-08-30T16:52:13+00:00

No effect size, interval or floor is shown because none was ever computed into the register. The source document’s numbers are authoritative; this record is not.

Decision the part no run can produce

No forward information at all, in units where the test was built to detect 0.07. This is the entry shared by four named strategies in the methodology, which is how two register entries closed eight named strategies and their variants.

source pinned at 4580cf2667b92398…

r1 controls passed

recorded
2026-08-03T21:28:58+00:00
engine_sha
01391280e3d5e1acb5470c33c7496463a478ca6b-dirty
spend
$0.00
metricvalue
MES.ce_fill_rate0.9207
MES.days1741
MES.fwd_mean_touched-0.00225
MES.gap_atr_median0.01861
MES.mirror_fill_rate0.9294
MNQ.ce_fill_rate0.8937
MNQ.days1741
MNQ.fwd_mean_touched0.00159
MNQ.gap_atr_median0.02134
MNQ.mirror_fill_rate0.9064
n3159
pooled.ci_high0.00903
pooled.ci_low-0.0104
pooled.cohen_d-0.0015
pooled.days_with_fvg3482
pooled.eod_mean0.00107
pooled.fwd_mean-0.00036
pooled.sd0.24361
pooled.unconditional_ci[-0.00956, 0.00859]
pooled.unconditional_mean-0.00033
Provenance the frozen script and its output, archived before the run
exp_ict_fvg.py6,534 B4e954d567f8d64ad…
output.txt945 B5c4cd3b907203973…
Config
{
 "days": 1849,
 "entry_ref": "CE (gap midpoint)",
 "fvg": "3-bar imbalance, first of RTH session",
 "horizon_bars": 30,
 "instruments": [
  "MES",
  "MNQ"
 ],
 "seed": 20260803,
 "unit": "daily ATR"
}
Note
ICT-P1. Strategy-free: signed forward return from a DERIVED CE touch. Primary conditional on touch; unconditional (no touch = 0) reported as pre-specified secondary. Mirror-level control for the fill rate.

SCRIPT: experiments/ICT-P1-FVG/r1/exp_ict_fvg.py
OUTPUT: experiments/ICT-P1-FVG/r1/output.txt
ICT-P2-CONTROLdrift -0.008, CI crosses zero; the corrected reading1 run
axis
price (1-minute bars, owned) -- control on ICT-P2-SWEEP
search space
1
alpha
0.0
outcome
drift -0.008, CI crosses zero; the corrected reading
effect
—
interval
—
supersedes
ICT-P2-SWEEP

Statement

P2's -0.159 ATR reversal reading is dominated by the reference price, not by forward drift. Decomposing total = penetration + drift, the penetration term accounts for the majority of it, and the drift term is small.

Mechanism written before the number existed

Stated from construction, before the decomposition is run. On an up-sweep bar price is ALREADY above the swept level when that bar closes. Measuring forward return from the level therefore includes the penetration distance, which is positive by construction. A market that froze the instant it swept would still print a positive continuation reading. The identity total = (C[i]-level) + (C[end]-C[i]) is exact, so this control cannot manufacture a better number -- the parts must sum to P2's r1 result. That is what makes it a control and not a second look at the hypothesis.

Registered because P2 r1 returned Cohen d = -0.43 where R1 established the gross signal on these instruments at +0.02 to +0.04R. An effect six times the declared detectable floor, appearing suddenly on the family we have already killed twice, is far more likely to be an artifact of measurement than a discovery. That is the #3a lesson: the fade also passed a sealed verdict before anyone asked what price it was using.

Resolved drift -0.008, CI crosses zero; the corrected reading documented

stated in prose, never computed into the register

stated in
docs/PROGRAMME-CONCLUSION.md
resolved
2026-08-30T16:52:17+00:00

No effect size, interval or floor is shown because none was ever computed into the register. The source document’s numbers are authoritative; this record is not.

Decision the part no run can produce

The corrected reading, registered at zero alpha because it is a control rather than a new question. What remains once the reference-price term is removed is indistinguishable from nothing: the sweep carries no forward information.

source pinned at 4580cf2667b92398…

r1 controls passed

recorded
2026-08-03T21:32:03+00:00
engine_sha
01391280e3d5e1acb5470c33c7496463a478ca6b-dirty
spend
$0.00
metricvalue
drift_ci[-0.01953, 0.00256]
drift_cohen_d-0.0341
drift_mean-0.00829
identity_holdsTrue
n3156
penetration_mean-0.15023
penetration_share0.9477
total_mean-0.15853
Provenance the frozen script and its output, archived before the run
exp_ict_sweep_decomp.py4,954 Ba805876cdf7955ba…
output.txt794 B83938c9f933460e2…
Config
{
 "horizon_bars": 30,
 "identity": "total = penetration + drift",
 "instruments": [
  "MES",
  "MNQ"
 ],
 "seed": 20260803
}
Note
Control on ICT-P2-SWEEP r1. Must reproduce its total exactly.

SCRIPT: experiments/ICT-P2-CONTROL/r1/exp_ict_sweep_decomp.py
OUTPUT: experiments/ICT-P2-CONTROL/r1/output.txt
ICT-P2-SWEEP-0.159 ATR, d -0.43 -- ARTIFACT, 94.8% reference price1 run
axis
price (1-minute bars, owned)
search space
3
alpha
0.05
outcome
-0.159 ATR, d -0.43 -- ARTIFACT, 94.8% reference price
effect
—
interval
—

Statement

A sweep of the prior session's RTH extreme carries directional forward information in the REVERSAL direction. After the first sweep of the prior day's high (low), the signed forward return measured short (long) has mean > 0 over 30 minutes, in ATR units.

Mechanism written before the number existed

The claim is that obvious highs and lows hold resting stop orders, that price is drawn to them because filling large orders requires the liquidity those stops provide, and that once taken the move reverses because the fuel is spent. Unlike the FVG claim this is a story about participants rather than a statistical property, which is why it scores 'fair' and not 'strong' on mechanism.

Tested as a primitive: 2022 Model, Judas Swing and Turtle Soup differ ONLY in which level is swept. Prior-day extreme is the cleanest and most objective of those levels, so it is the one measured. One measurement speaks to three named strategies.

This family is our failed-breakout fade under another name. Three pieces of counter-evidence are already on the record and are stated here BEFORE the run so they cannot be recalled selectively afterwards: R1 measured the gross signal at +0.02 to +0.04R against a +0.25R requirement; H-STOPWIDTH found expectancy flat across a four-rung stop ladder, so a tighter stop has already been probed on this family and did nothing; and Z0 priced four placeable orders on it, all negative. What is genuinely new is only the LEVEL (prior-day rather than opening-range) and the strategy-free outcome.

PRIOR: I expect zero or a small NEGATIVE, i.e. continuation rather than reversal, because a sweep of an obvious level is also a breakout, and the breakout reading is the more crowded one.

Power plan

{
 "alpha": 0.05,
 "cluster_size": 2,
 "effect": 0.07,
 "intra_cluster_r": 0.9,
 "power": 0.8,
 "test": "mean_shift"
}

Resolved -0.159 ATR, d -0.43 -- ARTIFACT, 94.8% reference price documented

stated in prose, never computed into the register

stated in
docs/PROGRAMME-CONCLUSION.md
resolved
2026-08-30T16:52:15+00:00

No effect size, interval or floor is shown because none was ever computed into the register. The source document’s numbers are authoritative; this record is not.

Decision the part no run can produce

Withdrawn as a finding. Six times the detectable floor, significant on both instruments, both sides and both horizons -- and 94.8% of it was the distance price had already travelled before the signal was observable. A market that froze solid at the trigger would have printed most of it. Superseded by ICT-P2-CONTROL; produced the calibration gate, which now catches this class on day one.

source pinned at 4580cf2667b92398…

r1 controls passed

recorded
2026-08-03T21:29:49+00:00
engine_sha
01391280e3d5e1acb5470c33c7496463a478ca6b-dirty
spend
$0.00
metricvalue
MES.days1740
MES.fwd_mean-0.16414
MES.sweep_rate0.9115
MNQ.days1740
MNQ.fwd_mean-0.15285
MNQ.sweep_rate0.9023
n3156
pooled.ci_high-0.14229
pooled.ci_low-0.175
pooled.cohen_d-0.4318
pooled.eod_mean_reversal-0.15696
pooled.fwd_mean_reversal-0.15853
pooled.sd0.36711
Provenance the frozen script and its output, archived before the run
exp_ict_sweep.py5,254 B337e72ee59a93db5…
output.txt740 B9042da2124d877ad…
Config
{
 "horizon_bars": 30,
 "instruments": [
  "MES",
  "MNQ"
 ],
 "level": "prior included RTH session high/low",
 "seed": 20260803,
 "sign": "reversal",
 "unit": "daily ATR"
}
Note
ICT-P2. Strategy-free reversal-signed forward return from the swept LEVEL. Ambiguous bars breaking both extremes in one minute are dropped rather than guessed -- intrabar order is unknowable.

SCRIPT: experiments/ICT-P2-SWEEP/r1/exp_ict_sweep.py
OUTPUT: experiments/ICT-P2-SWEEP/r1/output.txt
Q1-PAYOUT-PATHP(pass) 0.042 vs P(first payout) 0.012 -- 73% of qualifiers never paid1 run
axis
objective (capability, not a market question)
search space
1
alpha
0.0
outcome
P(pass) 0.042 vs P(first payout) 0.012 -- 73% of qualifiers never paid
effect
—
interval
—
supersedes
E2-OBJECTIVE

Statement

The engine can express P(first payout), not only P(pass). Demonstrated on the fade, whose entry is already known unobtainable, so that no decision rides on the number.

Mechanism written before the number existed

NOT A HYPOTHESIS TEST, and registered at ZERO ALPHA for that reason. The fade is dead on obtainability (#3a); no payout figure can revive it, so nothing rides on the result and no alpha is spent. What is being established is that the question can be asked at all.

The gap it closes: consistency rules differ BY STAGE and the engine held only the first. The tier we sealed has no evaluation consistency requirement; the qualified tier applies a largest-day rule to PAYOUT eligibility. A strategy can therefore qualify comfortably and never be paid, and ChallengeState alone could not express that. Every P(pass) this programme has quoted -- E2's 0.042, the 0.55 gate, the whole feasibility frontier -- measured getting through the door.

THE PAYOUT PARAMETERS ARE ILLUSTRATIVE AND UNVERIFIED. We verified the evaluation geometry and never verified the funded-stage terms. Every number this produces is conditional on parameters that a user must confirm before any of it is cited. Stated here rather than discovered in a footnote later.

Resolved P(pass) 0.042 vs P(first payout) 0.012 -- 73% of qualifiers never paid documented

stated in prose, never computed into the register

stated in
docs/Q1-PAYOUT.md
resolved
2026-08-30T16:52:18+00:00

No effect size, interval or floor is shown because none was ever computed into the register. The source document’s numbers are authoritative; this record is not.

Decision the part no run can produce

The objective was incomplete, and incomplete in the direction that flatters. Every probability the programme had quoted measured passing the evaluation; the funded stage that follows applies a stricter rule on an account whose drawdown floor has already ratcheted up. Passing is a one-time hurdle, getting paid requires surviving and then accumulating again -- so the standard metric flatters weak strategies, and flatters them worse the weaker they are. Registered at zero alpha: a capability demonstration, not an edge hunt.

source pinned at 0052f2af885e78a5…

r1 controls passed

recorded
2026-08-09T20:48:28+00:00
engine_sha
bcbe4cfcc74a84314e2dfd7a3c1efe27b7bc87a8-dirty
spend
$0.00
metricvalue
eval_horizon90
funded_horizon180
n1652
obtainable_limit.n1652
obtainable_limit.p_breach_while_funded0.0571
obtainable_limit.p_first_payout0.0115
obtainable_limit.p_pass0.0424
obtainable_limit.p_two_cycles0.0
payout_params_verifiedFalse
sealed_unobtainable.n1652
sealed_unobtainable.p_breach_while_funded0.5117
sealed_unobtainable.p_first_payout0.1659
sealed_unobtainable.p_pass0.2579
sealed_unobtainable.p_two_cycles0.0841
Provenance the frozen script and its output, archived before the run
exp_q1_payout.py5,276 B92ca821ad32e4eba…
output.txt1,221 Bb41c460fdf5476a5…
Config
{
 "eval_horizon": 90,
 "funded_horizon": 180,
 "instrument": "MES",
 "payout": {
  "VERIFIED": false,
  "consistency_frac": 0.4,
  "min_days": 5,
  "payout_frac": 1.0,
  "threshold": 2000,
  "trailing_dd": 2000
 },
 "risk_usd": 175.0
}
Note
Q1 capability demonstration, zero alpha. Payout parameters are ILLUSTRATIVE and unverified — the figures show a shape, not a venue.

SCRIPT: experiments/Q1-PAYOUT-PATH/r1/exp_q1_payout.py
OUTPUT: experiments/Q1-PAYOUT-PATH/r1/output.txt
CALCS: none — the script called no audited estimator
Q2-ATTRIBUTION4.2% pass, 44.0% breach, 35.4% timeout ahead, 16.4% timeout underwater2 runs
axis
objective (capability, not a market question)
search space
1
alpha
0.0
outcome
4.2% pass, 44.0% breach, 35.4% timeout ahead, 16.4% timeout underwater
effect
—
interval
—
supersedes
E2-OBJECTIVE

Statement

Every failed attempt can be assigned exactly one principal reason, and those reasons plus P(pass) sum to one. Demonstrated by decomposing E2's already-recorded runs.

Mechanism written before the number existed

NOT A HYPOTHESIS TEST — zero alpha. The decomposition asks nothing new of the market; it re-reads runs E2 already produced. No number here can revive the fade.

The gap: E2 reported P(pass) 0.042 and P(breach) 0.440, which do not sum to one, so roughly half of all attempts ended in some unnamed way. 'Unnamed' contains at least three genuinely different outcomes with genuinely different remedies -- an attempt that ran out of horizon while still alive wants MORE DAYS, one that breached wants LESS SIZE and more days would only have breached sooner, and one blocked by a consistency rule wants neither because it had the money and could not claim it. A single probability is a scoreboard; this is a diagnosis.

The rule that makes it honest: EXACTLY ONE REASON PER PATH, chosen by a DECLARED precedence rather than by whichever check happens to run first -- the latter reads as an ordering detail and behaves as a definition. Shares are computed over ALL attempts including passes, so P(pass) plus the shares sums to one; reporting shares of failures only would invite double-counting.

One reason exists purely to protect a past lesson: NEVER_TRADED. Verdict #1 was void for instrument failure, and a run whose equity never moved must be named as that rather than counted as a market result.

Resolved 4.2% pass, 44.0% breach, 35.4% timeout ahead, 16.4% timeout underwater documented

stated in prose, never computed into the register

stated in
docs/Q2-ATTRIBUTION.md
resolved
2026-08-30T16:52:20+00:00

No effect size, interval or floor is shown because none was ever computed into the register. The source document’s numbers are authoritative; this record is not.

Decision the part no run can produce

A single probability could not have said this. Given only P(pass)=0.042 a sensible reader concludes the horizon was too short -- which is wrong in the most expensive possible direction, because more days on a breaching strategy produce more breaches. The decomposition points at less size instead. Registered at zero alpha as a capability demonstration on already-recorded runs.

source pinned at 7d7976dde6a5519c…

r2-timeout-split controls passed

recorded
2026-08-09T20:58:57+00:00
engine_sha
68784dd80e300e8e196ccd0b817fa68ad1fd520e-dirty
spend
$0.00
metricvalue
horizon90
n1652
obtainable_limit.n1652
obtainable_limit.p_pass0.0424
obtainable_limit.reasons.breach_drawdown726
obtainable_limit.reasons.timeout_in_profit585
obtainable_limit.reasons.timeout_underwater271
obtainable_limit.shares.breach_drawdown0.4395
obtainable_limit.shares.timeout_in_profit0.3541
obtainable_limit.shares.timeout_underwater0.164
sealed_unobtainable.n1652
sealed_unobtainable.p_pass0.2579
sealed_unobtainable.reasons.breach_drawdown454
sealed_unobtainable.reasons.timeout_in_profit656
sealed_unobtainable.reasons.timeout_underwater116
sealed_unobtainable.shares.breach_drawdown0.2748
sealed_unobtainable.shares.timeout_in_profit0.3971
sealed_unobtainable.shares.timeout_underwater0.0702
Provenance the frozen script and its output, archived before the run
exp_q2_attribution.py3,304 Bf3bf4488252a974a…
output.txt1,807 Bf96301dd8db38298…
Config
{
 "change": "TIMEOUT split into in_profit vs underwater",
 "horizon": 90,
 "instrument": "MES",
 "risk_usd": 175.0,
 "supersedes_run": "r1"
}
Note
Q2 r2. r1 put every surviving-but-short attempt in one TIMEOUT bucket. Its own example detail showed an attempt '4,032 short of the 3,000 target' — i.e. DOWN 1,032, not nearly there. 'Ran out of days' implies the remedy is a longer horizon; for a losing strategy more days produce more BREACHES. Not estimand-shopping: this is a zero-alpha descriptive bucketing, no tested quantity changed, and r1 stays on the record.

SCRIPT: experiments/Q2-ATTRIBUTION/r2-timeout-split/exp_q2_attribution.py
OUTPUT: experiments/Q2-ATTRIBUTION/r2-timeout-split/output.txt
CALCS: none — the script called no audited estimator

r1 controls passed

recorded
2026-08-09T20:57:45+00:00
engine_sha
68784dd80e300e8e196ccd0b817fa68ad1fd520e-dirty
spend
$0.00
metricvalue
horizon90
n1652
obtainable_limit.n1652
obtainable_limit.p_pass0.0424
obtainable_limit.reasons.breach_drawdown726
obtainable_limit.reasons.timeout856
obtainable_limit.shares.breach_drawdown0.4395
obtainable_limit.shares.timeout0.5182
sealed_unobtainable.n1652
sealed_unobtainable.p_pass0.2579
sealed_unobtainable.reasons.breach_drawdown454
sealed_unobtainable.reasons.timeout772
sealed_unobtainable.shares.breach_drawdown0.2748
sealed_unobtainable.shares.timeout0.4673
Provenance the frozen script and its output, archived before the run
exp_q2_attribution.py3,304 Bf3bf4488252a974a…
output.txt1,116 Bb06a5ef34528db47…
Config
{
 "horizon": 90,
 "instrument": "MES",
 "precedence": [
  "breach_drawdown",
  "never_traded",
  "consistency",
  "min_days",
  "timeout"
 ],
 "risk_usd": 175.0
}
Note
Q2 capability demonstration, zero alpha. Decomposes E2's own runs.

SCRIPT: experiments/Q2-ATTRIBUTION/r1/exp_q2_attribution.py
OUTPUT: experiments/Q2-ATTRIBUTION/r1/output.txt
CALCS: none — the script called no audited estimator
T1-TARGET-LADDERnull1 run
axis
price (1-minute bars, owned; no new data)
search space
6
alpha
0.05
outcome
null
effect
-0.0213
interval
[-0.0893, +0.0523]

Statement

At a FIXED stop distance, raising the target multiple (target_r) improves expected value on a funded-account contract. PRIMARY: EV in dollars on the best-ranked contract is monotone non-decreasing in target_r across the declared ladder {1.0, 1.5, 2.0, 3.0, 4.0, 6.0}. SECONDARY: net R per trade across the same ladder, with a cluster-bootstrap interval per rung.

Mechanism written before the number existed

A funded-account evaluation is a BARRIER-CROSSING problem -- reach +target before a trailing drawdown ratchets up behind you -- not an expectancy problem. For a barrier, asymmetry dominates the mean: FEASIBILITY.md's frontier shows the required win rate falling from 60-65% at a 1.0R payoff to 45% at 2.0R.

The evidence for running it forwards comes from running it backwards. The archived H-STOPWIDTH ladder widens the STOP at a fixed target, which mechanically LOWERS the payoff ratio: MES goes 2.09 -> 1.21 -> 0.79 -> 0.54 across x1..x8. Net R per trade IMPROVES over that range (-0.046 -> -0.013) because cost in R falls 15x as the fixed per-contract cost is amortised over a larger per-contract risk -- and yet EV through the contract menu COLLAPSES (-$38 -> -$269). Expectancy and EV disagree, and the barrier is why.

So the mechanism predicts the mirror image: moving the TARGET outward at a fixed stop raises the payoff ratio, and if the market's hit-rate decay with target distance is slower than the payoff gain, EV rises. This axis has never been tested. The sealed grid varied k_stop and width_max and recorded 'Fixed, not searched: target = far side'.

BOTH DIRECTIONS, INTERPRETED IN ADVANCE. A rising EV curve means asymmetry is the untested lever and the axis reopens -- NOT that the ORB family reopens, which is closed on its own evidence. A flat or falling curve means hit-rate decay exactly offsets the payoff gain, and the 'exhausted' conclusion then extends to the one dimension that had never been searched, which is a stronger negative than the programme currently holds.

Power plan

{
 "alpha": 0.05,
 "cluster_size": 2,
 "effect": 0.05,
 "intra_cluster_r": 0.9,
 "power": 0.8,
 "search_space_size": 6,
 "test": "mean_shift"
}

Resolved null measured

read from a scored result in an archived run

from run
2026-08-30-ladder-v2 ladder.t6.0
effect
-0.02128
95% CI
[-0.0893, +0.0523]
floor
0.11361893769086785
resolved
2026-08-31T11:43:18+00:00
Decision the part no run can produce

MECHANISM CONFIRMED IN DIRECTION, INSUFFICIENT IN MAGNITUDE. The ladder is monotone increasing exactly as registered: net R per trade rises from -0.1028 at target_r=1.0 to -0.0213 at 6.0, a span of +0.0816R. That span is LARGER THAN THE ENTIRE GROSS SIGNAL nine families produced (+0.02 to +0.04R), so asymmetry is a real lever and the axis was worth searching. It is not enough. The ladder converts a DETECTABLE loss into an INDISTINGUISHABLE one -- verdicts walk -1.95x floor, -1.53x, -1.17x, 'precise but immaterial', then null, null. The best rung sits at -0.0213R against a +0.003R break-even on the best-ranked contract. It does not reach profitability. AND THIS DESIGN CANNOT RESOLVE WHETHER IT COULD. The detectable floor GROWS with target_r (0.0526 -> 0.1136) because variance rises with the target distance, so the test loses power precisely where the answer becomes interesting. At the top rung the floor is 38x the break-even requirement. The upper interval bound (+0.0523) clears break-even and the estimate does not: that is inconclusive, not null, and pretending otherwise would be the error this lab exists to catch. The correction it forces on the record: 'stop width did not matter at all' is wrong. Trade GEOMETRY matters more than any signal this programme found. It was never searched because the sealed grid fixed the target at the far side.

source pinned at b85cd8e779a6ee2f…

2026-08-30-ladder-v2 controls passed

recorded
2026-08-31T11:42:35+00:00
engine_sha
da556eaa427a3591b4b7bd6df7263ee7c4108c7b-dirty
spend
$0.00
metricvalue
best_target_r6.0
ladder.t1.0.ci[-0.1323610022746174, -0.07047785662557962]
ladder.t1.0.crosses_zeroFalse
ladder.t1.0.d-0.10284912191827084
ladder.t1.0.estimate-0.10284912191827084
ladder.t1.0.floor0.05262286279943919
ladder.t1.0.floor_multiples-1.9544569878354645
ladder.t1.0.n3384
ladder.t1.0.n_eff1781
ladder.t1.0.nameT1 target_r=1.0 - net R per trade
ladder.t1.0.notewin rate 0.476; ORB vehicle, stop_range=0.5, risk=$175
ladder.t1.0.unitR
ladder.t1.0.verdictdetectable
ladder.t1.5.ci[-0.13395610533552763, -0.05774732086542981]
ladder.t1.5.crosses_zeroFalse
ladder.t1.5.d-0.0973464412360689
ladder.t1.5.estimate-0.0973464412360689
ladder.t1.5.floor0.06357818583823024
ladder.t1.5.floor_multiples-1.5311295840331047
ladder.t1.5.n3384
ladder.t1.5.n_eff1781
ladder.t1.5.nameT1 target_r=1.5 - net R per trade
ladder.t1.5.notewin rate 0.385; ORB vehicle, stop_range=0.5, risk=$175
ladder.t1.5.unitR
ladder.t1.5.verdictdetectable
ladder.t2.0.ci[-0.12810560011677785, -0.0397759305814457]
ladder.t2.0.crosses_zeroFalse
ladder.t2.0.d-0.08511039344815942
ladder.t2.0.estimate-0.08511039344815942
ladder.t2.0.floor0.07297645225758571
ladder.t2.0.floor_multiples-1.1662720071365542
ladder.t2.0.n3384
ladder.t2.0.n_eff1781
ladder.t2.0.nameT1 target_r=2.0 - net R per trade
ladder.t2.0.notewin rate 0.330; ORB vehicle, stop_range=0.5, risk=$175
ladder.t2.0.unitR
ladder.t2.0.verdictdetectable
ladder.t3.0.ci[-0.11864270901508335, -0.01423587617806432]
ladder.t3.0.crosses_zeroFalse
ladder.t3.0.d-0.06694465552178319
ladder.t3.0.estimate-0.06694465552178319
ladder.t3.0.floor0.08789775560052443
ladder.t3.0.floor_multiples-0.7616196234410311
ladder.t3.0.n3384
ladder.t3.0.n_eff1781
ladder.t3.0.nameT1 target_r=3.0 - net R per trade
ladder.t3.0.notewin rate 0.268; ORB vehicle, stop_range=0.5, risk=$175
ladder.t3.0.unitR
ladder.t3.0.verdictprecise but immaterial
ladder.t4.0.ci[-0.09563783861690606, 0.023625276691543632]
ladder.t4.0.crosses_zeroTrue
ladder.t4.0.d-0.03717261904761905
ladder.t4.0.estimate-0.03717261904761905
ladder.t4.0.floor0.1000020079817115
ladder.t4.0.floor_multiples-0.37171872643214554
ladder.t4.0.n3384
ladder.t4.0.n_eff1781
ladder.t4.0.nameT1 target_r=4.0 - net R per trade
ladder.t4.0.notewin rate 0.242; ORB vehicle, stop_range=0.5, risk=$175
ladder.t4.0.unitR
ladder.t4.0.verdictnull
ladder.t6.0.ci[-0.08926305648169772, 0.052316549028921924]
ladder.t6.0.crosses_zeroTrue
ladder.t6.0.d-0.021279128672745704
ladder.t6.0.estimate-0.021279128672745704
ladder.t6.0.floor0.11361893769086785
ladder.t6.0.floor_multiples-0.1872850521683413
ladder.t6.0.n3384
ladder.t6.0.n_eff1781
ladder.t6.0.nameT1 target_r=6.0 - net R per trade
ladder.t6.0.notewin rate 0.223; ORB vehicle, stop_range=0.5, risk=$175
ladder.t6.0.unitR
ladder.t6.0.verdictnull
monotone_increasingTrue
n20304
n_eff10686
span_R0.08156999324552514
Provenance the frozen script and its output, archived before the run
calcs.json1,795 B57c889e46d7fa5d5…
exp_t1_target_ladder.py5,091 B6f77d67f1658b3c1…
output.txt3,268 B20812ecd5f04f558…
Config
{
 "cluster_size": 2,
 "instruments": [
  "MES",
  "MNQ"
 ],
 "intra_cluster_r": 0.9,
 "ladder": [
  1.0,
  1.5,
  2.0,
  3.0,
  4.0,
  6.0
 ],
 "range_minutes": 15,
 "risk_usd": 175.0,
 "seed": 20260830,
 "stop_range": 0.5,
 "supersedes_run": "2026-08-30-ladder (refused by the power gate: no top-level n emitted)",
 "vehicle": "ORB (closed family)"
}
Note
T1. Target ladder at a fixed stop -- the axis the sealed grid recorded as "Fixed, not searched".

SCRIPT: experiments/T1-TARGET-LADDER/2026-08-30-ladder-v2/exp_t1_target_ladder.py
OUTPUT: experiments/T1-TARGET-LADDER/2026-08-30-ladder-v2/output.txt
CALCS: experiments/T1-TARGET-LADDER/2026-08-30-ladder-v2/calcs.json
T2-COST-CONDITIONEDnull1 run
axis
price (1-minute bars, owned; no new data)
search space
4
alpha
0.05
outcome
null
effect
+0.0230
interval
[-0.0361, +0.0840]

Statement

Selecting trades by their COST IN R improves net expectancy on the retained subset. PRIMARY: mean net R per trade on trades whose cost-in-R falls below a threshold exceeds the unconditioned mean by more than the detectable floor, across the declared threshold ladder {no filter, 0.10, 0.07, 0.05}.

Mechanism written before the number existed

Computed today, not assumed. Cost in R = (fixed per-contract cost) / (per-contract risk), and per-contract risk = (stop_dist + slippage) * multiplier + 2 * commission. The contract count cancels, so cost in R is INVARIANT to risk per trade -- but it varies 15x with stop distance on MES (0.273 at a 2-point stop to 0.018 at 40 points), because a fixed cost is amortised over a larger per-contract risk.

Stop distance is set by the opening-range height, which is known BEFORE the trade is taken. So cost in R is an ex-ante observable, and filtering on it removes the most cost-burdened trades without touching the target, the stop multiple, or the signal.

This is the one way to raise net expectancy that does NOT buy it with variance. T1 established that pushing the target outward raises the mean (+0.0816R across the ladder) while raising the standard deviation from 0.79 to 1.71, so the detectable floor grew 0.053 -> 0.114 and the test lost power exactly where the answer became interesting. Filtering on cost changes the mean through the cost term while leaving the per-trade return distribution's shape alone.

BOTH DIRECTIONS, INTERPRETED IN ADVANCE. A monotone improvement across the ladder means cost burden is a selectable property and the geometry axis has a second lever. A flat or falling curve means the trades with low cost-in-R are also the trades with worse gross outcomes -- large ranges being harder to trade -- and the two effects cancel, which closes cost-conditioning and is a stronger negative than the programme holds today.

Power plan

{
 "alpha": 0.05,
 "cannot_detect": "the 0.024R gap to EV break-even",
 "cluster_size": 2,
 "effect": 0.1,
 "intra_cluster_r": 0.9,
 "power": 0.8,
 "search_space_size": 4,
 "test": "mean_shift"
}

Resolved null measured

read from a scored result in an archived run

from run
2026-08-30-cost-ladder ladder.c0.05
effect
+0.02296
95% CI
[-0.0361, +0.0840]
floor
0.11543800752505227
resolved
2026-08-31T12:17:32+00:00
Decision the part no run can produce

MONOTONE, POSITIVE AT THE TOP, AND NOT ESTABLISHED. The registered primary was that filtering beats the unconditioned mean by more than the detectable floor. It does not: the improvement is +0.1081R against a floor of 0.1154R -- 94% of the way, and short. The registered answer is therefore NO. What the ladder shows is nonetheless the strongest thing this programme has produced. Net R rises monotonically as the cost filter tightens: -0.0851 (detectable loss, no filter) -> -0.0321 -> -0.0115 -> +0.0230 at cost_R <= 0.05, keeping 36% of trades. That top rung is the FIRST POSITIVE NET EXPECTANCY the programme has recorded, and it sits above the +0.003R break-even on the best-ranked contract. IT IS NOT A FINDING AND MUST NOT BE READ AS ONE. The interval is [-0.0361, +0.0840] and crosses zero: verdict null. A positive point estimate whose interval spans zero is exactly the shape that produced three false results in this lab already, and the difference here is only that the sign is pleasing. WHAT WOULD RESOLVE IT, computed rather than hoped. The floor scales as 1/sqrt(n_eff). Adding MGC and MCL multiplies effective sample by 1.85x -- metals and energy correlate far less with equity index than MES and MNQ do with each other -- dropping the floor from 0.1154 to 0.0848, BELOW the measured +0.1081 improvement. For $19.16, already inside the data cap, this becomes resolvable. That is the first time in this programme that a purchase has been justified by a measured result rather than a hunch.

source pinned at fabbe0cd5667221c…

2026-08-30-cost-ladder controls passed

recorded
2026-08-31T12:16:44+00:00
engine_sha
2863c7020bc8be0b06e14cfdcd29468f9152ec83-dirty
spend
$0.00
metricvalue
best_thresholdc0.05
exceeds_floorFalse
improvement_over_nofilter0.10807532003869774
ladder.c0.05.ci[-0.03611102992428878, 0.08402355326712602]
ladder.c0.05.crosses_zeroTrue
ladder.c0.05.d0.02296492659053833
ladder.c0.05.estimate0.02296492659053833
ladder.c0.05.floor0.11543800752505227
ladder.c0.05.floor_multiples0.1989373091488477
ladder.c0.05.n1226
ladder.c0.05.n_eff645
ladder.c0.05.nameT2 cost_R<=0.05 - net R per trade
ladder.c0.05.notekept 36.2% of trades
ladder.c0.05.unitR
ladder.c0.05.verdictnull
ladder.c0.07.ci[-0.06391831850416482, 0.0406088726423246]
ladder.c0.07.crosses_zeroTrue
ladder.c0.07.d-0.011516884531590416
ladder.c0.07.estimate-0.011516884531590416
ladder.c0.07.floor0.09657798063502403
ladder.c0.07.floor_multiples-0.11924958935633218
ladder.c0.07.n1836
ladder.c0.07.n_eff966
ladder.c0.07.nameT2 cost_R<=0.07 - net R per trade
ladder.c0.07.notekept 54.3% of trades
ladder.c0.07.unitR
ladder.c0.07.verdictnull
ladder.c0.1.ci[-0.08037502328603563, 0.017911689229394644]
ladder.c0.1.crosses_zeroTrue
ladder.c0.1.d-0.0321101027705992
ladder.c0.1.estimate-0.0321101027705992
ladder.c0.1.floor0.086290625752713
ladder.c0.1.floor_multiples-0.37211577144681496
ladder.c0.1.n2377
ladder.c0.1.n_eff1251
ladder.c0.1.nameT2 cost_R<=0.1 - net R per trade
ladder.c0.1.notekept 70.2% of trades
ladder.c0.1.unitR
ladder.c0.1.verdictnull
ladder.nofilter.ci[-0.12810560011677785, -0.0397759305814457]
ladder.nofilter.crosses_zeroFalse
ladder.nofilter.d-0.08511039344815942
ladder.nofilter.estimate-0.08511039344815942
ladder.nofilter.floor0.07297645225758571
ladder.nofilter.floor_multiples-1.1662720071365542
ladder.nofilter.n3384
ladder.nofilter.n_eff1781
ladder.nofilter.nameT2 cost_R<=None - net R per trade
ladder.nofilter.notekept 100.0% of trades
ladder.nofilter.unitR
ladder.nofilter.verdictdetectable
n3384
n_eff1781
Provenance the frozen script and its output, archived before the run
calcs.json1,359 Bb278820795741e7a…
exp_t2_cost_conditioned.py5,983 Bbbfedf9d202867c5…
output.txt2,292 Bf5b548d42ba4fa94…
Config
{
 "cannot_detect": "the 0.024R gap to EV break-even; smallest available floor is 0.0526R",
 "cluster_size": 2,
 "instruments": [
  "MES",
  "MNQ"
 ],
 "intra_cluster_r": 0.9,
 "ladder": [
  "nofilter",
  0.1,
  0.07,
  0.05
 ],
 "range_minutes": 15,
 "risk_usd": 175.0,
 "seed": 20260830,
 "stop_range": 0.5,
 "target_r": 2.0,
 "vehicle": "ORB (closed family)"
}
Note
T2. Cost-in-R is an ex-ante observable (it follows from the opening-range height), so filtering on it raises the mean through the cost term without touching target, stop multiple or signal -- the one lever found so far that does not buy the mean with variance. Registered with its power limit stated: this cannot resolve +EV and does not claim to.

SCRIPT: experiments/T2-COST-CONDITIONED/2026-08-30-cost-ladder/exp_t2_cost_conditioned.py
OUTPUT: experiments/T2-COST-CONDITIONED/2026-08-30-cost-ladder/output.txt
CALCS: experiments/T2-COST-CONDITIONED/2026-08-30-cost-ladder/calcs.json
T3-COST-FILTER-REPLICATIONNOT RUN - programme closed before the data was bought0 runs
axis
price (1-minute bars) on metals and energy -- NEW instruments, not owned
search space
1
alpha
0.05
outcome
NOT RUN - programme closed before the data was bought
effect
—
interval
—

Statement

The cost-in-R filter that improved net expectancy on index futures also improves it on metals and energy. PRIMARY, single pre-specified test with NO freedom: on MGC and MCL alone, mean net R per trade at cost_R <= 0.05 exceeds the unfiltered mean on the same instruments by more than the detectable floor. The threshold 0.05, the vehicle (ORB, stop_range=0.5, target_r=2.0), the risk ($175) and the range window (15 min) are all FIXED at T2's values and may not be re-searched. SECONDARY: the pooled four-instrument estimate, reported only if the validity gates below pass.

Mechanism written before the number existed

T2 established the direction on MES and MNQ: net R rose monotonically as the cost filter tightened (-0.0851 unfiltered -> +0.0230 at cost_R <= 0.05, keeping 36% of trades), an improvement of +0.1081R against a floor of 0.1154R -- 94% of the way and SHORT, with an interval crossing zero.

The mechanism is arithmetic rather than a story about participants: cost in R = fixed per-contract cost / per-contract risk, so it falls as the stop widens, and the stop follows the opening-range height which is known before the trade. Nothing in that reasoning is specific to equity index futures, so it should hold on metals and energy -- and if it does not, the mechanism was never what produced T2's number.

BOTH DIRECTIONS, INTERPRETED IN ADVANCE. Replication means the effect is a property of trade geometry rather than of two correlated index contracts, and the pooled sample then resolves what T2 could not. Failure means the effect is either instrument-specific or was noise dressed as a monotone trend -- and a monotone ladder of four rungs is exactly what noise looks like often enough to matter. EITHER ANSWER IS WORTH $19.16; the failure is the more informative one, because it would retire the only positive number this programme has produced.

Power plan

{
 "alpha": 0.05,
 "cluster_size": 2,
 "effect": 0.1081,
 "intra_cluster_r": 0.9,
 "power": 0.8,
 "precondition": "measure cross-asset rho before scoring",
 "search_space_size": 1,
 "test": "mean_shift"
}

Resolved NOT RUN - programme closed before the data was bought documented

stated in prose, never computed into the register

stated in
docs/RESULTS.md
resolved
2026-08-31T14:32:16+00:00

No effect size, interval or floor is shown because none was ever computed into the register. The source document’s numbers are authoritative; this record is not.

Decision the part no run can produce

CLOSED WITHOUT RUNNING, and the reason is the finding rather than an abandonment. T3 was registered before any purchase precisely so that buying data could not also buy the freedom to design the test around it. That discipline held: the hypothesis is on the record, unrun, with its four validity gates intact. It is closed because the ROUTE is closed. Base rates from a major provider's own performance disclosure put ~5.6% of evaluation attempts at a paid trader and evaluation fees at 70-95% of provider revenue -- a fee business. Buying $19.16 of metals and energy bars to resolve one open question inside a route being abandoned would be spending real money on tidiness. WHAT IT WOULD HAVE ANSWERED, recorded so it is not lost. T2 and T4 share the same 3,384 trades: the cost filter selects by setup, so changing the target changed the outcome definition and not the sample. No further re-analysis of owned data can distinguish 'a property of trade geometry' from 'a property of these two instruments'. Only new instruments could, and that is what T3 was for. The question is therefore OPEN, not answered -- and if the cost-filter mechanism is ever carried to another venue, this is the test that has to run first. Gate 1 also never passed: `occams/instruments.py` carries MGC and MCL as RECORDED-NOT-VERIFIED behind a refusal, because the exchange's published specs were unreachable and slippage cannot be known without a trades tape the $19.16 quote does not cover.

source pinned at 49951ec62612c283…

T4-COST-FILTER-AT-LOW-TARGETnull1 run
axis
price (1-minute bars, owned; no new data)
search space
4
alpha
0.05
outcome
null
effect
-0.0218
interval
[-0.0635, +0.0195]

Statement

The cost-in-R filter improves net expectancy at target_r=1.0 as it did at 2.0. PRIMARY: across the same declared threshold ladder {no filter, 0.10, 0.07, 0.05}, mean net R per trade is monotone non-decreasing as the threshold tightens, and the improvement from unfiltered to the tightest rung exceeds the detectable floor.

Mechanism written before the number existed

A ROBUSTNESS CHECK, not a search for a better number, and it is registered that way so it cannot be reported as the latter.

cost_R = fixed per-contract cost / per-contract risk, and both terms come from the opening-range height. It is therefore a property of the SETUP and independent of the target: the filter selects the SAME TRADES at any target_r. Only the outcome distribution differs. So if cost filtering is a real mechanism it must improve expectancy at 1.0 as it did at 2.0, and if it works only at the target it was found on, it is an interaction or it is noise.

There is a second, statistical reason to prefer 1.0. T1 measured the standard deviation of net R per trade at each target: 0.79 at 1.0 rising to 1.71 at 6.0. The detectable floor scales with sd, so the same filter at 1.0 is measured against a floor roughly 25% lower than T2's -- attacking the floor through VARIANCE rather than through sample size, which is the only lever left that costs nothing.

BOTH DIRECTIONS, INTERPRETED IN ADVANCE. Monotone and clearing its floor CORROBORATES T2 on an independent target and makes the mechanism claim substantially harder to dismiss. Flat, non-monotone, or failing its floor UNDERMINES T2: a four-rung monotone ladder that does not reproduce one target step away is what noise looks like, and this programme has already published three results that looked real and were not.

Power plan

{
 "alpha": 0.05,
 "cluster_size": 2,
 "effect": 0.1,
 "intra_cluster_r": 0.3818,
 "power": 0.8,
 "search_space_size": 4,
 "test": "mean_shift"
}

Resolved null measured

read from a scored result in an archived run

from run
2026-08-30-low-target ladder.c0.05
effect
-0.02179
95% CI
[-0.0635, +0.0195]
floor
0.06945045557523696
resolved
2026-08-31T13:11:48+00:00
Decision the part no run can produce

REGISTERED PRIMARY SATISFIED -- the first time in this programme. The ladder is monotone (-0.1028 -> -0.0494 -> -0.0386 -> -0.0218) and the improvement from unfiltered to tightest is +0.0811R against a floor of 0.0695R. It CLEARS, where T2's +0.1081 against 0.1154 did not. The variance lever worked as predicted from T1's measurements: the top-rung floor is 0.0695 here against T2's 0.0984 under the same measured rho -- 29% lower, close to the ~25% predicted from sd falling 1.05 to 0.79. Attacking the floor through variance rather than sample size cost nothing and was the difference between clearing and not. THE STRATEGY STILL LOSES. The best rung is -0.0218R with an interval [-0.0635, +0.0195] crossing zero: verdict null. What is established is that FILTERING IMPROVES EXPECTANCY, not that the filtered strategy is profitable. Those remain different claims and only the first has cleared a bar. THIS IS NOT A REPLICATION, AND MUST NOT BE COUNTED AS ONE. cost_R is a property of the setup, so the filter selects the SAME 3,384 trades here as in T2 -- only the outcome definition changed. T2 and T4 are one filter evaluated against two targets, not two independent tests. The corroboration is real but weaker than it looks: it shows the effect is not an artifact of one target choice, and it cannot show the effect is not an artifact of these two instruments. That is precisely what T3 and the MGC/MCL purchase are for, and this result strengthens the case for it rather than removing it.

source pinned at 7d18955e7df100d9…

2026-08-30-low-target controls passed

recorded
2026-08-31T13:11:03+00:00
engine_sha
4efa7c3a79cbfcd50e564dc712eadea996dfcdd7-dirty
spend
$0.00
metricvalue
best_thresholdc0.05
exceeds_floorTrue
improvement_over_nofilter0.08105991194390588
ladder.c0.05.ci[-0.06345782047920316, 0.0194966425222209]
ladder.c0.05.crosses_zeroTrue
ladder.c0.05.d-0.02178920997436495
ladder.c0.05.estimate-0.02178920997436495
ladder.c0.05.floor0.06945045557523696
ladder.c0.05.floor_multiples-0.3137374664268443
ladder.c0.05.n1226
ladder.c0.05.n_eff887
ladder.c0.05.nameT4 cost_R<=0.05 - net R per trade
ladder.c0.05.notekept 36.2% of trades
ladder.c0.05.unitR
ladder.c0.05.verdictnull
ladder.c0.07.ci[-0.07581366029693601, -0.001317963327001998]
ladder.c0.07.crosses_zeroFalse
ladder.c0.07.d-0.0386484593837535
ladder.c0.07.estimate-0.0386484593837535
ladder.c0.07.floor0.05859399141997629
ladder.c0.07.floor_multiples-0.6595976557858658
ladder.c0.07.n1836
ladder.c0.07.n_eff1328
ladder.c0.07.nameT4 cost_R<=0.07 - net R per trade
ladder.c0.07.notekept 54.3% of trades
ladder.c0.07.unitR
ladder.c0.07.verdictprecise but immaterial
ladder.c0.1.ci[-0.08474646162629965, -0.013315194486885654]
ladder.c0.1.crosses_zeroFalse
ladder.c0.1.d-0.04939419436264199
ladder.c0.1.estimate-0.04939419436264199
ladder.c0.1.floor0.052568626722908414
ladder.c0.1.floor_multiples-0.9396135573980465
ladder.c0.1.n2377
ladder.c0.1.n_eff1720
ladder.c0.1.nameT4 cost_R<=0.1 - net R per trade
ladder.c0.1.notekept 70.2% of trades
ladder.c0.1.unitR
ladder.c0.1.verdictprecise but immaterial
ladder.nofilter.ci[-0.1323610022746174, -0.07047785662557962]
ladder.nofilter.crosses_zeroFalse
ladder.nofilter.d-0.10284912191827084
ladder.nofilter.estimate-0.10284912191827084
ladder.nofilter.floor0.04488494774988622
ladder.nofilter.floor_multiples-2.291394489114261
ladder.nofilter.n3384
ladder.nofilter.n_eff2448
ladder.nofilter.nameT4 cost_R<=None - net R per trade
ladder.nofilter.notekept 100.0% of trades
ladder.nofilter.unitR
ladder.nofilter.verdictdetectable
n3384
n_eff2448
Provenance the frozen script and its output, archived before the run
calcs.json1,359 B0926b4d0024be879…
exp_t4_cost_low_target.py6,579 B38e08469b8e24509…
output.txt2,353 B99e28a84a969b4a0…
Config
{
 "cluster_size": 2,
 "corroborates_or_undermines": "T2-COST-CONDITIONED at target_r=2.0",
 "instruments": [
  "MES",
  "MNQ"
 ],
 "intra_cluster_r": 0.3818,
 "ladder": [
  "nofilter",
  0.1,
  0.07,
  0.05
 ],
 "range_minutes": 15,
 "rho_note": "MEASURED on 1,637 paired days, not the asserted 0.9",
 "risk_usd": 175.0,
 "seed": 20260830,
 "stop_range": 0.5,
 "target_r": 1.0,
 "vehicle": "ORB (closed family)"
}
Note
T4. Robustness check on T2, registered as one. cost_R is a property of the SETUP (it follows from the opening-range height), so the filter selects the same trades at any target -- if the mechanism is real it must work here too. Also attacks the floor through variance rather than sample size: T1 measured sd 0.79 at target_r=1.0 against 1.05 at 2.0.

SCRIPT: experiments/T4-COST-FILTER-AT-LOW-TARGET/2026-08-30-low-target/exp_t4_cost_low_target.py
OUTPUT: experiments/T4-COST-FILTER-AT-LOW-TARGET/2026-08-30-low-target/output.txt
CALCS: experiments/T4-COST-FILTER-AT-LOW-TARGET/2026-08-30-low-target/calcs.json
X1-COMPRESSIONdetectable1 run
axis
price (1-minute bars, owned)
search space
2
alpha
0.05
outcome
detectable
effect
-0.0890
interval
[-0.1345, -0.0431]

Statement

Volatility compression predicts subsequent expansion. PRIMARY: the count of consecutive contracting RTH sessions ending at day N-1 is POSITIVELY rank-correlated with day N's realised RTH range divided by its trailing ATR. SECONDARY: day N's opening-range height over its trailing 20-day median is NEGATIVELY rank-correlated with the post-range expansion of the same day.

Mechanism written before the number existed

Volatility clusters and mean-reverts. Low-volatility states are followed by higher-volatility ones more often than chance. This is the strongest prior available to this programme because it is a DOCUMENTED STATISTICAL PROPERTY of financial time series rather than a story about what participants are thinking -- every other family in the catalogue rests on the latter.

WHY THE PRIMARY USES PRIOR DAYS ONLY, decided on design grounds before any number exists. The obvious predictor -- today's opening range -- is contaminated: the opening range and the rest of the same session are both measurements of that day's volatility, and intraday volatility PERSISTS. A quiet open genuinely predicts a quiet day. That confound would swamp any across-day mean reversion and the result would be uninterpretable. The consecutive-contraction count is computed entirely from COMPLETED PRIOR SESSIONS, so it cannot contain any part of the outcome. This is the C2 lesson (a predictor partly ALGEBRAIC with its outcome) and the ICT-P2 lesson (a measurement contaminating its own answer) applied at the design stage rather than discovered afterwards.

TWO DIRECTIONS, BOTH INTERPRETED IN ADVANCE, so neither can be reinterpreted after the fact:
  * PRIMARY POSITIVE -> the mechanism holds: a longer run of contracting sessions is followed by a larger range.
  * PRIMARY NEGATIVE -> the mechanism is refuted in the sharpest way: compression predicts MORE compression, i.e. volatility persists across days rather than mean-reverting at this horizon.
  * PRIMARY ~ZERO -> compression carries no information about what follows, and the family is dead.
  * SECONDARY POSITIVE is EXPECTED EVEN IF THE MECHANISM IS FALSE, because of the same-day persistence named above. It is reported as a confound check, not as evidence.

The outcome is strategy-free -- a realised range, with no entry, no stop and no target -- because a P&L outcome is how the fade's unobtainable fill entered the record. The eventual entry, if this ever advances, is a STOP order beyond the compressed range, which arms while price is inside it and is therefore obtainable by construction. That obtainability is why this family was worth queuing first.

PRIOR: genuinely uncertain, which is unusual here and is the point. Volatility clustering is real and well documented; whether it is exploitable at a ONE-DAY horizon on two of the most heavily traded futures on earth is a different question, and R1 established that public price structure on these instruments carries +0.02 to +0.04R gross against a +0.25R requirement. I expect a small positive that fails to clear the frontier.

Power plan

{
 "alpha": 0.05,
 "cluster_size": 2,
 "effect": 0.07,
 "intra_cluster_r": 0.9,
 "power": 0.8,
 "test": "correlation"
}

Resolved detectable measured

read from a scored result in an archived run

from run
r1 primary
effect
-0.08900
95% CI
[-0.1345, -0.0431]
floor
0.07
resolved
2026-08-30T16:38:17+00:00
Decision the part no run can produce

Family closed on MES and MNQ at a one-day horizon. The result is a REVERSAL, not a null: compression predicts more compression, so the half of the mechanism that is a documented property of financial time series (volatility clusters) survives, and the half the strategy needed (compression is stored energy about to release) does not. This was the best-scoring family in the catalogue and the strongest prior available to the programme. Does NOT close other horizons, other instruments, or the queue -- serial order stands, and X2 proceeded next against a bar raised by this registration.

source pinned at 41a78536c4277c56…

r1 controls passed

recorded
2026-08-09T18:57:27+00:00
engine_sha
3ac8953b059df7fd75adeb5b1effaba891b0bc17-dirty
spend
$0.00
metricvalue
MES.max_run7
MES.mean_run0.805
MES.n1721
MES.primary_r-0.1058
MNQ.max_run6
MNQ.mean_run0.765
MNQ.n1721
MNQ.primary_r-0.0703
buckets.run_00.9603
buckets.run_10.9349
buckets.run_20.8379
buckets.run_3plus0.8266
n3442
primary.ci[-0.13450737620661182, -0.04311178644524402]
primary.crosses_zeroFalse
primary.d-0.08899691318194508
primary.estimate-0.08899691318194508
primary.floor0.07
primary.floor_multiples-1.2713844740277866
primary.n3442
primary.n_eff1811
primary.nameX1 PRIMARY — contraction run (prior days) -> next-day range / ATR
primary.notepositive = the mechanism holds; negative = volatility persists across days instead
primary.unitSpearman r
primary.verdictdetectable
secondary_confound.ci[0.30925935277567107, 0.3900976254762833]
secondary_confound.crosses_zeroFalse
secondary_confound.d0.35033072804508375
secondary_confound.estimate0.35033072804508375
secondary_confound.floor0.07
secondary_confound.floor_multiples5.004724686358339
secondary_confound.n3442
secondary_confound.n_eff1811
secondary_confound.nameX1 SECONDARY (CONFOUND CHECK) — opening-range ratio -> same-day post-range expansion
secondary_confound.notea POSITIVE reading is expected even if the mechanism is false: intraday volatility persists, so a quiet open means a quiet day
secondary_confound.unitSpearman r
secondary_confound.verdictdetectable
Provenance the frozen script and its output, archived before the run
calcs.json4,608 Bd7cfce3bf2c012e3…
exp_x1_compression.py6,366 Be6611d5e544eb86b…
output.txt2,459 Bfb6f7977df9c6aac…
Config
{
 "floor": 0.07,
 "instruments": [
  "MES",
  "MNQ"
 ],
 "median_window": 20,
 "outcome_primary": "day N RTH range / trailing ATR",
 "outcome_secondary": "same-day post-range expansion / ATR",
 "predictor_primary": "consecutive contracting RTH sessions ending N-1",
 "predictor_secondary": "opening-range height / trailing 20d median",
 "unit": "Spearman r"
}
Note
X1. First experiment on the completed maths engine (M1-M7): all statistics from occams.stats, calibration-gated, calcs.json written. PRIMARY uses PRIOR SESSIONS ONLY -- the same-day opening range is contaminated by intraday volatility persistence, and that was decided on design grounds before any number existed. Both directions were interpreted in the registration.

SCRIPT: experiments/X1-COMPRESSION/r1/exp_x1_compression.py
OUTPUT: experiments/X1-COMPRESSION/r1/output.txt
CALCS: experiments/X1-COMPRESSION/r1/calcs.json
X1-RECONCILEreproduces verdict #3 to 3 d.p.; instrument trusted1 run
axis
instrument-trust
search space
1
alpha
0.0
outcome
reproduces verdict #3 to 3 d.p.; instrument trusted
effect
—
interval
—

Statement

The comparator's -0.046R and the sealed verdict's +0.1R can be reconciled, and the discrepancy is explained rather than assumed.

Mechanism written before the number existed

Registered as a gate, not a bet: a lab whose two implementations disagree by 0.15R across 3,295 setups cannot evaluate anything, and cross-implementation disagreement is the defect class that voided verdict #1.

Resolved reproduces verdict #3 to 3 d.p.; instrument trusted documented

stated in prose, never computed into the register

stated in
docs/PROGRAMME-CONCLUSION.md
resolved
2026-08-30T16:52:03+00:00

No effect size, interval or floor is shown because none was ever computed into the register. The source document’s numbers are authoritative; this record is not.

Decision the part no run can produce

The audited engine (M1-M7) reproduces a sealed verdict computed by the earlier code to three decimal places, so replacing three divergent bootstrap copies with one shared implementation did not silently change what the programme had already concluded. This is what licenses re-scoring the archive rather than re-running it.

source pinned at 4580cf2667b92398…

2026-08-01-decomposition controls passed

recorded
2026-08-03T10:19:24+00:00
engine_sha
bd93c839402d1f682e3ead85ec9886570d618ebd-dirty
spend
$0.00
metricvalue
gap_decomposition.MES.delta0.129
gap_decomposition.MES.limit_with_retest-0.049
gap_decomposition.MES.never_retested_pct5.0
gap_decomposition.MES.sealed_assumed_fill0.08
gap_decomposition.MNQ.delta0.193
gap_decomposition.MNQ.limit_with_retest-0.007
gap_decomposition.MNQ.never_retested_pct4.9
gap_decomposition.MNQ.sealed_assumed_fill0.186
sealed_reproduces_verdict.MES_OOS0.107
sealed_reproduces_verdict.MES_lockbox0.141
sealed_reproduces_verdict.MNQ_OOS0.171
sealed_reproduces_verdict.matchEXACT
sealed_reproduces_verdict.verdict_v3_statedMES +0.107 OOS / +0.141 lockbox; MNQ +0.171 OOS
Config
{
 "engine": "occams/fade.py via the sealed loader",
 "k_stop": 0.2,
 "range_minutes": 15,
 "risk_usd": 175.0,
 "width_max": null
}
Note
X1 CLOSED. The sealed engine reproduces verdict v3 to three decimal places, so the engine is correct and the comparator was wrong. THE CAUSE IS THE ENTRY MECHANISM, and it is not a bookkeeping detail: fade.py books a fill AT the boundary at the failure close, which in reality requires an order already RESTING there -- a stop armed during the breakout, while price is outside the range and the boundary is a valid stop level. Protocol #4a instead places a LIMIT after the failure close. That fills 95% of the time, so it is not a missed-trade problem; it is a TIMING problem. Entering on the retest forfeits 0.129R (MES) and 0.193R (MNQ) -- the entire measured edge. CONSEQUENCE: the live paper campaign is currently running a variant with NO edge, and protocol #4a needs superseding by #4b. I identified this reading on 2026-07-31, labelled it 'most faithful to the engine', and recommended against it for operational convenience. That was the wrong call and this is the measurement of what it cost.
X2-OVERNIGHT-GAPnull1 run
axis
price (1-minute bars, owned)
search space
2
alpha
0.05
outcome
null
effect
+0.0025
interval
[-0.0433, +0.0483]

Statement

The overnight gap carries DIRECTIONAL information about the session that follows. PRIMARY: the signed gap, (RTH open minus prior RTH close) / trailing ATR, is POSITIVELY rank-correlated with the signed forward return over the first session hour, measured from the close of the 09:30-09:31 bar.

Mechanism written before the number existed

The gap prices everything that happened while the market was shut. If that repricing were complete at the bell there would be nothing left to trade; the claim is that absorption is NOT instantaneous, so a large gap leaves the market still moving toward the new level during the first hour. That is the underreaction reading, and it predicts CONTINUATION.

The competing folklore reading is overreaction -- 'gaps fill' -- which predicts the opposite sign. Both are widely believed, which is exactly why the direction is committed to here rather than after the number exists.

DIRECTIONS INTERPRETED IN ADVANCE:
  * PRIMARY POSITIVE -> incomplete absorption; the mechanism holds.
  * PRIMARY NEGATIVE -> overreaction; gaps fade rather than extend. A real finding, and the opposite strategy.
  * PRIMARY ~ZERO -> the gap is fully priced at the bell and carries no directional information. Family dead.

REFERENCE PRICE, decided by design and not by default (M2). The gap is known at 09:30 from the OPEN; the measurement starts from the CLOSE of the first minute, by which time the gap is observable and an order could have been placed. Predictor and outcome therefore share no bar. Measuring from the open would fold the first minute's reaction into both, which is the ICT-P2 defect in a new place.

CONFOUND, NAMED BEFORE THE RUN, exactly as in X1. Gap SIZE almost certainly predicts subsequent VOLATILITY -- a big gap means a busy day -- and that is volatility persistence, not directional information. It is registered as a secondary CONTROL and a large positive there is expected even if the primary is zero. In X1 the analogous control came back at five times the detectable floor while the hypothesis itself was refuted; it will not be read as evidence here either.

PRIOR: I expect approximately zero. Index futures trade nearly 23 hours, so the 'overnight' gap on MES and MNQ is a gap across a brief maintenance break rather than a genuine information vacuum -- the structural argument that makes gap strategies work in equities is much weaker here, and that is worth saying before rather than after. What keeps it in the queue is that it is structurally DIFFERENT from every family that has already died, all of which traded intraday price structure.

Power plan

{
 "alpha": 0.05,
 "cluster_size": 2,
 "effect": 0.07,
 "intra_cluster_r": 0.9,
 "power": 0.8,
 "test": "correlation"
}

Resolved null measured

read from a scored result in an archived run

from run
r1 primary
effect
+0.00249
95% CI
[-0.0433, +0.0483]
floor
0.07
resolved
2026-08-30T16:38:45+00:00
Decision the part no run can produce

Family closed. Both folklore readings die together: gaps neither run nor fill. The gap itself is real and predicts the day's VOLATILITY perfectly well, but carries no directional information -- the market prices the news at the bell and leaves the leftover motion pointing nowhere. The registered confound behaved exactly as predicted, which is what makes the null trustworthy rather than merely quiet. Taken with X1 this strengthens R1's conclusion rather than adding to it: public price structure on MES and MNQ is exhausted at this resolution. X3, the only non-public axis held, remained the last structurally different item in the queue.

source pinned at 79b06807bac76c5c…

r1 controls passed

recorded
2026-08-09T19:32:02+00:00
engine_sha
9348f0b0791065470d5b37fa5d05e8bcf45a3ffe-dirty
spend
$0.00
metricvalue
MES.median_abs_gap_atr0.2789
MES.n1740
MES.primary_r0.017
MNQ.median_abs_gap_atr0.2691
MNQ.n1740
MNQ.primary_r-0.0111
control_confound.ci[0.11811904581278164, 0.20730748898392282]
control_confound.crosses_zeroFalse
control_confound.d0.16304634511841032
control_confound.estimate0.16304634511841032
control_confound.floor0.07
control_confound.floor_multiples2.3292335016915757
control_confound.n3480
control_confound.n_eff1831
control_confound.nameX2 CONTROL (CONFOUND) — |gap| -> |forward move|
control_confound.notea large POSITIVE is expected even if the primary is zero: a big gap means a busy day. Volatility persistence, NOT directional information
control_confound.unitSpearman r
control_confound.verdictdetectable
n3480
primary.ci[-0.04332667111299732, 0.04829184237777737]
primary.crosses_zeroTrue
primary.d0.002487806296301832
primary.estimate0.002487806296301832
primary.floor0.07
primary.floor_multiples0.03554008994716903
primary.n3480
primary.n_eff1831
primary.nameX2 PRIMARY — signed gap -> signed forward return, first session hour
primary.notepositive = incomplete absorption (continuation); negative = overreaction (gaps fade)
primary.unitSpearman r
primary.verdictnull
secondary_to_close.ci[-0.04805215894118799, 0.04356645855893034]
secondary_to_close.crosses_zeroTrue
secondary_to_close.d-0.002247566717407056
secondary_to_close.estimate-0.002247566717407056
secondary_to_close.floor0.07
secondary_to_close.floor_multiples-0.03210809596295794
secondary_to_close.n3480
secondary_to_close.n_eff1831
secondary_to_close.nameX2 SECONDARY — same predictor, held to the session close
secondary_to_close.notepre-specified secondary horizon
secondary_to_close.unitSpearman r
secondary_to_close.verdictnull
Provenance the frozen script and its output, archived before the run
calcs.json9,731 Bd5886e2b7f1786b6…
exp_x2_gap.py6,049 B2583968fd02ff766…
output.txt3,481 B89bd866b03af969f…
Config
{
 "control": "|gap| -> |forward move|",
 "floor": 0.07,
 "horizon_bars": 60,
 "instruments": [
  "MES",
  "MNQ"
 ],
 "outcome": "signed forward return, 60 bars from the close of bar 0",
 "predictor": "(RTH open - prior RTH close) / trailing ATR, signed",
 "reference": "AT_PRICE \u2014 predictor and outcome share no bar"
}
Note
X2. Serial after X1 (closed DEAD). Strategy-free outcome. Both directions and the volatility-persistence confound were interpreted in the registration before any number existed.

SCRIPT: experiments/X2-OVERNIGHT-GAP/r1/exp_x2_gap.py
OUTPUT: experiments/X2-OVERNIGHT-GAP/r1/output.txt
CALCS: experiments/X2-OVERNIGHT-GAP/r1/calcs.json
X3-PRINT-SIZEnull1 run
axis
order flow (trades tape, 525 sessions x2, owned)
search space
2
alpha
0.05
outcome
null
effect
+0.0260
interval
[-0.0610, +0.1125]

Statement

The SIZE DISTRIBUTION of aggressive prints carries directional information. PRIMARY: the share of 09:45-10:00 volume transacted in prints at or above that window's 90th-percentile print size is POSITIVELY rank-correlated with the window-direction-signed forward return over the following hour, measured from the 10:00 bar close.

Mechanism written before the number existed

Aggressor IDENTITY is invisible in the tape. Aggressor SIZE is not. A move carried by a few large prints is a different event from one carried by many small ones: the first is consistent with a participant who has a reason and more to do, the second with retail flow and market-making churn. If size carries any signal about who is behind a move, a move dominated by large prints should extend further.

This is the ONLY NON-PUBLIC AXIS this programme holds. Every family that has died -- ORB, the fade, second push, ICT, X1, X2 -- traded price structure visible to everyone with a chart. That is why X3 stayed in the queue after two price families died.

MULTIPLICITY, STATED PLAINLY: H5 already spent alpha on this same tape, asking whether aggressor IMBALANCE (buy vs sell volume) predicts direction. That found r = -0.072, underpowered, unconfirmed. X3 is a DIFFERENT QUESTION OF THE SAME DATA -- distribution rather than balance -- and the register must show both so the multiplicity is visible rather than inferred.

THE POWER LIMITATION, AND IT IS SEVERE. We own 525 sessions per instrument, so at most 1,050 raw observations and 552 independent-equivalent after the MES/MNQ cluster discount. The 0.07 floor used in X1 and X2 is IMPOSSIBLE here and the power gate would rightly refuse it. The declared effect is therefore 0.13, which this sample can see with margin. The consequence must be stated BEFORE the run and not discovered after: **X3 can only refute a LARGE effect. A null here means 'no effect big enough for 552 effective observations to see', which is a materially weaker statement than X1's or X2's nulls.** Declaring a bigger effect because the sample is small is only honest if that consequence is carried with it.

DIRECTIONS INTERPRETED IN ADVANCE:
  * POSITIVE -> large prints mark informed flow that extends. The mechanism holds and it is the first thing in this programme that has.
  * NEGATIVE -> large prints mark exhaustion or absorption; a move carried by size is more likely to reverse. A real finding, opposite strategy.
  * ~ZERO -> print size carries no directional information at a scale this sample can detect.

CONFOUND, NAMED BEFORE THE RUN -- the third time, and by now the expected pattern. Large-print share almost certainly predicts subsequent VOLATILITY: a chunky tape is a busy tape. It is registered as a control and a large positive there is expected even if the primary is zero. X1's analogous control returned 5.00x the floor and X2's 2.33x, while both hypotheses were refuted.

SEPARATION OF PREDICTOR AND OUTCOME. Predictor is computed from 09:45-10:00 only; the outcome starts at the 10:00 close. They share no bar and no print. Momentum, if it exists in this window, shifts the MEAN signed return but cannot create a correlation with print-size share -- so the primary is clean of it by construction.

PRIOR: the most uncertain of the queue, which is why it was kept. H5's first look at this tape was the only suggestive result this programme has produced. I expect a small positive that fails to clear 0.13.

Power plan

{
 "alpha": 0.05,
 "cluster_size": 2,
 "effect": 0.13,
 "intra_cluster_r": 0.9,
 "power": 0.8,
 "test": "correlation"
}

Resolved null measured

read from a scored result in an archived run

from run
r1 primary
effect
+0.02596
95% CI
[-0.0610, +0.1125]
floor
0.13
resolved
2026-08-31T14:31:11+00:00
Decision the part no run can produce

UNDERPOWERED NULL, CLOSED UNRESOLVED RATHER THAN LEFT OPEN. The 2x2 verdict is null -- estimate +0.026 with an interval [-0.061, +0.113] that crosses zero and sits inside a 0.13 floor (0.111 under the measured rho). But the writeup was right to call the null WEAK: the upper bound admits an effect LARGER THAN ANYTHING SEVEN YEARS OF PRICE DATA PRODUCED, on the one non-public information axis this programme holds. That is inconclusive, not answered, and the distinction has been load-bearing throughout: an ambiguous null costs the same alpha as a real test and then tempts a second look at a larger sample, which is optional stopping. WHAT WOULD RESOLVE IT, computed: 3.14x the current effective sample, which is 1,124 additional sessions per instrument at $186-291 against $21.91 of headroom -- requiring the data cap to rise to ~$315-420 with the free credit already exhausted. IT IS BEING CLOSED WITHOUT THAT, because the programme is ending. The base rates make the funded-account route a fee business (~5.6% of attempts end paid; fees are 70-95% of provider revenue), so buying power to close one open question inside a route being abandoned would be spending real money on tidiness. Recorded as UNRESOLVED-AND-CLOSED so the register never reads as though this was answered.

source pinned at 611eece5cb87fbb5…

r1 controls passed

recorded
2026-08-09T19:46:13+00:00
engine_sha
7af786f38391564a58cc7eea2b2f9806b97eb50d-dirty
spend
$0.00
metricvalue
MES.n481
MES.p90_median7.0
MES.primary_r0.0262
MES.share_median0.4717
MNQ.n489
MNQ.p90_median4.0
MNQ.primary_r-0.0024
MNQ.share_median0.3779
control_confound.ci[-0.13871547935422063, 0.034460846853245214]
control_confound.crosses_zeroTrue
control_confound.d-0.05252218292358279
control_confound.estimate-0.05252218292358279
control_confound.floor0.13
control_confound.floor_multiples-0.4040167917198676
control_confound.n970
control_confound.n_eff510
control_confound.nameX3 CONTROL (CONFOUND) — large-print share -> |forward move|
control_confound.notea large POSITIVE is expected even if the primary is zero: a chunky tape is a busy tape. Third experiment running, third such control
control_confound.unitSpearman r
control_confound.verdictnull
n970
primary.ci[-0.060998801297726744, 0.11253676859209714]
primary.crosses_zeroTrue
primary.d0.02596459281848421
primary.estimate0.02596459281848421
primary.floor0.13
primary.floor_multiples0.19972763706526314
primary.n970
primary.n_eff510
primary.nameX3 PRIMARY — large-print share -> signed forward return, next hour
primary.notepositive = large prints mark flow that extends; negative = they mark exhaustion/absorption
primary.unitSpearman r
primary.verdictnull
unconditional_mean_fwd0.00162
Provenance the frozen script and its output, archived before the run
calcs.json8,999 Baa4bdb3fc3052058…
exp_x3_print_size.py7,410 B28888dee6945943b…
output.txt2,550 Bd678fa93c1746bf3…
Config
{
 "control": "share -> |forward move|",
 "direction": "sign of the 09:45->10:00 price move",
 "floor": 0.13,
 "instruments": [
  "MES",
  "MNQ"
 ],
 "outcome": "direction-signed forward return, 60 bars from the 10:00 close",
 "predictor": "share of 09:45-10:00 volume in prints >= window p90 size",
 "reference": "AT_PRICE \u2014 predictor and outcome share no bar or print",
 "sessions_owned": 525
}
Note
X3. Serial after X1 and X2, both DEAD. The only non-public axis we hold. POWER IS THE BINDING CONSTRAINT and the floor is 0.13, not 0.07 — a null here means 'no effect big enough for 552 effective observations to see', which is materially weaker than X1/X2. Stated in the registration before the run. H5 spent alpha on this same tape asking a different question; the multiplicity is on the register.

SCRIPT: experiments/X3-PRINT-SIZE/r1/exp_x3_print_size.py
OUTPUT: experiments/X3-PRINT-SIZE/r1/output.txt
CALCS: experiments/X3-PRINT-SIZE/r1/calcs.json
Z-ENTRY-IMPLEMENTABLEassumed +0.080/+0.186 R; all four placeable orders negative2 runs
axis
instrument-trust
search space
3
alpha
0.0
outcome
assumed +0.080/+0.186 R; all four placeable orders negative
effect
—
interval
—

Statement

The sealed engine's entry -- a fill AT the range boundary at the failure close -- is achievable by some real order, and therefore the measured +0.1R is tradeable.

Mechanism written before the number existed

Registered as a verification, not a bet. #4a's failure showed the entry mechanism is worth ~0.13R, so before amending the protocol again the question had to be asked directly: which actual order produces the engine's fill?

Resolved assumed +0.080/+0.186 R; all four placeable orders negative documented

stated in prose, never computed into the register

stated in
docs/PROGRAMME-CONCLUSION.md
resolved
2026-08-30T16:52:05+00:00

No effect size, interval or floor is shown because none was ever computed into the register. The source document’s numbers are authoritative; this record is not.

Decision the part no run can produce

The fade's +0.1R does not exist. The engine booked entries at a level that sat behind the market at the failure close, so no order a human could place would have filled there; every placeable order is negative on both instruments. Verdict #3's NO-GO stands and the reasoning beneath it does not. Produced the obtainability gate, which is now a precondition on every family.

source pinned at 4580cf2667b92398…

2026-08-01-sop-verified controls passed

recorded
2026-08-03T10:50:52+00:00
engine_sha
d0aa5952f073c3fd4021e0fbebb4e73f65e781e2-dirty
spend
$0.00
metricvalue
a_limit_after_close.MES-0.0559013939564053
a_limit_after_close.MNQ-0.0091290258842136
b_market_at_close.MES-0.07555058592002953
b_market_at_close.MNQ-0.012426456450683182
c_stop_armed_at_breakout.MES-0.11879829031727884
c_stop_armed_at_breakout.MNQ-0.035672385351370135
sealed engine (ASSUMED fill).MES0.08015468041603568
sealed engine (ASSUMED fill).MNQ0.18622067669172862
Provenance the frozen script and its output, archived before the run
exp_entry_obtainable.py4,933 B102c2b93d91caa29…
output.txt695 Bceab6c22348b0791…
Config
{
 "engine": "occams/fade.py via the sealed loader",
 "k_stop": 0.2,
 "risk_usd": 175.0,
 "span": "2019-2026 pooled"
}
Note
Canonical SOP-run of the Z0 finding. The earlier record predates the harness and its option (b) had no frozen script; this run is the reproducible one.

SCRIPT: experiments/Z-ENTRY-IMPLEMENTABLE/2026-08-01-sop-verified/exp_entry_obtainable.py
OUTPUT: experiments/Z-ENTRY-IMPLEMENTABLE/2026-08-01-sop-verified/output.txt

2026-08-01-three-orders controls passed

recorded
2026-08-03T10:28:40+00:00
engine_sha
f9d59a2f7b09b969497a88e83ed507d4abdf4c62-dirty
spend
$0.00
metricvalue
a_limit_after_close.MES-0.049
a_limit_after_close.MNQ-0.007
armed_stop_breakdown.failure_close_came_first.MES_R0.122
armed_stop_breakdown.failure_close_came_first.share0.397
armed_stop_breakdown.filled_before_any_failure_close.MES_R-0.277
armed_stop_breakdown.filled_before_any_failure_close.share0.603
b_market_at_close.MES-0.076
b_market_at_close.MNQ-0.012
c_stop_armed_at_breakout.MES-0.119
c_stop_armed_at_breakout.MNQ-0.036
sealed_assumed_fill.MES0.08
sealed_assumed_fill.MNQ0.186
Config
{
 "engine": "occams/fade.py via the sealed loader",
 "k_stop": 0.2,
 "risk_usd": 175.0,
 "span": "2019-2026 pooled"
}
Note
HYPOTHESIS REJECTED, and the consequence is large. EVERY implementable order is negative on both instruments; only the engine's ASSUMED fill is positive.
WHY: at the failure close price is INSIDE the range, so the boundary is a price you cannot obtain. A resting stop fills earlier -- 60.3% of armed fills occur before any failure close, and those lose -0.277R. A market order takes the close, which is worse than the boundary. A limit enters on a retest, which is late. The engine books a fill at the boundary while conditioning on a close that had not yet happened when price was there.
That is LOOK-AHEAD, and it is the entire measured edge: the armed-stop subset where the close DID come first reads +0.122R (MES) / +0.226R (MNQ), but that subset is only selectable after the fact.
IMPLICATION: verdict v3's 'edge real (+0.1R, persists OOS)' is most likely an artifact of an unobtainable entry price. Protocol #4 is validating a non-edge, and H4 would be trying to improve something that is not there. This needs a DATED ADDENDUM to the verdict -- the verdicts are append-only and are never edited.
Z04-ORB-OBTAINABLECLEAN - ORB does not share the fade's defect1 run
axis
instrument-trust
search space
1
alpha
0.0
outcome
CLEAN - ORB does not share the fade's defect
effect
—
interval
—

Statement

The ORB verdicts (#1 and #2) do NOT share the fade's defect: every entry the ORB engine books is producible by an order that could actually have been placed at that moment.

Mechanism written before the number existed

The fade booked a fill at the range boundary at the failure close, when the boundary was already behind the market. ORB arms stops OUTSIDE the range while price is inside it, which is the correct side for a stop, so the same failure should not apply. But #1 was already VOIDED for a different instrument defect, so the family's entries have never been verified and reading the code is not checking it.

Resolved CLEAN - ORB does not share the fade's defect audit

a judgement with no estimate and no floor

from run
2026-08-01-audit .
resolved
2026-08-30T16:50:09+00:00
Decision the part no run can produce

ORB is CLEARED of the defect that voided the fade. 6,774 armed orders audited across both instruments: zero unplaceable, zero fill disagreements. ORB arms its stops OUTSIDE the range while price is inside it, so the order is on the correct side when placed -- structurally unlike the fade, whose entry sat behind the market at the failure close. Sealed verdicts #1 and #2 stand as recorded.

source pinned at 5cc2a49189f78d64…

2026-08-01-audit controls passed

recorded
2026-08-03T12:43:16+00:00
engine_sha
940d036a70059f263f4deaa2836198f1f5e360a8-dirty
spend
$0.00
metricvalue
MES.days1741
MES.examples[]
MES.fill_match2766
MES.fill_mismatch0
MES.placeable3470
MES.unplaceable0
MNQ.days1741
MNQ.examples[]
MNQ.fill_match2523
MNQ.fill_mismatch0
MNQ.placeable3304
MNQ.unplaceable0
n6774
verdictCLEAN - ORB does not share the fade's defect
Provenance the frozen script and its output, archived before the run
exp_orb_obtainable.py3,862 Bc8c6152d09bc012d…
output.txt740 B51dfe5c2e31a618a…
Config
{
 "checks": [
  "placeable when armed",
  "fill matches the engine"
 ],
 "family": "ORB",
 "instruments": [
  "MES",
  "MNQ"
 ],
 "risk_usd": 175.0,
 "span": "2019-2026",
 "stop_range": 0.5,
 "target_r": 2.0
}
Note
OUTCOME: CLEAN. 6,774 armed orders across 3,482 instrument-days, ZERO unplaceable and ZERO fill disagreements. ORB does NOT share the fade's defect.
WHY it is clean, and the contrast is the useful part: ORB arms its stops OUTSIDE the opening range while price is still INSIDE it. A buy stop above the high and a sell stop below the low are on the correct side at the moment they are placed, so the fill is produced by the order rather than assumed alongside it. The fade booked a fill at the boundary AFTER price had returned inside, when the boundary was already behind the market and nothing was resting there.
The engine's gap handling also checks out: it books max(level, open) + slippage, which is exactly what a real stop produces when a bar opens through the level -- conservative, and matching the gate on all 5,289 fills tested.
CONSEQUENCE: verdicts #1 and #2 are CLEARED of this defect. Their outcomes stand as recorded -- #1 VOID on the instrument defect it was already voided for, #2 NO-GO because the family never worked. The entry-obtainability problem is specific to the fade, not systemic across the sealed record.

SCRIPT: experiments/Z04-ORB-OBTAINABLE/2026-08-01-audit/exp_orb_obtainable.py
OUTPUT: experiments/Z04-ORB-OBTAINABLE/2026-08-01-audit/output.txt