← Research
Predictability Controls

The market tells you how far, not which way

We asked the same 1,200 questions of four futures markets twice. Which way will price go? How far will it travel? The identical features, sessions and statistics answer the second question 175 times and the first one never.

2,400 tests, four markets, nineteen years on gold
0 of 300 directional relationships replicate across the four markets
175 of 300 relationships predicting the size of the move replicate
28 famous price levels tested against a placebo; none of them beat it

Where this came from

This began as an attempt to find something rather than to refute something. We built a table of everything knowable about a session at or shortly after its open — the gap, the overnight range and where the open sat inside it, the previous day's structure, the first five, fifteen, thirty and sixty minutes — and tested all of it against everything that happened afterwards. Halfway through it became obvious that the interesting result was not any single relationship. It was that the whole exercise gave a completely different answer depending on which of two questions we asked, and that the difference is invisible unless you run both.

The same features, asked two ways

Every test in this study pairs one thing known early in the session with one thing that happens later. The pairing is the same in both halves of the study and so is everything else: the same four markets, the same sessions, the same rank correlation, the same number of tests, the same exposure to the problem that if you run enough tests, something will look significant by luck. Only the question at the end changes.

Direction asks whether the feature predicts the sign of the forward move — up or down. Size asks whether it predicts the range of the forward move — how far price travels, in either direction. Twelve hundred tests each.

Two rules keep the answers honest. No predictor window may overlap the outcome it is tested against, so nothing is ever correlated with a period containing itself. And everything is scaled by an average range taken from twenty sessions that exclude both today and yesterday, so the thing doing the normalising is never also one of the things being tested.

The same 1,200 tests, asked two waysEvery test sorted by how strong it came out, weakest on the left. The dashed line is what pure noise produces. The direction curve sits on it; the size curve is nowhere near it.0246810the usual significance bar0%25%50%75%100%tests, sorted from weakest to strongestdirectionsizenoisestrength of each test, on the same scalethe size curve leaves the frame at 10 and runs to 38median strength0.692.98noise puts it at 0.67share past the bar6.4%68.2%noise puts it at 4.6%ES 1,248 sessions, NQ 1,246, CL 1,237, all 2021-09-30 to 2026-08-27; GC 3,549 sessions,2010-07-08 to 2026-09-09. 1-minute bars, day session only. Spearman rank correlation on eachpredictor / outcome pair, with no window allowed to overlap its own outcome.
Every test sorted from weakest to strongest. The direction curve lies on the noise curve for its entire length. The size curve leaves the chart.

The direction row is the noise row

Under pure noise, the median test comes out at 0.67 and about 4.6% of tests clear the usual significance bar of 2. The 1,200 directional tests produced a median of 0.69 and 6.4% past the bar. That is the noise distribution with a rounding error attached.

The strongest single directional result reached 4.15. That sounds impressive until you compute what the largest of 1,200 independent noise draws is expected to be, which is 3.77. Finding one 4.15 in a search this size is not a discovery; it is the arithmetic of searching.

So we applied the test that actually decides things. A relationship counts only if it shows up with the same sign on at least three of four unrelated markets — an equity index, a tech index, crude oil and gold. Three hundred relationships could be checked that way. None of the directional ones replicated. Not a weakened version. Zero.

What survived the requirement to replicateA relationship only counts if it appears with the same sign on at least three of four unrelated markets. Direction produces nothing; size produces 175 of the 300 candidates.from every test run, down to the ones that hold on four marketsTests run2,400Tests run: 2,400every predictor paired with every non-overlapping outcome window, both questions, four marketsRelationships that could be checked on all four600Relationships that could be checked on all four: 600the same pair present in every market, so replication is even possibleDIRECTION relationships that replicatednoneDIRECTION relationships that replicated: 0same sign everywhere and past |t| = 2 on at least three of the fourSIZE relationships that replicated175SIZE relationships that replicated: 175same rule, same bar, same exposure to multiple testingES 1,248 sessions, NQ 1,246, CL 1,237, all 2021-09-30 to 2026-08-27; GC 3,549 sessions,2010-07-08 to 2026-09-09. 1-minute bars, day session only. A relationship is one predictorpaired with one outcome window. Both questions draw on the same features, the same sessionsand the same number of tests, so they carry identical exposure to multiple comparisons.
The same funnel applied to both questions. Replication is where a search this size either survives or stops being interesting.

Size is nowhere near the noise

The identical 1,200 tests, asked about the size of the move instead of its direction, have a median of 2.98 and put 68% of results past the significance bar. Half of them clear 3. The strongest reaches 38. And 175 of the 300 shared relationships replicate across all four markets.

The strongest single relationship is the plainest one available: the range of the first fifteen minutes against the range of the next thirty. Rank correlation 0.59 on ES, 0.52 on NQ, 0.50 on gold, 0.41 on crude, with t-statistics from 16 to 35.

There is an obvious way for that to be an illusion. Both sides are divided by the same volatility measure, and if that measure is noisy, dividing two unrelated quantities by it will correlate them all on its own. So we broke it three ways. Undo the denominator entirely and correlate the raw, unnormalised ranges: the correlation goes up, to 0.64–0.67 in every market. Partial the volatility measure out of both sides: it barely moves. Sort sessions into fifths by prevailing volatility and measure inside each fifth, where the denominator is nearly constant: it holds in all twenty buckets.

The strongest relationship, and three ways of trying to break itThe first fifteen minutes’ range against the next thirty minutes’ range. Both sides are divided by the same volatility measure, so the correlation is re-measured with that denominator removed and inside narrow volatility buckets.0.000.250.500.75ES, as measured: rho 0.5880.59ES, denominator removed: rho 0.6720.67ES, volatility partialled out: rho 0.5740.57ESn = 1,248t = 25.7NQ, as measured: rho 0.5240.52NQ, denominator removed: rho 0.6490.65NQ, volatility partialled out: rho 0.5120.51NQn = 1,246t = 21.7GC, as measured: rho 0.5040.50GC, denominator removed: rho 0.6560.66GC, volatility partialled out: rho 0.4990.50GCn = 3,549t = 34.7CL, as measured: rho 0.4130.41CL, denominator removed: rho 0.6430.64CL, volatility partialled out: rho 0.3940.39CLn = 1,237t = 16.0rank correlation, opening range against the range that followsas measureddenominator removedvolatility partialled outES 1,248 sessions, NQ 1,246, CL 1,237, all 2021-09-30 to 2026-08-27; GC 3,549 sessions,2010-07-08 to 2026-09-09. 1-minute bars, day session only. If the shared denominator weremanufacturing the result, the middle bar would collapse. It is the largest of the three inevery market.
If the shared denominator were manufacturing this result, the middle bar would collapse. It is the tallest of the three in every market.
It holds inside every volatility bucketThe same correlation measured separately within each fifth of sessions sorted by prevailing volatility, where the shared denominator is nearly constant and cannot be doing the work.0.000.250.500.75ES, bucket 1: rho 0.489ES, bucket 2: rho 0.570ES, bucket 3: rho 0.538ES, bucket 4: rho 0.653ES, bucket 5: rho 0.617NQ, bucket 1: rho 0.358NQ, bucket 2: rho 0.447NQ, bucket 3: rho 0.546NQ, bucket 4: rho 0.593NQ, bucket 5: rho 0.596GC, bucket 1: rho 0.416GC, bucket 2: rho 0.493GC, bucket 3: rho 0.515GC, bucket 4: rho 0.540GC, bucket 5: rho 0.523CL, bucket 1: rho 0.353CL, bucket 2: rho 0.433CL, bucket 3: rho 0.403CL, bucket 4: rho 0.406CL, bucket 5: rho 0.402ES 0.62NQ 0.60GC 0.52CL 0.4012345quietest fifth of sessionsloudest fifthrank correlation within each volatility quintileES 1,248 sessions, NQ 1,246, CL 1,237, all 2021-09-30 to 2026-08-27; GC 3,549 sessions,2010-07-08 to 2026-09-09. 1-minute bars, day session only.
The same correlation inside each fifth of sessions sorted by how volatile things already were. Twenty buckets, twenty positive readings.

This part is not a discovery, and that is the point

What the size result describes is volatility clustering, which has been documented since Mandelbrot noticed in 1963 that large changes are followed by large changes.1 Engle built the statistical machinery for it in 19822 and it has been measured intraday for decades.4 We are not claiming it. We are using it as the positive control: a search that cannot find a documented effect is broken, and a search that finds it exactly where it should be is working. That is what makes the directional null credible rather than merely disappointing.

What a size forecast is actually worth

Sort every session into fifths by the range of its first fifteen minutes, then measure the range the rest of the session delivered. The relationship is monotone in all four markets — each louder fifth of openings is followed by a wider remainder than the fifth below it — and the top fifth runs 2.15× the bottom fifth on ES, 1.83× on NQ, 1.73× on crude and 1.60× on gold.

That is a real, stable, cross-market fact about the next few hours of your trading day, and it says nothing whatsoever about which way to bet. Its uses are all defensive: how wide a stop has to be to survive normal noise, how large a position can be without the normal range taking it out, whether a target is reachable before the close, and whether to expect a session worth trading at all.

A loud open means a loud day, in every marketSessions sorted into fifths by the range of their first fifteen minutes, against the range the rest of the session went on to deliver. No direction anywhere in this chart.0.00.40.81.21.6ES, opening quintile 1: rest-of-session range 0.66 ATRES, opening quintile 2: rest-of-session range 0.87 ATRES, opening quintile 3: rest-of-session range 0.98 ATRES, opening quintile 4: rest-of-session range 1.05 ATRES, opening quintile 5: rest-of-session range 1.42 ATRNQ, opening quintile 1: rest-of-session range 0.68 ATRNQ, opening quintile 2: rest-of-session range 0.90 ATRNQ, opening quintile 3: rest-of-session range 0.95 ATRNQ, opening quintile 4: rest-of-session range 1.06 ATRNQ, opening quintile 5: rest-of-session range 1.24 ATRGC, opening quintile 1: rest-of-session range 0.75 ATRGC, opening quintile 2: rest-of-session range 0.89 ATRGC, opening quintile 3: rest-of-session range 0.94 ATRGC, opening quintile 4: rest-of-session range 1.00 ATRGC, opening quintile 5: rest-of-session range 1.20 ATRCL, opening quintile 1: rest-of-session range 0.76 ATRCL, opening quintile 2: rest-of-session range 0.91 ATRCL, opening quintile 3: rest-of-session range 0.95 ATRCL, opening quintile 4: rest-of-session range 1.01 ATRCL, opening quintile 5: rest-of-session range 1.32 ATRES 2.2xNQ 1.8xGC 1.6xCL 1.7x12345quietest opening fifteen minutesloudestrange of the rest of the session, in units of average daily rangetop vs bottomES 1,248 sessions, NQ 1,246, CL 1,237, all 2021-09-30 to 2026-08-27; GC 3,549 sessions,2010-07-08 to 2026-09-09. 1-minute bars, day session only. Monotone in every market: eachlouder fifth of openings is followed by a wider rest of the session than the fifth below it.
Opening range against the range of the rest of the session. Monotone in every market on 19 years of gold and five of the others.

The control that decides it: do famous levels do anything?

A directional null invites an objection: you were testing the wrong things. The real information is not in gaps and opening ranges, it is at the prices traders actually watch — the prior day's high and low, the overnight high and low, the value area, the point of control.

The naive version of that test is worthless. Price reacts at every price, and levels close to the open get touched more often than levels far away, so anything you measure at a real level is contaminated by how far away it happened to be. The only question with an answer is comparative: when price arrives at a famous level, does it behave differently than when it arrives at an arbitrary price the same distance away?

So each real level is paired with a placebo — the same side of the open, at a distance borrowed from a different randomly chosen session, ten draws per level. The distribution of distances is preserved exactly. The only thing destroyed is the identity of the price. We then measure the same thing at both: on the first touch, entering at the next bar's open, did price reject by a quarter of the average daily range before continuing by the same amount?

Across 28 market-and-level combinations and 19,042 real touches, the average difference between the real level and its placebo is −1.6 percentage points. Three cells clear the significance bar, and all three point the wrong way: ES at the point of control (−4.7 points), crude at the point of control (−5.4) and crude at the overnight low (−4.7). Price rejects less at the famous level than at a random price the same distance from the open.

Famous price levels against a fake one the same distance awayEach point is how much more often price rejected at the real level than at a placebo borrowed from another session at the identical distance from the open. Zero means the level’s identity made no difference.-8-40+4+8percentage points more rejection than the placebothe level did less than a random pricethe level did morepoint of controlES point of control: real 44.1% vs placebo 48.8%, difference -4.7% (t -2.19), 590 touchesNQ point of control: real 46.8% vs placebo 47.8%, difference -1.0% (t -0.45), 558 touchesGC point of control: real 51.3% vs placebo 50.1%, difference +1.3% (t +0.85), 1,255 touchesCL point of control: real 44.9% vs placebo 50.3%, difference -5.4% (t -2.23), 461 touchesvalue area highES value area high: real 45.2% vs placebo 48.5%, difference -3.4% (t -1.42), 485 touchesNQ value area high: real 47.2% vs placebo 48.3%, difference -1.2% (t -0.46), 439 touchesGC value area high: real 48.2% vs placebo 49.2%, difference -1.0% (t -0.63), 1,114 touchesCL value area high: real 46.7% vs placebo 49.7%, difference -3.0% (t -1.16), 411 touchesvalue area lowES value area low: real 47.7% vs placebo 48.8%, difference -1.1% (t -0.46), 434 touchesNQ value area low: real 44.0% vs placebo 48.8%, difference -4.8% (t -1.87), 411 touchesGC value area low: real 49.4% vs placebo 49.9%, difference -0.5% (t -0.32), 1,084 touchesCL value area low: real 48.8% vs placebo 49.4%, difference -0.6% (t -0.23), 412 touchesprior-day highES prior-day high: real 46.9% vs placebo 48.1%, difference -1.2% (t -0.48), 448 touchesNQ prior-day high: real 45.3% vs placebo 48.4%, difference -3.0% (t -1.20), 430 touchesGC prior-day high: real 46.6% vs placebo 49.3%, difference -2.7% (t -1.71), 1,058 touchesCL prior-day high: real 50.5% vs placebo 50.3%, difference +0.2% (t +0.07), 386 touchesprior-day lowES prior-day low: real 50.2% vs placebo 48.3%, difference +1.9% (t +0.74), 418 touchesNQ prior-day low: real 43.2% vs placebo 48.0%, difference -4.7% (t -1.83), 407 touchesGC prior-day low: real 48.9% vs placebo 51.5%, difference -2.6% (t -1.60), 1,005 touchesCL prior-day low: real 49.5% vs placebo 49.3%, difference +0.2% (t +0.07), 366 touchesovernight highES overnight high: real 48.4% vs placebo 48.2%, difference +0.2% (t +0.11), 616 touchesNQ overnight high: real 50.2% vs placebo 48.1%, difference +2.2% (t +1.03), 623 touchesGC overnight high: real 51.3% vs placebo 49.4%, difference +1.9% (t +1.48), 1,608 touchesCL overnight high: real 47.6% vs placebo 49.0%, difference -1.3% (t -0.61), 569 touchesovernight lowES overnight low: real 47.0% vs placebo 50.7%, difference -3.8% (t -1.78), 609 touchesNQ overnight low: real 46.2% vs placebo 49.0%, difference -2.8% (t -1.34), 636 touchesGC overnight low: real 50.5% vs placebo 49.6%, difference +1.0% (t +0.77), 1,656 touchesCL overnight low: real 43.9% vs placebo 48.6%, difference -4.7% (t -2.12), 553 touchesESNQGCCL28 cells tested · 3 clear |t| = 2 · 0 of those in the direction the folklore predictsES 1,248 sessions, NQ 1,246, CL 1,237, all 2021-09-30 to 2026-08-27; GC 3,549 sessions,2010-07-08 to 2026-09-09. 1-minute bars, day session only. Rejection is a move of 0.25 ATRback toward the session open before the same distance is travelled through the level, measuredat the first touch and entered at the next bar’s open. Each real level is matched against 10placebo draws. Bars are two standard errors.
Seven levels, four markets. If these prices had special properties, this chart would lean right. It leans slightly left.

The backwards sign has a mechanism, and it is not mysterious. Famous levels are exactly where resting stop orders sit, because that is where everyone was told to put them. Arriving at a level therefore brings price to a pocket of liquidity that executes through it. The folklore has the sign of its own mechanism backwards.

One honest limitation: this measures whether the level is special on average at the first touch, on a fixed definition of rejection. It does not test a discretionary trader's claim that they can tell a good touch from a bad one in real time. That claim needs different data than ours — it is a question about order flow, and nothing in a one-minute bar can see absorption or resting size.

What to do with this

What we’d test next

  1. Whether a size forecast improves a system that already has an entry. We have shown the forecast exists; we have not shown that sizing or stop placement driven by it beats a fixed rule after costs. That is a separate study and it is the honest next one.
  2. Whether the level result changes with order-flow confirmation rather than price alone. Two other studies in this lab have terminated at the same requirement, which is starting to look like the boundary of what bar data can answer.
  3. Whether direction is predictable at all on a horizon longer than a session. Everything here is intraday. The null we measured is an intraday null, and it should not be quoted as anything wider.

Method

Data
1-minute futures bars, day session only. ES 1,248 sessions, NQ 1,246 and CL 1,237, all 2021-09-30 to 2026-08-27; GC 3,549 sessions, 2010-07-08 to 2026-09-09. Contract roll days are dropped. A session needs at least 200 bars and a previous session with a complete overnight to be included.
Features
Twelve things knowable at the open — the gap, overnight range, overnight return, where the open sits inside the overnight range, position against the prior day's value area and point of control, prior-day close position, prior-day range, day of week, overnight volume, and the ratio of the last two sessions' ranges — plus the return, range and closing position of the first 5, 15, 30 and 60 minutes. Outcomes are forward returns and forward ranges measured from the 15, 30 and 60-minute marks out to 15, 30, 60 and 120 minutes and to the close.
Rules that keep it honest
A predictor measured over the first j minutes may only be tested against an outcome beginning at minute k ≥ j, so no window overlaps its own outcome. Everything is divided by a 20-session average range that excludes today and yesterday. Correlations are Spearman rank correlations, so a handful of violent days cannot create a relationship. Nothing counts as a finding unless it replicates with the same sign on at least three of the four markets.
The level test
Seven levels per market: point of control, value area high and low, prior-day high and low, overnight high and low. The point of control and value area are built from the prior session's volume at one-tick resolution, taking the 70% band. A level is skipped if it sits closer than 0.05 or further than 3.0 average daily ranges from the open. Rejection is a move of 0.25 average daily ranges back toward the session open before the same distance is travelled through the level, measured at the first touch and entered at the next bar's open. Each real level is matched against ten placebo draws taken from other sessions at the same distance.
Limits
Everything here is intraday and everything is measured on price alone. Three of the four markets cover five years; only gold covers nineteen. No result here is a trading rule: there are no costs, no fills and no positions anywhere in this study, because nothing in it is a strategy. The size result is a replication of documented work, not a discovery, and is reported as such.

References

  1. Mandelbrot, Benoit. “The Variation of Certain Speculative Prices.” The Journal of Business 36, no. 4 (1963): 394–419. Link
  2. Engle, Robert F. “Autoregressive Conditional Heteroscedasticity with Estimates of the Variance of United Kingdom Inflation.” Econometrica 50, no. 4 (1982): 987–1007. Link
  3. Bollerslev, Tim. “Generalized Autoregressive Conditional Heteroskedasticity.” Journal of Econometrics 31, no. 3 (1986): 307–327. Link
  4. Andersen, Torben G., and Tim Bollerslev. “Intraday Periodicity and Volatility Persistence in Financial Markets.” Journal of Empirical Finance 4, nos. 2–3 (1997): 115–158. Link
  5. Timmermann, Allan, and Clive W. J. Granger. “Efficient Market Hypothesis and Forecasting.” International Journal of Forecasting 20, no. 1 (2004): 15–27. Link
  6. Sullivan, Ryan, Allan Timmermann, and Halbert White. “Data-Snooping, Technical Trading Rule Performance, and the Bootstrap.” The Journal of Finance 54, no. 5 (1999): 1647–1691. Link
  7. Harvey, Campbell R., Yan Liu, and Heqing Zhu. “… and the Cross-Section of Expected Returns.” The Review of Financial Studies 29, no. 1 (2016): 5–68. Link
Historical behaviour of futures contracts, not a strategy or a recommendation. Full disclaimer.