The market tells you how far, not which way
We asked the same 1,200 questions of four futures markets twice. Which way will price go? How far will it travel? The identical features, sessions and statistics answer the second question 175 times and the first one never.
Where this came from
This began as an attempt to find something rather than to refute something. We built a table of everything knowable about a session at or shortly after its open — the gap, the overnight range and where the open sat inside it, the previous day's structure, the first five, fifteen, thirty and sixty minutes — and tested all of it against everything that happened afterwards. Halfway through it became obvious that the interesting result was not any single relationship. It was that the whole exercise gave a completely different answer depending on which of two questions we asked, and that the difference is invisible unless you run both.
The same features, asked two ways
Every test in this study pairs one thing known early in the session with one thing that happens later. The pairing is the same in both halves of the study and so is everything else: the same four markets, the same sessions, the same rank correlation, the same number of tests, the same exposure to the problem that if you run enough tests, something will look significant by luck. Only the question at the end changes.
Direction asks whether the feature predicts the sign of the forward move — up or down. Size asks whether it predicts the range of the forward move — how far price travels, in either direction. Twelve hundred tests each.
Two rules keep the answers honest. No predictor window may overlap the outcome it is tested against, so nothing is ever correlated with a period containing itself. And everything is scaled by an average range taken from twenty sessions that exclude both today and yesterday, so the thing doing the normalising is never also one of the things being tested.
The direction row is the noise row
Under pure noise, the median test comes out at 0.67 and about 4.6% of tests clear the usual significance bar of 2. The 1,200 directional tests produced a median of 0.69 and 6.4% past the bar. That is the noise distribution with a rounding error attached.
The strongest single directional result reached 4.15. That sounds impressive until you compute what the largest of 1,200 independent noise draws is expected to be, which is 3.77. Finding one 4.15 in a search this size is not a discovery; it is the arithmetic of searching.
So we applied the test that actually decides things. A relationship counts only if it shows up with the same sign on at least three of four unrelated markets — an equity index, a tech index, crude oil and gold. Three hundred relationships could be checked that way. None of the directional ones replicated. Not a weakened version. Zero.
Size is nowhere near the noise
The identical 1,200 tests, asked about the size of the move instead of its direction, have a median of 2.98 and put 68% of results past the significance bar. Half of them clear 3. The strongest reaches 38. And 175 of the 300 shared relationships replicate across all four markets.
The strongest single relationship is the plainest one available: the range of the first fifteen minutes against the range of the next thirty. Rank correlation 0.59 on ES, 0.52 on NQ, 0.50 on gold, 0.41 on crude, with t-statistics from 16 to 35.
There is an obvious way for that to be an illusion. Both sides are divided by the same volatility measure, and if that measure is noisy, dividing two unrelated quantities by it will correlate them all on its own. So we broke it three ways. Undo the denominator entirely and correlate the raw, unnormalised ranges: the correlation goes up, to 0.64–0.67 in every market. Partial the volatility measure out of both sides: it barely moves. Sort sessions into fifths by prevailing volatility and measure inside each fifth, where the denominator is nearly constant: it holds in all twenty buckets.
This part is not a discovery, and that is the point
What the size result describes is volatility clustering, which has been documented since Mandelbrot noticed in 1963 that large changes are followed by large changes.1 Engle built the statistical machinery for it in 19822 and it has been measured intraday for decades.4 We are not claiming it. We are using it as the positive control: a search that cannot find a documented effect is broken, and a search that finds it exactly where it should be is working. That is what makes the directional null credible rather than merely disappointing.
What a size forecast is actually worth
Sort every session into fifths by the range of its first fifteen minutes, then measure the range the rest of the session delivered. The relationship is monotone in all four markets — each louder fifth of openings is followed by a wider remainder than the fifth below it — and the top fifth runs 2.15× the bottom fifth on ES, 1.83× on NQ, 1.73× on crude and 1.60× on gold.
That is a real, stable, cross-market fact about the next few hours of your trading day, and it says nothing whatsoever about which way to bet. Its uses are all defensive: how wide a stop has to be to survive normal noise, how large a position can be without the normal range taking it out, whether a target is reachable before the close, and whether to expect a session worth trading at all.
The control that decides it: do famous levels do anything?
A directional null invites an objection: you were testing the wrong things. The real information is not in gaps and opening ranges, it is at the prices traders actually watch — the prior day's high and low, the overnight high and low, the value area, the point of control.
The naive version of that test is worthless. Price reacts at every price, and levels close to the open get touched more often than levels far away, so anything you measure at a real level is contaminated by how far away it happened to be. The only question with an answer is comparative: when price arrives at a famous level, does it behave differently than when it arrives at an arbitrary price the same distance away?
So each real level is paired with a placebo — the same side of the open, at a distance borrowed from a different randomly chosen session, ten draws per level. The distribution of distances is preserved exactly. The only thing destroyed is the identity of the price. We then measure the same thing at both: on the first touch, entering at the next bar's open, did price reject by a quarter of the average daily range before continuing by the same amount?
Across 28 market-and-level combinations and 19,042 real touches, the average difference between the real level and its placebo is −1.6 percentage points. Three cells clear the significance bar, and all three point the wrong way: ES at the point of control (−4.7 points), crude at the point of control (−5.4) and crude at the overnight low (−4.7). Price rejects less at the famous level than at a random price the same distance from the open.
The backwards sign has a mechanism, and it is not mysterious. Famous levels are exactly where resting stop orders sit, because that is where everyone was told to put them. Arriving at a level therefore brings price to a pocket of liquidity that executes through it. The folklore has the sign of its own mechanism backwards.
One honest limitation: this measures whether the level is special on average at the first touch, on a fixed definition of rejection. It does not test a discretionary trader's claim that they can tell a good touch from a bad one in real time. That claim needs different data than ours — it is a question about order flow, and nothing in a one-minute bar can see absorption or resting size.
What to do with this
- Stop asking price data which way. Nineteen years, four markets, 1,200 tests, every condition we could construct from the open. If a directional edge lived in the shape of the session so far, this search was built to find it, and it found the noise distribution.
- Ask it how far, because it will tell you. The size of the coming move is forecastable from the size of the recent one, comfortably, in every market. Everything you do with that answer is risk management rather than selection: stop distance, position size, target feasibility, whether today is worth trading.
- Never test a level without a placebo. Every claim of the form "price respects X" needs a matched fake X to be worth anything. Ours reversed the conclusion completely: without the control the levels look active, and with it they underperform a randomly chosen price.
- Make replication the bar, not significance. One market produced a 4.15 on direction. Four markets produced nothing. A result that does not survive being asked of an unrelated instrument was a property of the sample, not of markets.
What we’d test next
- Whether a size forecast improves a system that already has an entry. We have shown the forecast exists; we have not shown that sizing or stop placement driven by it beats a fixed rule after costs. That is a separate study and it is the honest next one.
- Whether the level result changes with order-flow confirmation rather than price alone. Two other studies in this lab have terminated at the same requirement, which is starting to look like the boundary of what bar data can answer.
- Whether direction is predictable at all on a horizon longer than a session. Everything here is intraday. The null we measured is an intraday null, and it should not be quoted as anything wider.
Method
- Data
- 1-minute futures bars, day session only. ES 1,248 sessions, NQ 1,246 and CL 1,237, all 2021-09-30 to 2026-08-27; GC 3,549 sessions, 2010-07-08 to 2026-09-09. Contract roll days are dropped. A session needs at least 200 bars and a previous session with a complete overnight to be included.
- Features
- Twelve things knowable at the open — the gap, overnight range, overnight return, where the open sits inside the overnight range, position against the prior day's value area and point of control, prior-day close position, prior-day range, day of week, overnight volume, and the ratio of the last two sessions' ranges — plus the return, range and closing position of the first 5, 15, 30 and 60 minutes. Outcomes are forward returns and forward ranges measured from the 15, 30 and 60-minute marks out to 15, 30, 60 and 120 minutes and to the close.
- Rules that keep it honest
- A predictor measured over the first j minutes may only be tested against an outcome beginning at minute k ≥ j, so no window overlaps its own outcome. Everything is divided by a 20-session average range that excludes today and yesterday. Correlations are Spearman rank correlations, so a handful of violent days cannot create a relationship. Nothing counts as a finding unless it replicates with the same sign on at least three of the four markets.
- The level test
- Seven levels per market: point of control, value area high and low, prior-day high and low, overnight high and low. The point of control and value area are built from the prior session's volume at one-tick resolution, taking the 70% band. A level is skipped if it sits closer than 0.05 or further than 3.0 average daily ranges from the open. Rejection is a move of 0.25 average daily ranges back toward the session open before the same distance is travelled through the level, measured at the first touch and entered at the next bar's open. Each real level is matched against ten placebo draws taken from other sessions at the same distance.
- Limits
- Everything here is intraday and everything is measured on price alone. Three of the four markets cover five years; only gold covers nineteen. No result here is a trading rule: there are no costs, no fills and no positions anywhere in this study, because nothing in it is a strategy. The size result is a replication of documented work, not a discovery, and is reported as such.
References
- Mandelbrot, Benoit. “The Variation of Certain Speculative Prices.” The Journal of Business 36, no. 4 (1963): 394–419. Link
- Engle, Robert F. “Autoregressive Conditional Heteroscedasticity with Estimates of the Variance of United Kingdom Inflation.” Econometrica 50, no. 4 (1982): 987–1007. Link
- Bollerslev, Tim. “Generalized Autoregressive Conditional Heteroskedasticity.” Journal of Econometrics 31, no. 3 (1986): 307–327. Link
- Andersen, Torben G., and Tim Bollerslev. “Intraday Periodicity and Volatility Persistence in Financial Markets.” Journal of Empirical Finance 4, nos. 2–3 (1997): 115–158. Link
- Timmermann, Allan, and Clive W. J. Granger. “Efficient Market Hypothesis and Forecasting.” International Journal of Forecasting 20, no. 1 (2004): 15–27. Link
- Sullivan, Ryan, Allan Timmermann, and Halbert White. “Data-Snooping, Technical Trading Rule Performance, and the Bootstrap.” The Journal of Finance 54, no. 5 (1999): 1647–1691. Link
- Harvey, Campbell R., Yan Liu, and Heqing Zhu. “… and the Cross-Section of Expected Returns.” The Review of Financial Studies 29, no. 1 (2016): 5–68. Link