Quantitative market microstructure research
Closing auction price discovery
Independent research · Aug 2026
A recurring 15:50 SPX observation became a point-in-time study of constituent closing-auction imbalance, information timing, and executable ES outcomes.
The mechanism was real. The trading hypothesis was not supported.
- Primary unconditional SPX test
- 498 days
- Constituent imbalance records
- 63.4M
- ES analytical rows after exact deduplication
- 241.3M
- Minimum absolute SPY-weight coverage
- 99.64%
- Strategies promoted
- 0
Historical research. No live trading performance or SPXW option profitability is claimed.
Research path
Observation · Replication · Mechanism · Timing · Prediction · Decision
01 / Observation
A repeated pattern was not enough.
I kept seeing the same late-day picture: SPX appeared to weaken around 15:50 ET, sometimes recovering near 15:55. The practical idea was straightforward: enter before the move and exit after it.
The research question had to be stricter: does the 15:50–15:55 window have a persistent bearish bias, and is any information available early enough to act on it?
The initial trade idea motivated the study. It was not treated as evidence.
Initial hypothesis
Initial hypothesis, not a result
15:45
Potential entry
15:50
Suspected sell pressure
15:55
Possible exit or reversal
Timeline of the initial hypothesis: potential entry at 15:45, suspected sell pressure at 15:50, and a possible exit or reversal at 15:55. This is a hypothesis schematic, not a measured result.
02 / Replication
The bearish anomaly did not replicate.
An older 249-day sample showed a 55.02% bearish rate from 15:50 to 15:55, with a one-sided p-value of 0.064. It was suggestive, but it did not meet the frozen evidence standard.
In the newer 498-day primary sample, the bearish rate was 47.99%, the one-sided p-value was 0.827, and mean return was +0.675 bps.
The periods were kept separate rather than pooled.
Result: no convincing unconditional 15:50 downside effect.
That rejection changed the project. The next question was no longer “how do I trade the pattern?” It was “why does this time of day look structurally different at all?”
Unconditional 15:50–15:55 bearish rate
Older 249-day sample bearish rate 55.02 percent, one-sided p 0.064. Newer 498-day primary sample bearish rate 47.99 percent, one-sided p 0.827. A 50 percent reference is shown. Samples are not pooled.
The pattern failed. The mechanism question remained.
03 / Reframing
So the question changed from pattern to mechanism.
Simple conditions did not rescue the original hypothesis. A predefined late-day pressure setup moved in the opposite direction. A descriptive Thursday effect failed on its validation block. SPY’s own Arca closing imbalance did not separate later SPX outcomes.
Rather than keep tuning thresholds around a weak pattern, I moved closer to the market structure of the close.
Is there a real closing-auction price-discovery event around 15:50, even if it is not an unconditional bearish anomaly?
Earlier tests
- Predefined pressure setup: did not support the original downside claim.
- Weekday discovery/validation: Thursday did not hold on the validation block.
- SPY Arca proxy: did not separate later SPX outcomes.
04 / Data system
Reconstructing the close at constituent level.
Closing pressure is distributed across hundreds of constituents and multiple listing venues. A single SPY imbalance feed was not enough to represent that cross-section.
I built a point-in-time SPY constituent universe from historical filings, linked holdings to stable security identifiers and primary listing venues, preserved historical ticker and venue transitions, and combined NYSE and Nasdaq closing-auction messages into a capitalization-weighted Constituent Imbalance Signal, or CSI.
The signal used only eligible messages received between 15:50:00 and 15:50:11 ET. Outcome data remained separated until the signal protocol and coverage checks were frozen.
- 63,396,782 imbalance records acquired after historical mapping patches
- 464 of 464 normal dates passed the frozen 95% coverage gate
- 99.6438% minimum absolute SPY-weight coverage
The data system mattered because the research question was now about what information existed, when it existed, and whether a trader could have known it yet.
Research pipeline
01
Historical holdings
Three SEC snapshots, 503 names each
02
Security + venue mapping
Time-varying ticker and listing intervals
03
XNYS / XNAS imbalance
Receive-time eligible auction messages
04
Security-level SNI
Signed imbalance intensity per name
05
Weighted CSI
Coverage-gated capitalization weights
06
ES BBO reconstruction
No-future-quote bid/ask anchors
07
Timing + execution tests
Frozen T0, spread, and fees
08
Research decision
CLOSE_AUCTION_RESEARCH
Eight-stage flow: historical holdings, security and venue mapping, XNYS and XNAS imbalance messages, security-level SNI, weighted CSI, ES BBO reconstruction, timing and execution tests, and the CLOSE_AUCTION_RESEARCH decision.
05 / Price discovery
The mechanism was real. Most of the move occurred before the full signal was actionable.
Once the full constituent signal was available, it no longer offered convincing predictive power for subsequent SPX moves. High-frequency ES data, however, showed where the visible 15:50 effect was coming from.
From 15:50:00 to 15:50:12, Sell-state ES averaged −2.50 bps while Buy-state ES averaged +2.28 bps. The Sell-minus-Buy separation was −4.79 bps.
From the frozen actionable time, T0 = 15:50:12, to 15:51, the additional separation was only −1.04 bps.
82.20%
of the absolute two-segment Sell-vs-Buy contrast occurred before T0.
Timing localization, not realizable profit.
This is a timing result, not a profit claim. It says that the market was already absorbing the closing-auction information while the full constituent signal was still arriving.
A real mechanism can be visible in the data and still arrive too quickly to become a usable signal.
Sell and Buy cumulative mean ES return
Three observed anchors for Sell and Buy cumulative mean ES returns. From 15:50:00 to T0 at 15:50:12 the Sell-minus-Buy contrast is minus 4.79 basis points. From T0 to 15:51 the additional contrast is minus 1.04 basis points. 82.20 percent of the absolute two-segment contrast occurred before T0. This is timing localization, not realizable profit.
06 / Execution
A statistical relationship remained. The executable edge did not.
After T0, the remaining relationship was small but directionally aligned. Across 459 primary ES dates, Sell-state mean return from 15:50:12 to 15:51 was −0.537 bps and Buy-state mean return was +0.500 bps.
That is a statistical relationship. Execution is a different test.
Using one ES contract, the observed bid/ask, the signal direction, and a frozen $5 round-trip fee, the residual price relationship was not enough to support a fee-positive strategy.
Mean ES midquote return after T0
After T0, Sell-state mean ES return is −0.537 basis points, N=196. Buy-state mean return is +0.500 basis points, N=263. Bootstrap intervals are shown. This is an association measure, not executable P&L.
Mean executable P&L
+$2.37
pre-fee
−$2.63
after $5 round-trip cost
Total after $5: −$1,207.50·Win rate: 47.06%
Statistically detectable does not imply economically executable.
07 / Lead signal
Could the move be anticipated before 15:50?
If most of the response had already happened by T0, the natural next question was whether earlier market behavior anticipated the closing imbalance.
It did, to a point.
Using the same frozen pre-close ES price-path feature, its association with the eventual Combined CSI rose as 15:50 approached: 0.112 → 0.224 → 0.243 → 0.279 from 15:40 to 15:49.
But at 15:49, the same feature’s Spearman correlation with the subsequent 15:50–15:51 ES return was only 0.013.
The market could anticipate part of the imbalance without providing a reliable later price direction.
Twenty-five registered pre-15:50 strategies and models were evaluated. None passed the full promotion gates.
Spearman rho versus Combined CSI and later ES return
- Future Combined CSI · ρ = 0.279
- Later ES return · ρ = 0.013
Association between a frozen pre-close ES path feature and eventual Combined CSI rises from 0.112 at 15:40 to 0.279 at 15:49. The same feature's correlation with the subsequent 15:50 to 15:51 ES return is 0.013 at 15:49.
08 / Surprise
What if only the unexpected imbalance matters?
If part of the closing imbalance was already anticipated before 15:50, the relevant quantity might not be Actual CSI. It might be the residual.
The leakage-safe construction is:
CSI Surprise = Actual CSI − OOF Expected CSI
Expected CSI was built out of fold with chronological training, training-only preprocessing, and model selection based only on CSI forecast error.
Expected CSI had weak out-of-fold predictive power.
OOF R² = −0.0356
Surprise localized price discovery.
ρ = 0.558 pre-T0
ρ = 0.104 after T0
The primary post-T0 bootstrap interval crossed zero.
The effect still failed the economic test.
−$11.60 per historical trade after $5
−$3,910 cumulative
- 337 historical trades under the frozen sign-aligned Surprise rule
- Win rate after $5: 44.81%
Supporting detail
- Profit factor: 0.811
The decomposition clarified when price discovery occurred. It did not create an executable result.
CSI Surprise versus ES return
CSI Surprise Spearman rho is 0.558 before T0 and 0.104 from T0 to 15:51, where the bootstrap interval crosses zero. Later horizons 15:52 and 15:55 remain weaker.
Historical Surprise strategy, cumulative ES P&L after $5
Cumulative historical one-contract ES bid/ask P&L after a $5 round-trip fee ends at −$3,910. This is not live performance.
09 / Promotion gate
No candidate was allowed to bypass the promotion criteria.
The project allowed exploration, but it did not allow attractive historical results to bypass the research protocol.
Rules were registered before evaluation. Dates were never randomly shuffled. Model selection stayed inside training windows. Executable tests crossed the observed bid/ask spread. Cost sensitivity, minimum sample sizes, outlier checks, calendar stability, and multiple-testing controls remained in place even when a candidate looked promising.
Three separate registries ended in the same decision.
Promotion attrition
Strategy Discovery Lab
24
Registered
17
OOF trades ≥ 60
3
Positive OOF mean after $5
0
Promoted
Pre-15:50 Lead Signal
25
Registered
23
Historical trades ≥ 60
6
Positive OOF mean after $5
0
All frozen gates passed
CSI Surprise
2
Frozen rules
1
Mean after $5 > 0
0
Robust enough to promote
Strategy Discovery Lab: 24 registered, 17 with OOF N at least 60, 3 with positive OOF mean after $5, 0 promoted. Pre-15:50 Lead Signal: 25 registered, 23 with historical N at least 60, 6 with positive OOF mean after $5, 0 passed all frozen gates. CSI Surprise: 2 frozen rules, 1 with positive mean after $5, 0 robust enough to promote.
0 strategies advanced to forward validation
10 / Decision
What the research established.
Supported
- The primary unconditional 15:50 bearish anomaly did not replicate.
- A point-in-time constituent closing-auction signal was associated with immediate ES price discovery.
- Most of the measured Sell-vs-Buy separation occurred before the frozen actionable timestamp.
- Pre-15:50 price behavior contained information about the eventual constituent imbalance.
- No registered strategy passed the full promotion process.
Not established
- A general predictive rule for SPX or ES.
- A live, deployable, or independently validated trading strategy.
- SPXW or 0DTE option profitability.
- The 82.20% pre-T0 share as realizable profit.
The project closed when the remaining historical evidence no longer justified another round of threshold, timing, or feature search.
Research decision
CLOSE_AUCTION_RESEARCH
The trading hypothesis was rejected.
The research question was resolved.
Technical appendix
Methods, definitions, and controls
The main page keeps the research path readable. The sections below document the point-in-time universe, signal construction, information clock, execution model, out-of-fold Surprise design, promotion gates, limitations, and selected implementation details.
Point-in-time universe and mapping+
The constituent universe was not reconstructed from a current S&P 500 list.
Historical SPY holdings came from three SEC-filed snapshots dated 2024-09-30, 2025-09-30, and 2026-03-31. Each contained 503 common-stock holdings. Schedule of Investments rows were linked one-to-one with concurrent N-PORT holdings to retain CUSIP/ISIN identifiers and historical portfolio weights.
Primary listing and ticker mappings were treated as time-varying intervals. Audited transitions included PLTR’s NYSE-to-Nasdaq move, Fiserv’s ticker/venue transition, Marsh’s ticker change, FITB’s 2026 venue move, and EchoStar’s SATS/ECHO history.
A date could produce CSI only after at least 95% of mapped portfolio weight had a valid signal state.
- 464 / 464 normal dates passed
- minimum within-mapped coverage: 99.6877%
- minimum absolute SPY-weight coverage: 99.6438%
Missing constituents were renormalized only after the coverage gate passed.
Security SNI and aggregate CSI+
For constituent i:
sign_i = −1 for Sell/Ask, +1 for Buy/Bid, and 0 for no imbalance. SNI_i = sign_i × total_imbalance_qty_i / (paired_qty_i + total_imbalance_qty_i)
The daily capitalization-weighted Constituent Imbalance Signal is:
CSI_d = Σ(valid_weight_i,d × SNI_i,d) / Σ(valid_weight_i,d)
Signal eligibility: 15:50:00 ≤ ts_recv < 15:50:11 ET. The latest eligible message per security was selected by receive time.
CSI < 0 → Sell; CSI > 0 → Buy; CSI = 0 → Neutral. These labels describe aggregate imbalance sign; they are not trade recommendations.
Information clock and T0+
The research separates event time from information availability.
- ts_recv is the information-arrival clock for signal eligibility and BBO availability.
- ts_event is retained for source-order and latency audit.
- Market rules are expressed in America/New_York and converted with DST awareness.
- SPX minute opens are minute-bar observations, not executable quotes.
The constituent signal can use messages received through the frozen [15:50:00, 15:50:11) window. The actionable timestamp was frozen as T0 = 15:50:12 ET. T0 was fixed before ES outcomes were merged and was not moved after observing returns.
ES execution model+
Instrument: one continuous front ES futures contract resolved daily through ES.v.0. Multiplier $50 per index point; tick 0.25 points; tick value $12.50.
Valid BBO requires finite positive bid/ask and sizes, bid < ask, and prices on tick. Quote anchor: latest valid BBO with ts_recv ≤ target timestamp. No future quote and no interpolation.
Long → enter at ask, exit at bid. Short → enter at bid, exit at ask. Primary holding window: 15:50:12 → 15:51:00 ET. Round-trip fee scenarios: $2.50, $5.00 primary, $7.50.
SPX minute returns and ES midquote returns measure association. Bid/ask P&L measures a historical executable approximation. They are not interchangeable.
Expected CSI and Surprise+
Expected CSI used 22 frozen pre-15:50 ES and SPY Arca features at 15:45 or 15:49. Model families: M0 historical mean, M1 ridge, M2 elastic net.
Outer folds expanded chronologically from 2025-04-01 onward. Inner chronological splits selected hyperparameters by CSI validation RMSE. Imputation, scaling, tuning, fitting, and residual-scale estimation all stayed inside training data. ES returns and strategy P&L were not used to choose the Expected CSI model.
CSI Surprise = Actual CSI − OOF Expected CSI
The primary Combined 15:49 model produced OOF R² = −0.0356, so the expectation itself was weak.
Research governance+
The project distinguished frozen/confirmatory tests within the available sample, exploratory historical research, and descriptive/diagnostic analysis. Later OOF procedures reduce leakage but do not turn reused history into untouched forward validation.
Controls included frozen rule registries, expanding date-blocked folds, no random date shuffling, minimum sample gates, bid/ask execution, fee sensitivity, block-bootstrap uncertainty where specified, Holm or BH multiple-testing adjustment, outlier and degraded-date sensitivity, calendar-block stability, and no outcome-driven timestamp or threshold changes.
A candidate could be promoted only if all applicable frozen gates passed. Passing historical gates would only justify a prospective protocol; it would not prove live profitability. No candidate qualified.
Limitations+
All results are historical. OOF evaluation is not an untouched prospective test. The strongest CSI/price relationship occurred while information was arriving and cannot be presented as a forecast known before the move.
Midquote and index returns do not include spread or fees. Historical bid/ask execution still does not model queue position, partial fills, market impact, or production latency.
ES is a futures proxy. SPX is a cash index. SPXW options add strike, expiry, convexity, implied volatility, theta, and option-specific liquidity. This completed branch has no SPXW profitability result.
Historical SPY weights were held constant between three filing snapshots. The old and newer unconditional samples differ materially and were not pooled. Final Surprise inference has 337 OOF dates, not the 498 dates used for the primary unconditional SPX test. Isolated low p-values must be read with researcher degrees of freedom in view.
Reproducibility+
This page documents architecture and selected methodology. It does not expose raw market data, credentials, proprietary vendor archives, local absolute paths, or paid-data files.
Selected code+
Receive-time-safe auction state
Only messages received inside the frozen signal window can enter CSI; late event timestamps cannot be treated as earlier tradable information.
start = pd.Timestamp(datetime.combine(day, wall_time(15, 50), tzinfo=ET)).tz_convert(UTC)
end = start + pd.Timedelta(seconds=11)
eligible = frame[
frame["auction_type"].isin(CLOSING_TYPES[dataset])
& frame["ts_recv"].ge(start)
& frame["ts_recv"].lt(end)
].copy()
latest = (
eligible.sort_values(["symbol", "ts_recv", "ts_event"], kind="stable")
.drop_duplicates("symbol", keep="last")
.copy()
)
paired = pd.to_numeric(latest["paired_qty"], errors="coerce")
total = pd.to_numeric(latest["total_imbalance_qty"], errors="coerce")
denominator = paired + total
valid = (
latest["side"].isin(VALID_SIDES)
& paired.notna() & total.notna() & paired.ge(0) & total.ge(0)
& (denominator.gt(0) | (denominator.eq(0) & latest["side"].eq("N")))
)
latest = latest[valid].copy()
latest["sign"] = latest["side"].map(SIGN).astype(float)
latest["sni"] = np.where(
denominator[valid].eq(0),
0.0,
latest["sign"] * total[valid] / denominator[valid],
)No-future-quote execution
The strategy test crosses the observed spread and never selects a quote received after the frozen anchor.
start = pd.Timestamp(f"{date} 15:50:00", tz=ET).tz_convert("UTC")
t0 = pd.Timestamp(f"{date} 15:50:12", tz=ET).tz_convert("UTC")
t1 = pd.Timestamp(f"{date} 15:51:00", tz=ET).tz_convert("UTC")
start_rows = bbo[bbo["ts_recv"] >= start]
t0_rows = bbo[bbo["ts_recv"] <= t0]
t1_rows = bbo[bbo["ts_recv"] <= t1]
if start_rows.empty or t0_rows.empty or t1_rows.empty:
raise RuntimeError(f"Missing frozen BBO state on {date}")
start_row, t0_row, t1_row = start_rows.iloc[0], t0_rows.iloc[-1], t1_rows.iloc[-1]
if t0_row["ts_recv"] > t0 or t1_row["ts_recv"] > t1:
raise AssertionError(f"Future quote selected on {date}")
side = "Sell" if csi < 0 else ("Buy" if csi > 0 else "Neutral")
if side == "Sell":
execution_points = float(t0_row["bid_px_00"] - t1_row["ask_px_00"])
elif side == "Buy":
execution_points = float(t1_row["bid_px_00"] - t0_row["ask_px_00"])Leakage-safe Surprise construction
Expected CSI is selected by out-of-fold CSI forecast error only; ES outcomes never enter model selection.
def fit_outer_model(model_id, x_train, y_train, x_test):
candidates = []
for params in parameter_grid(model_id):
inner_pred, fold_rmse = inner_predictions(model_id, params, x_train, y_train)
candidates.append((float(np.mean(fold_rmse)), params, inner_pred))
candidates.sort(key=lambda item: item[0])
inner_rmse, params, inner_pred = candidates[0]
residual = y_train[np.isfinite(inner_pred)] - inner_pred[np.isfinite(inner_pred)]
residual_sd = float(np.std(residual, ddof=1))
if not math.isfinite(residual_sd) or residual_sd <= 0:
raise RuntimeError("Invalid training-only residual scale")
fit = model_pipeline(model_id, params)
fit.fit(x_train, y_train)
prediction = fit.predict(x_test)
return prediction, residual_sd, params, inner_rmse
prediction, residual_sd, params, inner_rmse = fit_outer_model(
model_id, x_train, y_train, x_test
)
surprise = actual - predicted
surprise_z = surprise / residual_sdWhat comes next
What comes next
With the closing-auction hypothesis closed, the next research branch shifts from explaining one 15:50 mechanism to studying actual SPXW 0DTE option behavior across broader market states.
No results from that unfinished work are included here.