ZixuanChen
IWorkIIArchiveIIINotesIVAbout
WorkArchiveNotesAboutContact
IWork

Quantitative market microstructure research

Closing auction price discovery

Independent research · Aug 2026

A recurring 15:50 SPX observation became a point-in-time study of constituent closing-auction imbalance, information timing, and executable ES outcomes.

The mechanism was real. The trading hypothesis was not supported.
Primary unconditional SPX test
498 days
Constituent imbalance records
63.4M
ES analytical rows after exact deduplication
241.3M
Minimum absolute SPY-weight coverage
99.64%
Strategies promoted
0

Historical research. No live trading performance or SPXW option profitability is claimed.

Research pathTechnical appendixBack to Work
Research path

Observation · Replication · Mechanism · Timing · Prediction · Decision

  1. Observation
  2. Replication
  3. Data system
  4. Price discovery
  5. Execution
  6. Lead signal
  7. Surprise
  8. Promotion gate
  9. Decision
  10. Methods

01 / Observation

A repeated pattern was not enough.

I kept seeing the same late-day picture: SPX appeared to weaken around 15:50 ET, sometimes recovering near 15:55. The practical idea was straightforward: enter before the move and exit after it.

The research question had to be stricter: does the 15:50–15:55 window have a persistent bearish bias, and is any information available early enough to act on it?

The initial trade idea motivated the study. It was not treated as evidence.

Initial hypothesis

Initial hypothesis, not a result

  1. 15:45

    Potential entry

  2. 15:50

    Suspected sell pressure

  3. 15:55

    Possible exit or reversal

A schematic of the original trade idea. It is not a result.

Timeline of the initial hypothesis: potential entry at 15:45, suspected sell pressure at 15:50, and a possible exit or reversal at 15:55. This is a hypothesis schematic, not a measured result.

02 / Replication

The bearish anomaly did not replicate.

An older 249-day sample showed a 55.02% bearish rate from 15:50 to 15:55, with a one-sided p-value of 0.064. It was suggestive, but it did not meet the frozen evidence standard.

In the newer 498-day primary sample, the bearish rate was 47.99%, the one-sided p-value was 0.827, and mean return was +0.675 bps.

The periods were kept separate rather than pooled.

Result: no convincing unconditional 15:50 downside effect.

That rejection changed the project. The next question was no longer “how do I trade the pattern?” It was “why does this time of day look structurally different at all?”

Unconditional 15:50–15:55 bearish rate

50%Older sample55.02%N=249, p=0.064Newer primary47.99%N=498, p=0.82735%65%
The periods are separate samples and are not pooled.

Older 249-day sample bearish rate 55.02 percent, one-sided p 0.064. Newer 498-day primary sample bearish rate 47.99 percent, one-sided p 0.827. A 50 percent reference is shown. Samples are not pooled.

The pattern failed. The mechanism question remained.

03 / Reframing

So the question changed from pattern to mechanism.

Simple conditions did not rescue the original hypothesis. A predefined late-day pressure setup moved in the opposite direction. A descriptive Thursday effect failed on its validation block. SPY’s own Arca closing imbalance did not separate later SPX outcomes.

Rather than keep tuning thresholds around a weak pattern, I moved closer to the market structure of the close.

Is there a real closing-auction price-discovery event around 15:50, even if it is not an unconditional bearish anomaly?
Earlier tests
  • Predefined pressure setup: did not support the original downside claim.
  • Weekday discovery/validation: Thursday did not hold on the validation block.
  • SPY Arca proxy: did not separate later SPX outcomes.

04 / Data system

Reconstructing the close at constituent level.

Closing pressure is distributed across hundreds of constituents and multiple listing venues. A single SPY imbalance feed was not enough to represent that cross-section.

I built a point-in-time SPY constituent universe from historical filings, linked holdings to stable security identifiers and primary listing venues, preserved historical ticker and venue transitions, and combined NYSE and Nasdaq closing-auction messages into a capitalization-weighted Constituent Imbalance Signal, or CSI.

The signal used only eligible messages received between 15:50:00 and 15:50:11 ET. Outcome data remained separated until the signal protocol and coverage checks were frozen.

  • 63,396,782 imbalance records acquired after historical mapping patches
  • 464 of 464 normal dates passed the frozen 95% coverage gate
  • 99.6438% minimum absolute SPY-weight coverage

The data system mattered because the research question was now about what information existed, when it existed, and whether a trader could have known it yet.

Research pipeline

  1. 01

    Historical holdings

    Three SEC snapshots, 503 names each

  2. 02

    Security + venue mapping

    Time-varying ticker and listing intervals

  3. 03

    XNYS / XNAS imbalance

    Receive-time eligible auction messages

  4. 04

    Security-level SNI

    Signed imbalance intensity per name

  5. 05

    Weighted CSI

    Coverage-gated capitalization weights

  6. 06

    ES BBO reconstruction

    No-future-quote bid/ask anchors

  7. 07

    Timing + execution tests

    Frozen T0, spread, and fees

  8. 08

    Research decision

    CLOSE_AUCTION_RESEARCH

A constructed point-in-time research system, not a single CSV backtest. Completeness of the pipeline is not evidence that the hypothesis held.

Eight-stage flow: historical holdings, security and venue mapping, XNYS and XNAS imbalance messages, security-level SNI, weighted CSI, ES BBO reconstruction, timing and execution tests, and the CLOSE_AUCTION_RESEARCH decision.

05 / Price discovery

The mechanism was real. Most of the move occurred before the full signal was actionable.

Once the full constituent signal was available, it no longer offered convincing predictive power for subsequent SPX moves. High-frequency ES data, however, showed where the visible 15:50 effect was coming from.

From 15:50:00 to 15:50:12, Sell-state ES averaged −2.50 bps while Buy-state ES averaged +2.28 bps. The Sell-minus-Buy separation was −4.79 bps.

From the frozen actionable time, T0 = 15:50:12, to 15:51, the additional separation was only −1.04 bps.

82.20%

of the absolute two-segment Sell-vs-Buy contrast occurred before T0.

Timing localization, not realizable profit.

This is a timing result, not a profit claim. It says that the market was already absorbing the closing-auction information while the full constituent signal was still arriving.

A real mechanism can be visible in the data and still arrive too quickly to become a usable signal.

Sell and Buy cumulative mean ES return

82.2% of measured separationSellBuyT0 15:50:1215:50:0015:50:1215:51:00
82.20% refers to measured Sell-vs-Buy contrast before T0, not realizable profit. Anchors are 15:50:00, 15:50:12, and 15:51:00 only; segments are straight connections, not a continuous tape.

Three observed anchors for Sell and Buy cumulative mean ES returns. From 15:50:00 to T0 at 15:50:12 the Sell-minus-Buy contrast is minus 4.79 basis points. From T0 to 15:51 the additional contrast is minus 1.04 basis points. 82.20 percent of the absolute two-segment contrast occurred before T0. This is timing localization, not realizable profit.

06 / Execution

A statistical relationship remained. The executable edge did not.

After T0, the remaining relationship was small but directionally aligned. Across 459 primary ES dates, Sell-state mean return from 15:50:12 to 15:51 was −0.537 bps and Buy-state mean return was +0.500 bps.

That is a statistical relationship. Execution is a different test.

Using one ES contract, the observed bid/ask, the signal direction, and a frozen $5 round-trip fee, the residual price relationship was not enough to support a fee-positive strategy.

Mean ES midquote return after T0

Sell-0.537 bps · N=196Buy+0.500 bps · N=2630 bps
Bootstrap 95% intervals on mean ES midquote return, 15:50:12–15:51. Dollars in the economics bridge are a separate executable test and do not share this axis.

After T0, Sell-state mean ES return is −0.537 basis points, N=196. Buy-state mean return is +0.500 basis points, N=263. Bootstrap intervals are shown. This is an association measure, not executable P&L.

Mean executable P&L

+$2.37

pre-fee

→

−$2.63

after $5 round-trip cost

Total after $5: −$1,207.50·Win rate: 47.06%

Statistically detectable does not imply economically executable.

07 / Lead signal

Could the move be anticipated before 15:50?

If most of the response had already happened by T0, the natural next question was whether earlier market behavior anticipated the closing imbalance.

It did, to a point.

Using the same frozen pre-close ES price-path feature, its association with the eventual Combined CSI rose as 15:50 approached: 0.112 → 0.224 → 0.243 → 0.279 from 15:40 to 15:49.

But at 15:49, the same feature’s Spearman correlation with the subsequent 15:50–15:51 ES return was only 0.013.

The market could anticipate part of the imbalance without providing a reliable later price direction.

Twenty-five registered pre-15:50 strategies and models were evaluated. None passed the full promotion gates.

Spearman rho versus Combined CSI and later ES return

15:4015:4515:4715:49
  • Future Combined CSI · ρ = 0.279
  • Later ES return · ρ = 0.013
Same frozen feature, return from 15:30 to each checkpoint. Nasdaq CSI is omitted here; it does not change the main-page conclusion.

Association between a frozen pre-close ES path feature and eventual Combined CSI rises from 0.112 at 15:40 to 0.279 at 15:49. The same feature's correlation with the subsequent 15:50 to 15:51 ES return is 0.013 at 15:49.

08 / Surprise

What if only the unexpected imbalance matters?

If part of the closing imbalance was already anticipated before 15:50, the relevant quantity might not be Actual CSI. It might be the residual.

The leakage-safe construction is:

CSI Surprise = Actual CSI − OOF Expected CSI

Expected CSI was built out of fold with chronological training, training-only preprocessing, and model selection based only on CSI forecast error.

Expected CSI had weak out-of-fold predictive power.

OOF R² = −0.0356

Surprise localized price discovery.

ρ = 0.558 pre-T0

ρ = 0.104 after T0

The primary post-T0 bootstrap interval crossed zero.

The effect still failed the economic test.

−$11.60 per historical trade after $5

−$3,910 cumulative

  • 337 historical trades under the frozen sign-aligned Surprise rule
  • Win rate after $5: 44.81%
Supporting detail
  • Profit factor: 0.811

The decomposition clarified when price discovery occurred. It did not create an executable result.

CSI Surprise versus ES return

0.558Pre-T00.104T0→15:510.100T0→15:520.052T0→15:55
Primary post-T0 interval shown with bootstrap 95% interval. Pre-T0 association is contemporaneous with information arrival and is not a completed-signal trade.

CSI Surprise Spearman rho is 0.558 before T0 and 0.104 from T0 to 15:51, where the bootstrap interval crosses zero. Later horizons 15:52 and 15:55 remain weaker.

Historical Surprise strategy, cumulative ES P&L after $5

-$3,9102025-04-012026-08-14
Historical one-contract ES bid/ask execution under frozen rules; not live performance and not SPXW.

Cumulative historical one-contract ES bid/ask P&L after a $5 round-trip fee ends at −$3,910. This is not live performance.

09 / Promotion gate

No candidate was allowed to bypass the promotion criteria.

The project allowed exploration, but it did not allow attractive historical results to bypass the research protocol.

Rules were registered before evaluation. Dates were never randomly shuffled. Model selection stayed inside training windows. Executable tests crossed the observed bid/ask spread. Cost sensitivity, minimum sample sizes, outlier checks, calendar stability, and multiple-testing controls remained in place even when a candidate looked promising.

Three separate registries ended in the same decision.

Promotion attrition

Strategy Discovery Lab

  1. 24

    Registered

    →
  2. 17

    OOF trades ≥ 60

    →
  3. 3

    Positive OOF mean after $5

    →
  4. 0

    Promoted

Pre-15:50 Lead Signal

  1. 25

    Registered

    →
  2. 23

    Historical trades ≥ 60

    →
  3. 6

    Positive OOF mean after $5

    →
  4. 0

    All frozen gates passed

CSI Surprise

  1. 2

    Frozen rules

    →
  2. 1

    Mean after $5 > 0

    →
  3. 0

    Robust enough to promote

These registries are separate research families; their candidate counts are not direct measures of model quality.

Strategy Discovery Lab: 24 registered, 17 with OOF N at least 60, 3 with positive OOF mean after $5, 0 promoted. Pre-15:50 Lead Signal: 25 registered, 23 with historical N at least 60, 6 with positive OOF mean after $5, 0 passed all frozen gates. CSI Surprise: 2 frozen rules, 1 with positive mean after $5, 0 robust enough to promote.

0 strategies advanced to forward validation

10 / Decision

What the research established.

Supported

  • The primary unconditional 15:50 bearish anomaly did not replicate.
  • A point-in-time constituent closing-auction signal was associated with immediate ES price discovery.
  • Most of the measured Sell-vs-Buy separation occurred before the frozen actionable timestamp.
  • Pre-15:50 price behavior contained information about the eventual constituent imbalance.
  • No registered strategy passed the full promotion process.

Not established

  • A general predictive rule for SPX or ES.
  • A live, deployable, or independently validated trading strategy.
  • SPXW or 0DTE option profitability.
  • The 82.20% pre-T0 share as realizable profit.

The project closed when the remaining historical evidence no longer justified another round of threshold, timing, or feature search.

Research decision

CLOSE_AUCTION_RESEARCH

The trading hypothesis was rejected.

The research question was resolved.

Research path

  1. 01Observation
  2. 02Replication
  3. 03Data system
  4. 04Price discovery
  5. 05Execution
  6. 06Lead signal
  7. 07Surprise
  8. 08Promotion gate
  9. 09Decision
  10. 10Methods

Technical appendix

Methods, definitions, and controls

The main page keeps the research path readable. The sections below document the point-in-time universe, signal construction, information clock, execution model, out-of-fold Surprise design, promotion gates, limitations, and selected implementation details.

Point-in-time universe and mapping+

The constituent universe was not reconstructed from a current S&P 500 list.

Historical SPY holdings came from three SEC-filed snapshots dated 2024-09-30, 2025-09-30, and 2026-03-31. Each contained 503 common-stock holdings. Schedule of Investments rows were linked one-to-one with concurrent N-PORT holdings to retain CUSIP/ISIN identifiers and historical portfolio weights.

Primary listing and ticker mappings were treated as time-varying intervals. Audited transitions included PLTR’s NYSE-to-Nasdaq move, Fiserv’s ticker/venue transition, Marsh’s ticker change, FITB’s 2026 venue move, and EchoStar’s SATS/ECHO history.

A date could produce CSI only after at least 95% of mapped portfolio weight had a valid signal state.

  • 464 / 464 normal dates passed
  • minimum within-mapped coverage: 99.6877%
  • minimum absolute SPY-weight coverage: 99.6438%

Missing constituents were renormalized only after the coverage gate passed.

Security SNI and aggregate CSI+

For constituent i:

sign_i = −1 for Sell/Ask, +1 for Buy/Bid, and 0 for no imbalance. SNI_i = sign_i × total_imbalance_qty_i / (paired_qty_i + total_imbalance_qty_i)

The daily capitalization-weighted Constituent Imbalance Signal is:

CSI_d = Σ(valid_weight_i,d × SNI_i,d) / Σ(valid_weight_i,d)

Signal eligibility: 15:50:00 ≤ ts_recv < 15:50:11 ET. The latest eligible message per security was selected by receive time.

CSI < 0 → Sell; CSI > 0 → Buy; CSI = 0 → Neutral. These labels describe aggregate imbalance sign; they are not trade recommendations.

Information clock and T0+

The research separates event time from information availability.

  • ts_recv is the information-arrival clock for signal eligibility and BBO availability.
  • ts_event is retained for source-order and latency audit.
  • Market rules are expressed in America/New_York and converted with DST awareness.
  • SPX minute opens are minute-bar observations, not executable quotes.

The constituent signal can use messages received through the frozen [15:50:00, 15:50:11) window. The actionable timestamp was frozen as T0 = 15:50:12 ET. T0 was fixed before ES outcomes were merged and was not moved after observing returns.

ES execution model+

Instrument: one continuous front ES futures contract resolved daily through ES.v.0. Multiplier $50 per index point; tick 0.25 points; tick value $12.50.

Valid BBO requires finite positive bid/ask and sizes, bid < ask, and prices on tick. Quote anchor: latest valid BBO with ts_recv ≤ target timestamp. No future quote and no interpolation.

Long → enter at ask, exit at bid. Short → enter at bid, exit at ask. Primary holding window: 15:50:12 → 15:51:00 ET. Round-trip fee scenarios: $2.50, $5.00 primary, $7.50.

SPX minute returns and ES midquote returns measure association. Bid/ask P&L measures a historical executable approximation. They are not interchangeable.

Expected CSI and Surprise+

Expected CSI used 22 frozen pre-15:50 ES and SPY Arca features at 15:45 or 15:49. Model families: M0 historical mean, M1 ridge, M2 elastic net.

Outer folds expanded chronologically from 2025-04-01 onward. Inner chronological splits selected hyperparameters by CSI validation RMSE. Imputation, scaling, tuning, fitting, and residual-scale estimation all stayed inside training data. ES returns and strategy P&L were not used to choose the Expected CSI model.

CSI Surprise = Actual CSI − OOF Expected CSI

The primary Combined 15:49 model produced OOF R² = −0.0356, so the expectation itself was weak.

Research governance+

The project distinguished frozen/confirmatory tests within the available sample, exploratory historical research, and descriptive/diagnostic analysis. Later OOF procedures reduce leakage but do not turn reused history into untouched forward validation.

Controls included frozen rule registries, expanding date-blocked folds, no random date shuffling, minimum sample gates, bid/ask execution, fee sensitivity, block-bootstrap uncertainty where specified, Holm or BH multiple-testing adjustment, outlier and degraded-date sensitivity, calendar-block stability, and no outcome-driven timestamp or threshold changes.

A candidate could be promoted only if all applicable frozen gates passed. Passing historical gates would only justify a prospective protocol; it would not prove live profitability. No candidate qualified.

Limitations+

All results are historical. OOF evaluation is not an untouched prospective test. The strongest CSI/price relationship occurred while information was arriving and cannot be presented as a forecast known before the move.

Midquote and index returns do not include spread or fees. Historical bid/ask execution still does not model queue position, partial fills, market impact, or production latency.

ES is a futures proxy. SPX is a cash index. SPXW options add strike, expiry, convexity, implied volatility, theta, and option-specific liquidity. This completed branch has no SPXW profitability result.

Historical SPY weights were held constant between three filing snapshots. The old and newer unconditional samples differ materially and were not pooled. Final Surprise inference has 337 OOF dates, not the 498 dates used for the primary unconditional SPX test. Isolated low p-values must be read with researcher degrees of freedom in view.

Reproducibility+

This page documents architecture and selected methodology. It does not expose raw market data, credentials, proprietary vendor archives, local absolute paths, or paid-data files.

Selected code+

Receive-time-safe auction state

Only messages received inside the frozen signal window can enter CSI; late event timestamps cannot be treated as earlier tradable information.

start = pd.Timestamp(datetime.combine(day, wall_time(15, 50), tzinfo=ET)).tz_convert(UTC)
end = start + pd.Timedelta(seconds=11)
eligible = frame[
    frame["auction_type"].isin(CLOSING_TYPES[dataset])
    & frame["ts_recv"].ge(start)
    & frame["ts_recv"].lt(end)
].copy()
latest = (
    eligible.sort_values(["symbol", "ts_recv", "ts_event"], kind="stable")
    .drop_duplicates("symbol", keep="last")
    .copy()
)
paired = pd.to_numeric(latest["paired_qty"], errors="coerce")
total = pd.to_numeric(latest["total_imbalance_qty"], errors="coerce")
denominator = paired + total
valid = (
    latest["side"].isin(VALID_SIDES)
    & paired.notna() & total.notna() & paired.ge(0) & total.ge(0)
    & (denominator.gt(0) | (denominator.eq(0) & latest["side"].eq("N")))
)
latest = latest[valid].copy()
latest["sign"] = latest["side"].map(SIGN).astype(float)
latest["sni"] = np.where(
    denominator[valid].eq(0),
    0.0,
    latest["sign"] * total[valid] / denominator[valid],
)

No-future-quote execution

The strategy test crosses the observed spread and never selects a quote received after the frozen anchor.

start = pd.Timestamp(f"{date} 15:50:00", tz=ET).tz_convert("UTC")
t0 = pd.Timestamp(f"{date} 15:50:12", tz=ET).tz_convert("UTC")
t1 = pd.Timestamp(f"{date} 15:51:00", tz=ET).tz_convert("UTC")
start_rows = bbo[bbo["ts_recv"] >= start]
t0_rows = bbo[bbo["ts_recv"] <= t0]
t1_rows = bbo[bbo["ts_recv"] <= t1]
if start_rows.empty or t0_rows.empty or t1_rows.empty:
    raise RuntimeError(f"Missing frozen BBO state on {date}")
start_row, t0_row, t1_row = start_rows.iloc[0], t0_rows.iloc[-1], t1_rows.iloc[-1]
if t0_row["ts_recv"] > t0 or t1_row["ts_recv"] > t1:
    raise AssertionError(f"Future quote selected on {date}")

side = "Sell" if csi < 0 else ("Buy" if csi > 0 else "Neutral")
if side == "Sell":
    execution_points = float(t0_row["bid_px_00"] - t1_row["ask_px_00"])
elif side == "Buy":
    execution_points = float(t1_row["bid_px_00"] - t0_row["ask_px_00"])

Leakage-safe Surprise construction

Expected CSI is selected by out-of-fold CSI forecast error only; ES outcomes never enter model selection.

def fit_outer_model(model_id, x_train, y_train, x_test):
    candidates = []
    for params in parameter_grid(model_id):
        inner_pred, fold_rmse = inner_predictions(model_id, params, x_train, y_train)
        candidates.append((float(np.mean(fold_rmse)), params, inner_pred))
    candidates.sort(key=lambda item: item[0])
    inner_rmse, params, inner_pred = candidates[0]
    residual = y_train[np.isfinite(inner_pred)] - inner_pred[np.isfinite(inner_pred)]
    residual_sd = float(np.std(residual, ddof=1))
    if not math.isfinite(residual_sd) or residual_sd <= 0:
        raise RuntimeError("Invalid training-only residual scale")
    fit = model_pipeline(model_id, params)
    fit.fit(x_train, y_train)
    prediction = fit.predict(x_test)
    return prediction, residual_sd, params, inner_rmse

prediction, residual_sd, params, inner_rmse = fit_outer_model(
    model_id, x_train, y_train, x_test
)
surprise = actual - predicted
surprise_z = surprise / residual_sd

What comes next

What comes next

With the closing-auction hypothesis closed, the next research branch shifts from explaining one 15:50 mechanism to studying actual SPXW 0DTE option behavior across broader market states.

No results from that unfinished work are included here.

Historical research for portfolio and educational purposes. Not financial advice. Statistical association is not the same as executable alpha. No live trading performance or SPXW option profitability is claimed.

Zixuan Chen

Boston

adrian.zixuan.chen@gmail.comLinkedIn ↗

All text and images by Zixuan Chen, 2026

Index
IWorkIIArchiveIIINotesIVAboutVContact