Figure 9 · Three benchmarks, one conclusion

Three different ways of measuring the same answer

Three independent benchmarks for the announcement-day effect — an energy-peer synthetic counterfactual, a head-to-head against Chevron, and a properly-specified oil-factor model. All three land inside the bounded-null range.

Bounded-null range: ±1.5% 0 Synthetic counterfactual +0.021% 10-firm pre-registered donor pool Head-to-head vs Chevron +0.04 pp ExxonMobil minus Chevron, market-model adjusted Oil-augmented two-factor CI extends to −4.37% −2.19% Market index + Brent crude oil; Patell p = 0.049 −4% −3% −2% −1% +1% +2% +3% +4% Day-0 abnormal return (percentage points)
2 of 3
Benchmarks within the bounded-null range
158.5
F-stat showing the third benchmark is misspecified
0/18
Specs significant after multiple-test correction
Sources & methodology notes
Three-benchmark Day-0 abnormal return comparisonAR0(b) = RExxonMobil,0Rbenchmark,0(b),   b ∈ {matched-pair, XLE, SPY+BNO}
Matched-pair: +0.04 pp (t=0.05, p=0.958). XLE: −0.03%. SPY+BNO: −0.07%. The Bayesian credible interval (Fig 8) brackets all three.

Each Day-0 estimate in the chart is supported by a distinct methodological tradition in the empirical-asset-pricing literature. The five-decade canonical references for that literature, plus the multiple-testing correction applied across the 18-specification battery, are listed below.

  1. Stephen J. Brown & Jerold B. Warner, Using Daily Stock Returns: The Case of Event Studies, 14 J. Fin. Econ. 3 (1985). Foundational empirical-event-study paper establishing the daily-returns research design used throughout the chart. Documents the small-sample behavior of the market-model parametric test (row 2 here); a follow-up to their earlier monthly-data work (1980).
  2. A. Craig MacKinlay, Event Studies in Economics and Finance, 35 J. Econ. Literature 13 (1997). Survey reference for the event-study research design adopted across all three rows: 240-day estimation window, Day-0 abnormal-return framework, and the parametric / nonparametric / robust-inference families distinguished in the chart’s three benchmarks.
  3. Alberto Abadie, Alexis Diamond & Jens Hainmueller, Synthetic Control Methods for Comparative Case Studies, 105 J. Am. Stat. Ass’n 493 (2010). Underlies row 1 (synthetic counterfactual estimator). Establishes the donor-pool weighting protocol that produces the +0.021% Day-0 gap visualized at the top of the chart. doi.org/10.1198/jasa.2009.ap08746.
  4. James M. Patell, Corporate Forecasts of Earnings Per Share and Stock Price Behavior: Empirical Tests, 14 J. Acct. Res. 246 (1976). Source of the parametric Patell-z test statistic flagging the oil-augmented Day-0 estimate (row 3) as raw-significant (p = 0.049). The Patell-z standardizes the abnormal return by an out-of-sample-corrected residual variance.
  5. Charles J. Corrado, A Nonparametric Test for Abnormal Security-Price Performance in Event Studies, 23 J. Fin. Econ. 385 (1989). Provides the rank-based companion to the Patell-z test reported alongside row 3 (Corrado p = 0.107). The disagreement between Patell-z (significant) and Corrado-rank (not significant) is itself a misspecification signal — resolved by the nested F-test on the BNO factor.
  6. Mark L. Mitchell & Erik Stafford, Managerial Decisions and Long-Term Stock Price Performance, 73 J. Bus. 287 (2000). Source of the wild-bootstrap inference (10,000 resamples, p = 0.057 for the oil-augmented Day-0) reported as the third test on row 3. Bootstrap inference is robust to event-clustering and cross-sectional dependence, providing a third opinion that aligns with Corrado against the misspecified Patell-z.
  7. Yoav Benjamini & Yosef Hochberg, Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing, 57 J. Royal Stat. Soc. B 289 (1995). Multiple-testing correction applied across the 18-specification battery (3 benchmarks × 6 windows = 18). BH-adjusted minimum p = 1.0; 0 of 18 specifications classify significant. Bottom-right callout card.
  8. Andrew Gelman et al., Bayesian Data Analysis ch. 2 (3d ed. 2013). Reference for the bounded-null credibility band shaded across the chart (±1.5%, the 95% credible-interval half-width derived in Figure 4 and fn. 29). The band represents the range of Day-0 effects consistent with the pre-period gap distribution.
  9. Halbert White, A Heteroskedasticity-Consistent Covariance Matrix Estimator and a Direct Test for Heteroskedasticity, 48 Econometrica 817 (1980); MacKinnon & White, 29 J. Econometrics 305 (1985) (HC3 refinement). Source of the HC3-robust standard error reported in the article’s robustness panel (p = 0.281); confirms the matched-pair row’s inference is unchanged under heteroskedasticity-robust covariance.
  10. Shane Goodwin, Read the Fine Print: What ExxonMobil’s Proxy Actually Says About Texas Redomiciliation, Columbia Law School Blue Sky Blog (May 2026); replication kit on file with the SMU Corporate Governance Initiative. Companion paper. The three Day-0 point estimates (synthetic +0.021%; matched-pair +0.04 pp; oil-augmented −2.19%) and the BNO nested F-test (R² 0.22 to 0.55; F(1,217)=158.5; p < 10⁻¹⁶) are documented in fn. 27.

Data attribution. Underlying daily adjusted closing prices sourced from S&P Capital IQ (IQ_CLOSEPRICE_ADJ feed) over the 240-day estimation window ending T−1 = March 9, 2026, plus the [−5, +5] event window through March 17. Whisker bands on rows 1 and 2 are 95% confidence intervals derived respectively from the pre-period synthetic-control gap distribution (σpre = 0.7676%) and the matched-pair standard error (0.85 pp); row 3 whiskers are the two-factor (market-index plus Brent crude oil benchmark) model 95% CI.

Source: Author’s calculations from S&P Capital IQ daily adjusted closing prices; methodology footnote in Goodwin (May 2026).