Parallel inference validation
The methodological repair includes regression counterexamples, independent reference calculations and a bounded simulation. These are separate forms of evidence. Passing software tests does not establish statistical calibration.
Regression and reference checks
tests/test_methodological_contract.py covers stale results after data changes,
unidentified CRTs, stratum weights, exact and Monte Carlo randomisation tests,
bootstrap draw multiplicities, confidence levels and seed zero, observation
windows, missing/duplicate records, assignment changes, positive baselines,
float64 contrasts, unavailable balance checks, withdrawn sensitivity claims,
allocation reproducibility and reporting units.
An unstratified parallel contrast is checked against SciPy’s Welch test and interval. Exact allocation enumeration checks a known randomisation probability. The CRT outcome-unit coefficient is checked using an independent least-squares matrix calculation. These checks do not estimate coverage or false-positive rates.
Simulation design
Run from the checkout:
PYTHONPATH=src OPENBLAS_NUM_THREADS=1 OMP_NUM_THREADS=1 \
python scripts/validate_parallel_inference.py
The default study has 200 replicates in each of eight cells, with seed 20260911. The estimand is a constant additive assignment effect of zero or five outcome units. Each geo has eight observed dates, split into four baseline and four test dates. Outcomes include geo intercepts, a common time trend and AR(1) errors with correlation 0.6. Scenarios vary geo counts, variance and strata. Unequal strata contain 17 and 23 geos, with different secular trends. Allocation is complete within strata. The analysis uses the public DiD analyser.
The machine-readable results record bias, its Monte Carlo standard error, RMSE, coverage, interval width, rejection, failed analyses, Wilson Monte Carlo intervals, dependency versions and source hashes.
Observed results
Scenario |
Additive effect |
Valid fits |
Coverage |
Rejection rate |
|---|---|---|---|---|
Normal errors, 12 geos |
0 |
200 |
0.965 |
0.035 |
Normal errors, 12 geos |
5 |
200 |
0.930 |
0.155 |
Normal errors, 40 geos |
0 |
200 |
0.960 |
0.040 |
Normal errors, 40 geos |
5 |
200 |
0.965 |
0.590 |
Unequal variances, 40 geos |
0 |
200 |
0.940 |
0.060 |
Unequal variances, 40 geos |
5 |
200 |
0.960 |
0.400 |
Unequal strata, 40 geos |
0 |
200 |
0.945 |
0.055 |
Unequal strata, 40 geos |
5 |
200 |
0.985 |
0.490 |
No analysis failed in these cells. Null rejection ranges from 0.035 to 0.060. Coverage ranges from 0.930 to 0.985. At 200 replicates, a proportion near 0.05 has a Monte Carlo standard error of about 0.015. Interpret individual cells with their intervals. In particular, the high coverage in the final cell is not proof of exact calibration. This study supports the stated parallel scenarios, not universal nominal coverage.
Exclusions
This study does not validate CRT, BCa or cluster-bootstrap coverage. It does not include interference, non-adherence, missing outcomes, treatment-effect heterogeneity, heavy tails or outcome-dependent selection. The randomisation test has exact reference examples but no separate simulation study here.
Review the methodology and repair contract before extending these findings to a different design. Keep CRT reporting marked approximate and obtain design-specific evidence before using it for a decision.
See the implementation verification record for executed checks.