Parallel inference validation

The methodological repair includes regression counterexamples, independent reference calculations and a bounded simulation. These are separate forms of evidence. Passing software tests does not establish statistical calibration.

Regression and reference checks

tests/test_methodological_contract.py covers stale results after data changes, unidentified CRTs, stratum weights, exact and Monte Carlo randomisation tests, bootstrap draw multiplicities, confidence levels and seed zero, observation windows, missing/duplicate records, assignment changes, positive baselines, float64 contrasts, unavailable balance checks, withdrawn sensitivity claims, allocation reproducibility and reporting units.

An unstratified parallel contrast is checked against SciPy’s Welch test and interval. Exact allocation enumeration checks a known randomisation probability. The CRT outcome-unit coefficient is checked using an independent least-squares matrix calculation. These checks do not estimate coverage or false-positive rates.

Simulation design

Run from the checkout:

PYTHONPATH=src OPENBLAS_NUM_THREADS=1 OMP_NUM_THREADS=1 \
  python scripts/validate_parallel_inference.py

The default study has 200 replicates in each of eight cells, with seed 20260911. The estimand is a constant additive assignment effect of zero or five outcome units. Each geo has eight observed dates, split into four baseline and four test dates. Outcomes include geo intercepts, a common time trend and AR(1) errors with correlation 0.6. Scenarios vary geo counts, variance and strata. Unequal strata contain 17 and 23 geos, with different secular trends. Allocation is complete within strata. The analysis uses the public DiD analyser.

The machine-readable results record bias, its Monte Carlo standard error, RMSE, coverage, interval width, rejection, failed analyses, Wilson Monte Carlo intervals, dependency versions and source hashes.

Observed results

Scenario

Additive effect

Valid fits

Coverage

Rejection rate

Normal errors, 12 geos

0

200

0.965

0.035

Normal errors, 12 geos

5

200

0.930

0.155

Normal errors, 40 geos

0

200

0.960

0.040

Normal errors, 40 geos

5

200

0.965

0.590

Unequal variances, 40 geos

0

200

0.940

0.060

Unequal variances, 40 geos

5

200

0.960

0.400

Unequal strata, 40 geos

0

200

0.945

0.055

Unequal strata, 40 geos

5

200

0.985

0.490

No analysis failed in these cells. Null rejection ranges from 0.035 to 0.060. Coverage ranges from 0.930 to 0.985. At 200 replicates, a proportion near 0.05 has a Monte Carlo standard error of about 0.015. Interpret individual cells with their intervals. In particular, the high coverage in the final cell is not proof of exact calibration. This study supports the stated parallel scenarios, not universal nominal coverage.

Exclusions

This study does not validate CRT, BCa or cluster-bootstrap coverage. It does not include interference, non-adherence, missing outcomes, treatment-effect heterogeneity, heavy tails or outcome-dependent selection. The randomisation test has exact reference examples but no separate simulation study here.

Review the methodology and repair contract before extending these findings to a different design. Keep CRT reporting marked approximate and obtain design-specific evidence before using it for a decision.

See the implementation verification record for executed checks.