GeoStep 0.5 release simulation results
The frozen 0.5 screen completed on the source identified in
release-calibration-0.5.json. It retained 3,000 generated panels and 12,000
analysis calls with zero failed calls. This is bounded release evidence, not a
universal calibration claim.
Analytic DiD and Lift pass the regular core screen. In the unequal-strata positive-effect cell, analytic DiD is inconclusive on the coverage rule. BCa and stratified DiD remain experimental. CRT and staircase inference remain experimental; CRT and staircase power are unsupported.
Core interval results
There are 250 replicates per cell. The Monte Carlo intervals use the Wilson method. The positive intervention is five outcome units, not five per cent Lift.
Design |
Effect |
Method |
Interval |
Coverage |
95% Monte Carlo interval |
Coverage screen |
|---|---|---|---|---|---|---|
regular |
0 |
did |
analytic |
0.960 |
[0.928, 0.978] |
pass |
regular |
0 |
did |
bca |
0.936 |
[0.899, 0.960] |
inconclusive |
regular |
0 |
lift |
analytic |
0.956 |
[0.923, 0.975] |
pass |
regular |
0 |
lift |
bca |
0.932 |
[0.894, 0.957] |
inconclusive |
regular |
5 |
did |
analytic |
0.972 |
[0.943, 0.986] |
pass |
regular |
5 |
did |
bca |
0.940 |
[0.903, 0.963] |
pass |
regular |
5 |
lift |
analytic |
0.964 |
[0.933, 0.981] |
pass |
regular |
5 |
lift |
bca |
0.936 |
[0.899, 0.960] |
inconclusive |
unequal strata |
0 |
did |
analytic |
0.944 |
[0.908, 0.966] |
pass |
unequal strata |
0 |
did |
bca |
0.928 |
[0.889, 0.954] |
inconclusive |
unequal strata |
0 |
lift |
analytic |
0.952 |
[0.918, 0.972] |
pass |
unequal strata |
0 |
lift |
bca |
0.920 |
[0.880, 0.948] |
inconclusive |
unequal strata |
5 |
did |
analytic |
0.932 |
[0.894, 0.957] |
inconclusive |
unequal strata |
5 |
did |
bca |
0.880 |
[0.834, 0.915] |
inconclusive |
unequal strata |
5 |
lift |
analytic |
0.940 |
[0.903, 0.963] |
pass |
unequal strata |
5 |
lift |
bca |
0.880 |
[0.834, 0.915] |
inconclusive |
Stress findings
BCa Lift coverage was 0.848 in the unstable-baseline positive-effect cell, with a 95% Wilson interval of [0.798, 0.887]. It fails the protocol’s 0.90 minimum screen. BCa also remains inconclusive in several few-geo, unequal-variance and heterogeneous-effect cells. In the unstable-baseline positive DiD cell, analytic coverage was 0.908 with a 95% Wilson interval of [0.866, 0.938], so it is inconclusive and does not support arbitrary unstable baselines.
The few-geo sharp-null screens are inconclusive at this budget. Exact finite allocation references establish the implemented sharp-null calculation; they do not make it a test of zero average effect under arbitrary heterogeneity.
Retained evidence and limits
Frozen protocol fixes the estimands, design and decision rules.
Machine-readable summary records all 48 cells, resolved environment, source hashes and decision status.
Replicate records retain each estimate, truth, interval, p-value, input hash and failure status.
scripts/validate_release.pyreproduces the screen with seed 20260912, 250 repetitions and 999 resamples per applicable method.
The analytic-interval call uses a sharp-null randomisation p-value. The BCa call
uses its analytic Welch p-value. Compare interval coverage by interval_method
and rejection by p_value_method; the two answer different questions.
Missing outcomes, interference, non-adherence, delayed effects, heavy-tailed innovations and CRT schedules are outside this release screen. The separate parallel inference study does not replace these gates.