GeoStep 0.5 release simulation results

The frozen 0.5 screen completed on the source identified in release-calibration-0.5.json. It retained 3,000 generated panels and 12,000 analysis calls with zero failed calls. This is bounded release evidence, not a universal calibration claim.

Analytic DiD and Lift pass the regular core screen. In the unequal-strata positive-effect cell, analytic DiD is inconclusive on the coverage rule. BCa and stratified DiD remain experimental. CRT and staircase inference remain experimental; CRT and staircase power are unsupported.

Core interval results

There are 250 replicates per cell. The Monte Carlo intervals use the Wilson method. The positive intervention is five outcome units, not five per cent Lift.

Design

Effect

Method

Interval

Coverage

95% Monte Carlo interval

Coverage screen

regular

0

did

analytic

0.960

[0.928, 0.978]

pass

regular

0

did

bca

0.936

[0.899, 0.960]

inconclusive

regular

0

lift

analytic

0.956

[0.923, 0.975]

pass

regular

0

lift

bca

0.932

[0.894, 0.957]

inconclusive

regular

5

did

analytic

0.972

[0.943, 0.986]

pass

regular

5

did

bca

0.940

[0.903, 0.963]

pass

regular

5

lift

analytic

0.964

[0.933, 0.981]

pass

regular

5

lift

bca

0.936

[0.899, 0.960]

inconclusive

unequal strata

0

did

analytic

0.944

[0.908, 0.966]

pass

unequal strata

0

did

bca

0.928

[0.889, 0.954]

inconclusive

unequal strata

0

lift

analytic

0.952

[0.918, 0.972]

pass

unequal strata

0

lift

bca

0.920

[0.880, 0.948]

inconclusive

unequal strata

5

did

analytic

0.932

[0.894, 0.957]

inconclusive

unequal strata

5

did

bca

0.880

[0.834, 0.915]

inconclusive

unequal strata

5

lift

analytic

0.940

[0.903, 0.963]

pass

unequal strata

5

lift

bca

0.880

[0.834, 0.915]

inconclusive

Stress findings

BCa Lift coverage was 0.848 in the unstable-baseline positive-effect cell, with a 95% Wilson interval of [0.798, 0.887]. It fails the protocol’s 0.90 minimum screen. BCa also remains inconclusive in several few-geo, unequal-variance and heterogeneous-effect cells. In the unstable-baseline positive DiD cell, analytic coverage was 0.908 with a 95% Wilson interval of [0.866, 0.938], so it is inconclusive and does not support arbitrary unstable baselines.

The few-geo sharp-null screens are inconclusive at this budget. Exact finite allocation references establish the implemented sharp-null calculation; they do not make it a test of zero average effect under arbitrary heterogeneity.

Retained evidence and limits

  • Frozen protocol fixes the estimands, design and decision rules.

  • Machine-readable summary records all 48 cells, resolved environment, source hashes and decision status.

  • Replicate records retain each estimate, truth, interval, p-value, input hash and failure status.

  • scripts/validate_release.py reproduces the screen with seed 20260912, 250 repetitions and 999 resamples per applicable method.

The analytic-interval call uses a sharp-null randomisation p-value. The BCa call uses its analytic Welch p-value. Compare interval coverage by interval_method and rejection by p_value_method; the two answer different questions.

Missing outcomes, interference, non-adherence, delayed effects, heavy-tailed innovations and CRT schedules are outside this release screen. The separate parallel inference study does not replace these gates.