Release simulation results for 0.4.0rc2

The frozen screen supports a limited release candidate. BCa remains experimental. Stratified DiD also remains experimental because its positive-effect core cell was inconclusive under the prespecified coverage rule. No acceptance criterion was relaxed after inspecting results.

The second-candidate rerun reproduced all first-candidate summaries and numerical replicate records exactly. Targeted scale and cancellation regressions exercise the corrected randomisation-test implementation. The repair contract and retained first-candidate evidence describe the numerical boundaries of that comparison.

Method status

Method

Status for this candidate

Evidence boundary

Unstratified analytic DiD

Core screen passed

Complete randomisation and regular 40-geo panels

Analytic Lift

Core screens passed

Regular positive baselines; complete and stratified 40-geo panels

Stratified analytic DiD

Experimental

One core coverage screen inconclusive

BCa intervals

Experimental

Core screens inconclusive; undercoverage in several stress cells

Sharp-null randomisation tests

Core screens passed

Permitted complete or stratified assignments; no average-effect claim

CRT and staircase inference

Experimental

No design-specific calibration in this study

CRT or staircase power

Unsupported

No implemented contract

A passed screen is not a universal 95% coverage claim. The study’s minimum screen uses a 90% coverage floor and a 10% null-rejection ceiling with Monte Carlo uncertainty; these are release-screen tolerances, not recommended analysis levels. All analyses here requested 95% confidence. A numerical minimum of four geos for variance estimation is not evidence that four geos are adequate for a decision.

Core interval results

There are 250 replicates per cell. Monte Carlo intervals use the Wilson method. The status below assesses interval coverage alone; complete results retain bias, null rejection and the combined screen. The positive intervention is five outcome units, not five per cent Lift.

Design

Effect

Method

Interval

Coverage

95% Monte Carlo interval

Coverage screen

regular

0

did

analytic

0.960

[0.928, 0.978]

pass

regular

0

did

bca

0.936

[0.899, 0.960]

inconclusive

regular

0

lift

analytic

0.956

[0.923, 0.975]

pass

regular

0

lift

bca

0.932

[0.894, 0.957]

inconclusive

regular

5

did

analytic

0.972

[0.943, 0.986]

pass

regular

5

did

bca

0.940

[0.903, 0.963]

pass

regular

5

lift

analytic

0.964

[0.933, 0.981]

pass

regular

5

lift

bca

0.936

[0.899, 0.960]

inconclusive

unequal_strata

0

did

analytic

0.944

[0.908, 0.966]

pass

unequal_strata

0

did

bca

0.928

[0.889, 0.954]

inconclusive

unequal_strata

0

lift

analytic

0.952

[0.918, 0.972]

pass

unequal_strata

0

lift

bca

0.920

[0.880, 0.948]

inconclusive

unequal_strata

5

did

analytic

0.932

[0.894, 0.957]

inconclusive

unequal_strata

5

did

bca

0.880

[0.834, 0.915]

inconclusive

unequal_strata

5

lift

analytic

0.940

[0.903, 0.963]

pass

unequal_strata

5

lift

bca

0.880

[0.834, 0.915]

inconclusive

Stress findings

BCa coverage reached 0.848 for Lift with unstable baselines and a positive intervention. The corresponding Monte Carlo upper bound is below the protocol’s 0.90 floor. Other BCa cells also provide evidence of coverage below the requested 0.95, even when the minimum-floor screen is inconclusive.

Analytic intervals generally fared better, but some small-sample, heterogeneous effect and unstable-baseline cells were inconclusive. In the unstable-baseline positive DiD cell, coverage was 0.908 and its Monte Carlo upper bound was below 0.95. Do not infer that passing on stable baselines validates arbitrary data.

The small-sample sharp-null screens were also inconclusive at this simulation budget. Exact-enumeration unit references establish the algorithm’s finite-design contract separately. They do not turn the sharp-null test into a test of zero average effect under arbitrary effect heterogeneity.

Reproduction and retained evidence

The completed study has 3,000 generated panels and 12,000 public analysis calls, with zero failed calls. The four calls on each panel share its data; they are not 12,000 independent experiments. All bias screens passed. The positive-effect rejection rates are descriptive and are not comparative performance claims.

  • Frozen protocol defines the estimands, grid and criteria.

  • Machine-readable summary records every cell, environment, source hashes and decision status.

  • Replicate records retain each estimate, truth, interval, p-value, input hash and status.

  • scripts/validate_release.py reproduces the study with seed 20260912, 250 repetitions and 999 resamples per applicable method.

The analytic-interval call uses the sharp-null randomisation p-value; the BCa call retains the analytic Welch p-value. Compare interval coverage by the interval_method field and rejection by p_value_method in the replicate file. A BCa interval and its reported p-value need not make the same decision.

Missing outcomes, interference, non-adherence, delayed effects, heavy-tailed innovations and CRT schedules are not covered. The earlier DiD study is separate, dated evidence and does not replace these release gates.