Release simulation results for 0.4.0rc2
The frozen screen supports a limited release candidate. BCa remains experimental. Stratified DiD also remains experimental because its positive-effect core cell was inconclusive under the prespecified coverage rule. No acceptance criterion was relaxed after inspecting results.
The second-candidate rerun reproduced all first-candidate summaries and numerical replicate records exactly. Targeted scale and cancellation regressions exercise the corrected randomisation-test implementation. The repair contract and retained first-candidate evidence describe the numerical boundaries of that comparison.
Method status
Method |
Status for this candidate |
Evidence boundary |
|---|---|---|
Unstratified analytic DiD |
Core screen passed |
Complete randomisation and regular 40-geo panels |
Analytic Lift |
Core screens passed |
Regular positive baselines; complete and stratified 40-geo panels |
Stratified analytic DiD |
Experimental |
One core coverage screen inconclusive |
BCa intervals |
Experimental |
Core screens inconclusive; undercoverage in several stress cells |
Sharp-null randomisation tests |
Core screens passed |
Permitted complete or stratified assignments; no average-effect claim |
CRT and staircase inference |
Experimental |
No design-specific calibration in this study |
CRT or staircase power |
Unsupported |
No implemented contract |
A passed screen is not a universal 95% coverage claim. The study’s minimum screen uses a 90% coverage floor and a 10% null-rejection ceiling with Monte Carlo uncertainty; these are release-screen tolerances, not recommended analysis levels. All analyses here requested 95% confidence. A numerical minimum of four geos for variance estimation is not evidence that four geos are adequate for a decision.
Core interval results
There are 250 replicates per cell. Monte Carlo intervals use the Wilson method. The status below assesses interval coverage alone; complete results retain bias, null rejection and the combined screen. The positive intervention is five outcome units, not five per cent Lift.
Design |
Effect |
Method |
Interval |
Coverage |
95% Monte Carlo interval |
Coverage screen |
|---|---|---|---|---|---|---|
regular |
0 |
did |
analytic |
0.960 |
[0.928, 0.978] |
pass |
regular |
0 |
did |
bca |
0.936 |
[0.899, 0.960] |
inconclusive |
regular |
0 |
lift |
analytic |
0.956 |
[0.923, 0.975] |
pass |
regular |
0 |
lift |
bca |
0.932 |
[0.894, 0.957] |
inconclusive |
regular |
5 |
did |
analytic |
0.972 |
[0.943, 0.986] |
pass |
regular |
5 |
did |
bca |
0.940 |
[0.903, 0.963] |
pass |
regular |
5 |
lift |
analytic |
0.964 |
[0.933, 0.981] |
pass |
regular |
5 |
lift |
bca |
0.936 |
[0.899, 0.960] |
inconclusive |
unequal_strata |
0 |
did |
analytic |
0.944 |
[0.908, 0.966] |
pass |
unequal_strata |
0 |
did |
bca |
0.928 |
[0.889, 0.954] |
inconclusive |
unequal_strata |
0 |
lift |
analytic |
0.952 |
[0.918, 0.972] |
pass |
unequal_strata |
0 |
lift |
bca |
0.920 |
[0.880, 0.948] |
inconclusive |
unequal_strata |
5 |
did |
analytic |
0.932 |
[0.894, 0.957] |
inconclusive |
unequal_strata |
5 |
did |
bca |
0.880 |
[0.834, 0.915] |
inconclusive |
unequal_strata |
5 |
lift |
analytic |
0.940 |
[0.903, 0.963] |
pass |
unequal_strata |
5 |
lift |
bca |
0.880 |
[0.834, 0.915] |
inconclusive |
Stress findings
BCa coverage reached 0.848 for Lift with unstable baselines and a positive intervention. The corresponding Monte Carlo upper bound is below the protocol’s 0.90 floor. Other BCa cells also provide evidence of coverage below the requested 0.95, even when the minimum-floor screen is inconclusive.
Analytic intervals generally fared better, but some small-sample, heterogeneous effect and unstable-baseline cells were inconclusive. In the unstable-baseline positive DiD cell, coverage was 0.908 and its Monte Carlo upper bound was below 0.95. Do not infer that passing on stable baselines validates arbitrary data.
The small-sample sharp-null screens were also inconclusive at this simulation budget. Exact-enumeration unit references establish the algorithm’s finite-design contract separately. They do not turn the sharp-null test into a test of zero average effect under arbitrary effect heterogeneity.
Reproduction and retained evidence
The completed study has 3,000 generated panels and 12,000 public analysis calls, with zero failed calls. The four calls on each panel share its data; they are not 12,000 independent experiments. All bias screens passed. The positive-effect rejection rates are descriptive and are not comparative performance claims.
Frozen protocol defines the estimands, grid and criteria.
Machine-readable summary records every cell, environment, source hashes and decision status.
Replicate records retain each estimate, truth, interval, p-value, input hash and status.
scripts/validate_release.pyreproduces the study with seed 20260912, 250 repetitions and 999 resamples per applicable method.
The analytic-interval call uses the sharp-null randomisation p-value; the BCa
call retains the analytic Welch p-value. Compare interval coverage by the
interval_method field and rejection by p_value_method in the replicate file.
A BCa interval and its reported p-value need not make the same decision.
Missing outcomes, interference, non-adherence, delayed effects, heavy-tailed innovations and CRT schedules are not covered. The earlier DiD study is separate, dated evidence and does not replace these release gates.