REV12 METHOD DECISIONS - 8 OCTOBER 2026

These are post-analysis amendments. They are not amendments to the bytes of
the frozen protocol, and must not be described as decisions fixed before data.

1. SCORING
All characters for which unicodedata.decimal returns a value are digits.
Every emitted decimal digit participates in script classification. A digit
outside the input block and ASCII Western block is third-script substitution,
including fullwidth forms. Nondecimal numeric characters (for example a
superscript) are outside the original decimal-based outcome definition. Value
accuracy still extracts/concatenates decimal values exactly as specified.
This implements the stated outcomes without redefining exact transcription.

2. ANALYSIS SETS AND SIGNS
H1: Arabic-Indic input, neutral prompt, passed Western control, categories A
or B only; sign is converted minus preserved. The all-non-Western contrast is
separate and exploratory. H2: passed controls, all three non-Western scripts;
sign is preserve instruction minus neutral. The H2 script-fidelity outcome
is evaluated among digit-emitting outputs; selection can change with prompt.
That conditional contrast must not be given a causal interpretation. All
row weights are equal within each selected analysis set. Consequently the
observed-cluster H1 contrast weights models/values by their selected rows,
not equally by model or by all original 100 values.

3. FITTING
Bernoulli-logit model y ~ condition + (1|value) + (1|model).
Statsmodels 0.15.0 mean-field variational Bayes; Gaussian priors on fixed
coefficients with SD 2 and on log random-effect SDs with SD 1. The prior
choices are inherited from rev11 defaults and now made explicit. There are
only five model levels. The variance distribution and shrinkage are therefore
sensitive modeling choices, not precisely established population properties.
All fits use zero starting means and starting posterior SD 0.6, BFGS, maxiter
1000 and gradient tolerance 1e-5. Unsuccessful or nonfinite fits stop the
analysis. All refits and warnings are retained. The conditional odds-ratio
interval is a VB credible interval, not a bootstrap confidence interval.

4. ESTIMAND AMBIGUITY AND BOOTSTRAP
The frozen text combines averaging over observed values/models with cluster-
level parametric resampling, but does not resolve whether realized random
effects are conditioned on or integrated out. Rev12 does not declare that
ambiguity resolved by a programming change.

Observed-cluster target:
  100/n sum_i [expit(beta0+beta1+b_value(i)+b_model(i))
              -expit(beta0+b_value(i)+b_model(i))].
This is estimated by plugging in posterior coefficient/random-effect means.
It is not the posterior expectation of that nonlinear functional.
Conditional bootstrap: hold fitted random effects, exposure and cluster
memberships fixed; simulate Bernoulli outcomes; refit; recompute the same
functional using the newly fitted random effects.

Population-integrated sensitivity target:
  100 E_B[expit(beta0+beta1+B)-expit(beta0+B)],
  B ~ Normal(0, sigma_value^2 + sigma_model^2).
The expectation uses 80-node Gauss-Hermite quadrature, checked against 160
nodes on the original fit. Variance components use exp(2 * posterior mean
log-SD). This is a plug-in functional, not a posterior integral. Cluster
bootstrap redraws one effect per value and per model from the fitted Gaussian
distributions, simulates observations and refits the entire model. Every
replicate recomputes the SAME integrated target. Thus a changing empirical
cluster composition is not silently used as a fixed population estimand.
This target is exploratory and is not called the frozen observed-cluster
target. Extrapolating five selected systems to random new systems remains
substantively questionable even when the computation is coherent.

No randomization distribution for prompt assignment is claimed: the prompts
are fully crossed but no random allocation/order scheme was documented.
H1 remains observational, with sparse within-model overlap and no script-
conversion causal interpretation. A raw/adjusted difference is not proof that
between-model confounding is its sole source.

5. INTERVALS AND DECISIONS
Two-sided percentile intervals use the 2.5/97.5 or 5/95 percentiles. Basic
intervals use [2*estimate-q_upper, 2*estimate-q_lower]. Basic is not BCa; a
bootstrap displacement does not prove that basic has superior coverage.
Both methods are reported. Neither is selected because of its verdict.
Association uses a two-sided 95% interval excluding zero. TOST equivalence
requires the 90% interval strictly inside [-10,+10] pp. The 90% upper bound
below -10 is a directional one-sided 95% negative-threshold test. The positive
direction is reported separately; selecting either direction post hoc is not
claimed as a two-sided 5% test. The original text did not fully specify that
threshold-testing convention, so it is explicitly an amended reporting rule.
An equivalence conclusion within a broad 10pp margin would not establish
that the effect is exactly zero or exclude a useful improvement below 10pp.

6. CALIBRATION
The nested-bootstrap diagnostic uses 40 outer datasets for each scheme and
99 inner draws, at the fitted H1 parameter. Known generating contrasts allow
coverage to be counted. Wilson intervals describe Monte Carlo uncertainty.
This small experiment is not sufficient to establish nominal coverage,
particularly near null/threshold values, under variance misspecification,
different shrinkage priors, outcome selection, or a different true model.
It is a diagnostic and must not be used to select a favorable method.
The initial diagnostic raised a conditional-coverage concern. A refinement
keeps the SAME 40 outer datasets and increases inner draws to 399. It reduces
inner-quantile noise without increasing independent outer information. Both
versions are retained. Overflow warnings during optimization, if any, are
retained even when the final fit succeeds. Every warning-bearing full-run fit
is rechecked with L-BFGS-B from the same fixed start; this is a numerical
sensitivity diagnostic, not independent scientific validation.

7. H3 AND PRECISION
Exact paired McNemar tests are used because each pair of models saw the same
100 Arabic-Indic values. Ten observed-system pairs replace the planned 15
for six systems; Holm adjustment is over those ten pairs. The paired test
choice was not named in the frozen protocol and is disclosed here. Wilson
intervals are descriptive per-model intervals. No pairwise conclusion is
drawn merely from overlap/nonoverlap of separate intervals. The original
ad hoc ICC/design effect are withdrawn. Latent-logit correlations computed
from fitted variance components are provided, with separate shared-value,
shared-model, and shared-both cases. No single effective sample size is
asserted for this crossed design and nonlinear contrast.

8. ROBUSTNESS, EXPLORATION AND PROVENANCE
The robustness subset is exactly the archived first 25 generated values,
all short strings. All GeezaPro cells fail controls and remain excluded.
Robustness is descriptive, not an independent replication. Per-model H1
intervals and empty-output intervals are explicitly exploratory and use
value-cluster resampling. Sparse per-model arms (<10 observations) are not
assigned bootstrap intervals; this is a reporting guard, not a new exclusion
from the pooled model. The original frozen file and historical corrections
are retained. No external timestamp attestation is supplied, and no claim
that hashes alone prove pre-data registration is made.

9. REV12.1 REPORTING CLARIFICATION (no method change)
Observed H1 nominal 95% percentile coverage was 55%/65%, versus basic 85%/87.5%,
for 99/399 inner draws on the same 40 outer datasets. This weighs against the
percentile-dependent threshold claim, without certifying basic. The bootstrap
mean -22.03 pp is displaced -4.78 pp from the -17.25 pp estimate. Both targets
and both intervals are retained. The common negative H1 interval sign is
distinguished from calibrated significance. Secondary H2 script inference is
interval-sensitive. See the updated draft and report for complete disclosure.
Historical records may describe earlier interpretations and are superseded.
