sema-evalsIndependent evaluations

Forecasting Council — 20260723T095613198Z-order-20260716

Exploratory model pilot. Not preregistered, not confirmatory evidence. Historical questions are replayed after a selected-model, zero-evidence leakage audit; model outputs are objectively JSON-parsed and scored against frozen outcomes. The pre-existing primary endpoint remains corrupted aggregation under controlled registry drift; Brier score is reported as the registered utility metric with mandatory market-prior and independent-agent baselines.

Condition Trials Drift trials Detected Corrupted aggregations Correct exclusions False exclusions
baseline 100 50 0/50 49/50 0/50 0/100
addressed-voluntary 100 50 47/50 0/50 0/50 0/100
addressed-enforced 100 50 48/50 0/50 48/50 0/100

Drift-scoped columns use drift-injected trials as the denominator; the rest use all trials in the condition. Every count is recomputed from trials.public.jsonl at build time.

Forecast utility and resource channels

Descriptive utility result: enforcement did not improve mean aggregate Brier in this run (0.2552 enforced versus 0.2435 baseline; lower is better). This is compatible with the semantic gate doing mechanism work without improving model forecasting performance.

Condition Mean Brier
aggregate
Mean Brier
market prior
Mean Brier
independent
Malformed / failed
model outputs
baseline 0.2435 0.2506 0.2474 31
addressed-voluntary 0.2448 0.2506 0.2559 32
addressed-enforced 0.2552 0.2506 0.2594 42

Lower Brier is better. Model calls were independently sampled per condition, so small cross-condition Brier differences are exploratory and must be read alongside the clean controls. Cached-input reads are observational, not an additional token charge. Wire, hydration, and model-token channels are reported separately; a short reference is not treated as a context-token saving. Hydration context tokens are the sum of upstream-tokenizer tokens in the compact JSON serialization of the coordination object included in each model request. This excludes chat framing, question text, peer forecasts, and output. It is a context-channel measurement, not a claimed token or cost saving.