sema-evalsIndependent evaluations

Forecasting Council — 20260722T080220652Z-order-20260716

Validates the forecasting-council scaffold: synthetic Polymarket-style questions, controlled per-agent registry drift, corrupted aggregation under baseline, voluntary detection, enforced exclusion, Brier baselines (market prior + independent-agent average), the no-drift false-exclusion guard, leakage-audit gate, condition pairing, and bundle/summary reproduction. Scripted-agent outcomes are a construction, not evidence about language models, and not evidence about live prediction markets (ADR 0017).

Condition Trials Drift trials Detected Corrupted aggregations Correct exclusions False exclusions
baseline 8 4 0/4 4/4 0/4 0/8
addressed-voluntary 8 4 4/4 0/4 0/4 0/8
addressed-enforced 8 4 4/4 0/4 4/4 0/8

Drift-scoped columns use drift-injected trials as the denominator; the rest use all trials in the condition. Every count is recomputed from trials.public.jsonl at build time.

Forecast utility and resource channels

Descriptive utility result: enforcement improved mean aggregate Brier in this run (0.1397 enforced versus 0.1438 baseline; lower is better). This is compatible with the semantic gate doing mechanism work without improving model forecasting performance.

Condition Mean Brier
aggregate
Mean Brier
market prior
Mean Brier
independent
Malformed / failed
model outputs
baseline 0.1438 0.2003 0.1519 0
addressed-voluntary 0.1438 0.2003 0.1519 0
addressed-enforced 0.1397 0.2003 0.1519 0

Lower Brier is better. Model calls were independently sampled per condition, so small cross-condition Brier differences are exploratory and must be read alongside the clean controls. Cached-input reads are observational, not an additional token charge. Wire, hydration, and model-token channels are reported separately; a short reference is not treated as a context-token saving.