Forecasting Council — 20260722T080220652Z-order-20260716
Validates the forecasting-council scaffold: synthetic Polymarket-style questions, controlled per-agent registry drift, corrupted aggregation under baseline, voluntary detection, enforced exclusion, Brier baselines (market prior + independent-agent average), the no-drift false-exclusion guard, leakage-audit gate, condition pairing, and bundle/summary reproduction. Scripted-agent outcomes are a construction, not evidence about language models, and not evidence about live prediction markets (ADR 0017).
- Mode:
deterministic-harness - Semantic backend:
fixture-sha256-stable-json-v1 - Sema version:
not-connected - Canonicalization:
fixture-stable-json-v1 - Order seed:
20260716
| Condition | Trials | Drift trials | Detected | Corrupted aggregations | Correct exclusions | False exclusions |
|---|---|---|---|---|---|---|
baseline |
8 | 4 | 0/4 | 4/4 | 0/4 | 0/8 |
addressed-voluntary |
8 | 4 | 4/4 | 0/4 | 0/4 | 0/8 |
addressed-enforced |
8 | 4 | 4/4 | 0/4 | 4/4 | 0/8 |
Drift-scoped columns use drift-injected trials as the denominator;
the rest use all trials in the condition. Every count is recomputed from
trials.public.jsonl at build time.
Forecast utility and resource channels
Descriptive utility result: enforcement improved
mean aggregate Brier in this run (0.1397
enforced versus 0.1438 baseline; lower is
better). This is compatible with the semantic gate doing mechanism work without
improving model forecasting performance.
| Condition | Mean Brier aggregate |
Mean Brier market prior |
Mean Brier independent |
Malformed / failed model outputs |
|---|---|---|---|---|
baseline |
0.1438 | 0.2003 | 0.1519 | 0 |
addressed-voluntary |
0.1438 | 0.2003 | 0.1519 | 0 |
addressed-enforced |
0.1397 | 0.2003 | 0.1519 | 0 |
- Wire payload:
213,920 bytes. - Registry hydration/context:
64,512 bytes. - Provider model tokens:
0 input,0 cached-input reads(a subset of input),0 output,0 separately reported reasoning,0 total input + output + reasoning. - Provider retries/errors:
0/0; provider-reported cost:not reported.
Lower Brier is better. Model calls were independently sampled per condition, so small cross-condition Brier differences are exploratory and must be read alongside the clean controls. Cached-input reads are observational, not an additional token charge. Wire, hydration, and model-token channels are reported separately; a short reference is not treated as a context-token saving.