Forecasting Council — 20260723T095613198Z-order-20260716
Exploratory model pilot. Not preregistered, not confirmatory evidence. Historical questions are replayed after a selected-model, zero-evidence leakage audit; model outputs are objectively JSON-parsed and scored against frozen outcomes. The pre-existing primary endpoint remains corrupted aggregation under controlled registry drift; Brier score is reported as the registered utility metric with mandatory market-prior and independent-agent baselines.
- Mode:
model-pilot - Provider:
llm.chutes.ai - Model:
unsloth/Mistral-Nemo-Instruct-2407-TEE - Semantic backend:
semahash-python-api - Sema version:
0.3.0 - Canonicalization:
v2 - Order seed:
20260716
| Condition | Trials | Drift trials | Detected | Corrupted aggregations | Correct exclusions | False exclusions |
|---|---|---|---|---|---|---|
baseline |
100 | 50 | 0/50 | 49/50 | 0/50 | 0/100 |
addressed-voluntary |
100 | 50 | 47/50 | 0/50 | 0/50 | 0/100 |
addressed-enforced |
100 | 50 | 48/50 | 0/50 | 48/50 | 0/100 |
Drift-scoped columns use drift-injected trials as the denominator;
the rest use all trials in the condition. Every count is recomputed from
trials.public.jsonl at build time.
Forecast utility and resource channels
Descriptive utility result: enforcement did not improve
mean aggregate Brier in this run (0.2552
enforced versus 0.2435 baseline; lower is
better). This is compatible with the semantic gate doing mechanism work without
improving model forecasting performance.
| Condition | Mean Brier aggregate |
Mean Brier market prior |
Mean Brier independent |
Malformed / failed model outputs |
|---|---|---|---|---|
baseline |
0.2435 | 0.2506 | 0.2474 | 31 |
addressed-voluntary |
0.2448 | 0.2506 | 0.2559 | 32 |
addressed-enforced |
0.2552 | 0.2506 | 0.2594 | 42 |
- Historical source: SimpleFunctions Settled Prediction Markets
at revision
a27e3e9307266481d51e087fffd5bf934410e01c, licensedCC-BY-4.0; attribution: SimpleFunctions (simplefunctions.dev). - Wire payload:
2,505,168 bytes. - Registry hydration/context:
1,180,320 bytes. - Tokenizer-derived coordination hydration:
596,820 context tokensusingmistralai/Mistral-Nemo-Instruct-2407@04d8a90549d23fc6bd7f642064003592df51e9b3(tokenizer SHA-256e11c71726323d33da7b8d6f6f269f1988931c0a52b7122bcdd8c05042974e0db). - Provider model tokens:
1,299,880 input,727,154 cached-input reads(a subset of input),269,159 output,0 separately reported reasoning,1,569,039 total input + output + reasoning. - Provider retries/errors:
0/0; provider-reported cost:not reported.
Lower Brier is better. Model calls were independently sampled per condition, so small cross-condition Brier differences are exploratory and must be read alongside the clean controls. Cached-input reads are observational, not an additional token charge. Wire, hydration, and model-token channels are reported separately; a short reference is not treated as a context-token saving. Hydration context tokens are the sum of upstream-tokenizer tokens in the compact JSON serialization of the coordination object included in each model request. This excludes chat framing, question text, peer forecasts, and output. It is a context-channel measurement, not a claimed token or cost saving.