Forecasting Council — 20260723T135520372Z-order-20260716
Exploratory model pilot. Not preregistered, not confirmatory evidence. Historical questions are replayed after a selected-model audit bound to the exact dataset and, for the frozen-market-signal arm, the exact retained evidence bytes. Model outputs are objectively JSON-parsed and scored against frozen outcomes. The pre-existing primary endpoint remains corrupted aggregation under controlled registry drift; Brier score is reported as the registered utility metric with mandatory market-prior and independent-agent baselines. Conditions make independent model calls with no condition label in the model request; sampling/provider variation remains a nuisance and is preserved, not attributed to Sema.
- Mode:
model-pilot - Provider:
llm.chutes.ai - Model:
unsloth/Mistral-Nemo-Instruct-2407-TEE - Semantic backend:
semahash-python-api - Sema version:
0.3.0 - Canonicalization:
v2 - Order seed:
20260716
| Condition | Trials | Drift trials | Detected | Corrupted aggregations | Correct exclusions | False exclusions |
|---|---|---|---|---|---|---|
baseline |
20 | 10 | 0/10 | 10/10 | 0/10 | 0/20 |
addressed-voluntary |
20 | 10 | 9/10 | 0/10 | 0/10 | 0/20 |
addressed-enforced |
20 | 10 | 10/10 | 0/10 | 10/10 | 0/20 |
Drift-scoped columns use drift-injected trials as the denominator;
the rest use all trials in the condition. Every count is recomputed from
trials.public.jsonl at build time.
Forecast utility and resource channels
Descriptive utility result: enforcement did not improve
mean aggregate Brier in this run (0.3109
enforced versus 0.3046 baseline; lower is
better). This is compatible with the semantic gate doing mechanism work without
improving model forecasting performance.
Model semantic transformation: among parseable round forecasts,
aligned members reproduced the frozen source-market YES signal exactly in
498/505 cases
(mean absolute error 0.0016),
while polarity-drifted members reproduced its complement exactly in
50/56 cases
(mean absolute error 0.0168).
This shows model-side interpretation of the local definition; Sema's measured
role is to address, detect, and enforce that semantic mismatch at aggregation.
| Condition | Mean Brier aggregate |
Mean Brier market prior |
Mean Brier independent |
Malformed / failed model outputs |
|---|---|---|---|---|
baseline |
0.3046 | 0.3129 | 0.3129 | 15 |
addressed-voluntary |
0.2945 | 0.3129 | 0.3107 | 13 |
addressed-enforced |
0.3109 | 0.3129 | 0.3109 | 11 |
- Historical source: SimpleFunctions Settled Prediction Markets
at revision
a27e3e9307266481d51e087fffd5bf934410e01c, licensedCC-BY-4.0; attribution: SimpleFunctions (simplefunctions.dev). - Wire payload:
489,166 bytes. - Registry hydration/context:
236,490 bytes. - Retained frozen evidence:
4,560 bytessummed once per trial (deduplicated within a trial, not multiplied by member calls). - Tokenizer-derived coordination hydration:
118,620 context tokensusingmistralai/Mistral-Nemo-Instruct-2407@04d8a90549d23fc6bd7f642064003592df51e9b3(tokenizer SHA-256e11c71726323d33da7b8d6f6f269f1988931c0a52b7122bcdd8c05042974e0db). - Provider model tokens:
539,112 input,419,497 cached-input reads(a subset of input),49,418 output,0 separately reported reasoning,588,530 total input + output + reasoning. - Provider retries/errors:
0/0; provider-reported cost:not reported.
Lower Brier is better. Model calls were independently sampled per condition, so small cross-condition Brier differences are exploratory and must be read alongside the clean controls. Cached-input reads are observational, not an additional token charge. Wire, hydration, and model-token channels are reported separately; a short reference is not treated as a context-token saving. Hydration context tokens are the sum of upstream-tokenizer tokens in the compact JSON serialization of the coordination object included in each model request. This excludes chat framing, question text, frozen evidence, peer forecasts, and output. It is a context-channel measurement, not a claimed token or cost saving.