sema-evalsIndependent evaluations

Forecasting Council — 20260723T135520372Z-order-20260716

Exploratory model pilot. Not preregistered, not confirmatory evidence. Historical questions are replayed after a selected-model audit bound to the exact dataset and, for the frozen-market-signal arm, the exact retained evidence bytes. Model outputs are objectively JSON-parsed and scored against frozen outcomes. The pre-existing primary endpoint remains corrupted aggregation under controlled registry drift; Brier score is reported as the registered utility metric with mandatory market-prior and independent-agent baselines. Conditions make independent model calls with no condition label in the model request; sampling/provider variation remains a nuisance and is preserved, not attributed to Sema.

Condition Trials Drift trials Detected Corrupted aggregations Correct exclusions False exclusions
baseline 20 10 0/10 10/10 0/10 0/20
addressed-voluntary 20 10 9/10 0/10 0/10 0/20
addressed-enforced 20 10 10/10 0/10 10/10 0/20

Drift-scoped columns use drift-injected trials as the denominator; the rest use all trials in the condition. Every count is recomputed from trials.public.jsonl at build time.

Forecast utility and resource channels

Descriptive utility result: enforcement did not improve mean aggregate Brier in this run (0.3109 enforced versus 0.3046 baseline; lower is better). This is compatible with the semantic gate doing mechanism work without improving model forecasting performance.

Model semantic transformation: among parseable round forecasts, aligned members reproduced the frozen source-market YES signal exactly in 498/505 cases (mean absolute error 0.0016), while polarity-drifted members reproduced its complement exactly in 50/56 cases (mean absolute error 0.0168). This shows model-side interpretation of the local definition; Sema's measured role is to address, detect, and enforce that semantic mismatch at aggregation.

Condition Mean Brier
aggregate
Mean Brier
market prior
Mean Brier
independent
Malformed / failed
model outputs
baseline 0.3046 0.3129 0.3129 15
addressed-voluntary 0.2945 0.3129 0.3107 13
addressed-enforced 0.3109 0.3129 0.3109 11

Lower Brier is better. Model calls were independently sampled per condition, so small cross-condition Brier differences are exploratory and must be read alongside the clean controls. Cached-input reads are observational, not an additional token charge. Wire, hydration, and model-token channels are reported separately; a short reference is not treated as a context-token saving. Hydration context tokens are the sum of upstream-tokenizer tokens in the compact JSON serialization of the coordination object included in each model request. This excludes chat framing, question text, frozen evidence, peer forecasts, and output. It is a context-channel measurement, not a claimed token or cost saving.