sema-evalsIndependent evaluations

Forecasting Council

Five-agent forecast council demo: does coordination-term drift corrupt the aggregate, or is the drifted forecast detected and excluded, depending only on whether content-addressed references are honored?

Deterministic demos validate the pipeline; model-pilot runs are exploratory evidence about the named setup only. Not preregistered, not confirmatory evidence.

RunCreatedModeModelEvidence claim
20260723T135520372Z-order-20260716 2026-07-23T13:55:20.372Z model-pilot unsloth/Mistral-Nemo-Instruct-2407-TEE Exploratory model pilot. Not preregistered, not confirmatory evidence. Historical questions are replayed after a selected-model audit bound to the exact dataset and, for the frozen-market-signal arm, the exact retained evidence bytes. Model outputs are objectively JSON-parsed and scored against frozen outcomes. The pre-existing primary endpoint remains corrupted aggregation under controlled registry drift; Brier score is reported as the registered utility metric with mandatory market-prior and independent-agent baselines. Conditions make independent model calls with no condition label in the model request; sampling/provider variation remains a nuisance and is preserved, not attributed to Sema.
20260723T095613198Z-order-20260716 2026-07-23T09:56:13.198Z model-pilot unsloth/Mistral-Nemo-Instruct-2407-TEE Exploratory model pilot. Not preregistered, not confirmatory evidence. Historical questions are replayed after a selected-model, zero-evidence leakage audit; model outputs are objectively JSON-parsed and scored against frozen outcomes. The pre-existing primary endpoint remains corrupted aggregation under controlled registry drift; Brier score is reported as the registered utility metric with mandatory market-prior and independent-agent baselines.
20260722T080220652Z-order-20260716 2026-07-22T08:02:20.652Z deterministic-harness Validates the forecasting-council scaffold: synthetic Polymarket-style questions, controlled per-agent registry drift, corrupted aggregation under baseline, voluntary detection, enforced exclusion, Brier baselines (market prior + independent-agent average), the no-drift false-exclusion guard, leakage-audit gate, condition pairing, and bundle/summary reproduction. Scripted-agent outcomes are a construction, not evidence about language models, and not evidence about live prediction markets (ADR 0017).