Open, causal evaluations for content-addressed semantics and multi-agent
coordination. Each report is generated directly from a committed result artifact — a run
manifest plus a redacted public trial derivative — and every statistic is recomputed from
those trials at build time.
This site is deliberately independent: it reports what the trials show, including nulls and
harness checks, and never markets a result. Runs are labelled by mode, and the distinction is
structural, not editorial. A deterministic harness run validates the measurement machinery
and is not evidence about model behaviour. A model pilot is exploratory and not
confirmatory. Only a confirmatory run tests a preregistered hypothesis.
Disclosure: this site is maintained by a Sema core contributor and community moderator, and
some experiments test code the maintainer authored. Independence here is methodological, not
organizational — see the
full disclosure for the
safeguards (preregistration, raw records, build-time recomputation) that carry the evidential
weight.
A2A semantic-extension middleware demo: does cross-agent registry drift execute silently, or is it detected and halted, depending only on which A2A extension point is honored?
- Runs
- 3
- Latest
- 2026-07-23
- Models
a2a-drift-demo-v1, composer-2.5-fast, Mistral-Nemo-Instruct-2407-TEE
Baseline: 20/20 drifted tasks execute silently; enforced: 20/20 halted, 0 false halts.
Babel Relay measures what happens when the meaning of a shared term silently drifts while work passes through a relay of agents — and which mechanism, if any, surfaces the drift before it becomes a wrong outcome.
- Runs
- 3
- Latest
- 2026-07-15
- Models
MiniMax-M2.5-TEE, Mistral-Nemo-Instruct-2407-TEE
Latest addressed-enforced — silent divergence 0.0%, task success 98.9%
Five-agent forecast council demo: does coordination-term drift corrupt the aggregate, or is the drifted forecast detected and excluded, depending only on whether content-addressed references are honored?
- Runs
- 3
- Latest
- 2026-07-23
- Models
Mistral-Nemo-Instruct-2407-TEE
Baseline: 10/10 aggregates corrupted; enforced: 10/10 drifted forecasts excluded, 0 false exclusions.
Replays of the babel-relay drift scenarios through real agent harnesses with the sema ref-gate enforcing at the harness boundary; one experiment, multiple harness/model runs.
- Runs
- 8
- Latest
- 2026-07-22
- Models
claude-haiku-4-5, composer-2.5-fast, gemini-3.5-flash, gpt-5.4-nano-low, gpt-5.6-luna (reasoning effort low), kimi-k2.7-code
Voluntary-leak span across 8 runs: 3%–78%; enforcement 217/217 halted.
Vulnerability recall at a fixed false-positive budget on mutation-backed Solidity cases.
- Runs
- 1
- Latest
- 2026-07-22
- Models
- —
Instrumentation: 9 TP / 0 FP under the enforced condition, scorer-validated.
Deterministic pattern discovery, dependency resolution, and within-session reuse mechanism scaffold.
- Runs
- 1
- Latest
- 2026-07-23
- Models
scripted-discovery-executor-v1
Latest scripted discovery end-to-end success 100.0%
Token, byte, and reuse costs of carrying semantic patterns.
- Runs
- 2
- Latest
- 2026-07-23
- Models
sema-tax-simulator-v1, MiniMax-M2.5-TEE
Latest size/reuse arm — content beats prose bytes in 4/9 cells; latest curve best full-coverage score 0.996, score/1k tok 0.1318–0.4297
Seed-only workflow delivery and executable-validator mechanism scaffold.
- Runs
- 1
- Latest
- 2026-07-23
- Models
workflow-value-scripted-executor-v1
Dataset gate seed-only — model pilot blocked
x402 payer–seller demo: does payment-contract drift produce silent payment, or a refusal, depending only on whether the x402 extension surface is honored?
- Runs
- 3
- Latest
- 2026-07-23
- Models
Mistral-Nemo-Instruct-2407-TEE
Baseline: 17/30 drifted contracts pay silently; enforced: 30/30 refused, 7 false refusals; 0 provider failures, 6 malformed outputs.