sema-evalsIndependent evaluations

sema-evals

Open, causal evaluations for content-addressed semantics and multi-agent coordination. Each report is generated directly from a committed result artifact — a run manifest plus a redacted public trial derivative — and every statistic is recomputed from those trials at build time.

This site is deliberately independent: it reports what the trials show, including nulls and harness checks, and never markets a result. Runs are labelled by mode, and the distinction is structural, not editorial. A deterministic harness run validates the measurement machinery and is not evidence about model behaviour. A model pilot is exploratory and not confirmatory. Only a confirmatory run tests a preregistered hypothesis.

Disclosure: this site is maintained by a Sema core contributor and community moderator, and some experiments test code the maintainer authored. Independence here is methodological, not organizational — see the full disclosure for the safeguards (preregistration, raw records, build-time recomputation) that carry the evidential weight.

a2a-drift

A2A semantic-extension middleware demo: does cross-agent registry drift execute silently, or is it detected and halted, depending only on which A2A extension point is honored?

Runs
3
Latest
2026-07-23
Models
a2a-drift-demo-v1, composer-2.5-fast, Mistral-Nemo-Instruct-2407-TEE

Baseline: 20/20 drifted tasks execute silently; enforced: 20/20 halted, 0 false halts.

babel-relay

Babel Relay measures what happens when the meaning of a shared term silently drifts while work passes through a relay of agents — and which mechanism, if any, surfaces the drift before it becomes a wrong outcome.

Runs
3
Latest
2026-07-15
Models
MiniMax-M2.5-TEE, Mistral-Nemo-Instruct-2407-TEE

Latest addressed-enforced — silent divergence 0.0%, task success 98.9%

forecasting

Five-agent forecast council demo: does coordination-term drift corrupt the aggregate, or is the drifted forecast detected and excluded, depending only on whether content-addressed references are honored?

Runs
3
Latest
2026-07-23
Models
Mistral-Nemo-Instruct-2407-TEE

Baseline: 10/10 aggregates corrupted; enforced: 10/10 drifted forecasts excluded, 0 false exclusions.

hook-enforcement

Replays of the babel-relay drift scenarios through real agent harnesses with the sema ref-gate enforcing at the harness boundary; one experiment, multiple harness/model runs.

Runs
8
Latest
2026-07-22
Models
claude-haiku-4-5, composer-2.5-fast, gemini-3.5-flash, gpt-5.4-nano-low, gpt-5.6-luna (reasoning effort low), kimi-k2.7-code

Voluntary-leak span across 8 runs: 3%–78%; enforcement 217/217 halted.

security

Vulnerability recall at a fixed false-positive budget on mutation-backed Solidity cases.

Runs
1
Latest
2026-07-22
Models

Instrumentation: 9 TP / 0 FP under the enforced condition, scorer-validated.

sema-discovery

Deterministic pattern discovery, dependency resolution, and within-session reuse mechanism scaffold.

Runs
1
Latest
2026-07-23
Models
scripted-discovery-executor-v1

Latest scripted discovery end-to-end success 100.0%

sema-tax

Token, byte, and reuse costs of carrying semantic patterns.

Runs
2
Latest
2026-07-23
Models
sema-tax-simulator-v1, MiniMax-M2.5-TEE

Latest size/reuse arm — content beats prose bytes in 4/9 cells; latest curve best full-coverage score 0.996, score/1k tok 0.1318–0.4297

workflow-value

Seed-only workflow delivery and executable-validator mechanism scaffold.

Runs
1
Latest
2026-07-23
Models
workflow-value-scripted-executor-v1

Dataset gate seed-only — model pilot blocked

x402-contract-drift

x402 payer–seller demo: does payment-contract drift produce silent payment, or a refusal, depending only on whether the x402 extension surface is honored?

Runs
3
Latest
2026-07-23
Models
Mistral-Nemo-Instruct-2407-TEE

Baseline: 17/30 drifted contracts pay silently; enforced: 30/30 refused, 7 false refusals; 0 provider failures, 6 malformed outputs.