sema-evalsIndependent evaluations

Security — 20260722T080228609Z-order-20260716

Validates the security fixture catalog, train/heldout leakage guard, condition ladder, deterministic scorer, enforced-decision gate, randomization, and bundle/summary reproduction. Scripted-auditor outcomes are a construction, not evidence about language models (ADR 0014).

Condition Trials True positives False positives False negatives Parse failures Enforcement refusals
baseline 18 0 0 9 0 0
equal-prose 18 9 0 0 0 0
addressed-voluntary 18 9 0 0 0 0
addressed-enforced 18 9 0 0 0 0

Totals are summed over all cases in the condition and recomputed from trials.public.jsonl at build time. Rate-style endpoints (recall at the FP budget) live in the committed summary.json.