Real Codex execution / GPT-5.6 / independent evidence challenge
GPT-5.6 can challenge the evidence. It cannot rewrite the facts.
Codex Control Tower gives GPT-5.6 neutral claims and bounded raw evidence without giving it the reconciler's locked claim-status field or expected answer. Pinned Codex CLI runs in an empty ephemeral workspace and any tool event rejects the run. Only after validated SUPPORTS, CONTRADICTS, or INSUFFICIENT output does local code compare both views.
- Deterministic PASS, WARN, FAIL, NOT_RUN, and SIMULATED states remain immutable.
- Mission PASS is a structural precheck, not deterministic semantic truth.
- Compatible uncertainty stays distinct from agreement; conflict becomes HUMAN REVIEW REQUIRED.
- Full-file and included-content hashes, model identity, report age, and commit provenance stay inspectable.
Evidence boundary: InvoiceFlow Mini is a fictional sample project. The scans, two tests, evidence bundle, hashes, and recorded GPT-5.6 Codex execution are real project outputs.
Inspect the source and evidence on GitHub