Primer delivery map: from concepts to independent practice¶
The completion target is a book that teaches a reader to design evaluations, improve datasets, qualify graders, inspect executions, and operate an application through CI/CD, shadow, canary, expansion, restriction, and rollback. Interview preparation uses the same evidence and reasoning.
Existing chapters, source mappings, and examples remain part of the curriculum. This map adds explicit implementation and learning requirements. A chapter title or passing scaffold test does not prove depth or runtime validity.
Required learning contract¶
Final book acceptance checkpoint — 7 September 2026¶
The later source-depth pass has closed the named teaching gaps in the consolidated checkpoint below. The source ledger now maps all 118 interview questions, the eight course labs and all 46 ranked frontier-review entries to practice. Added worked solutions cover evaluation budgets, trajectory comparison modes, judgment-versus-ranking validity, adaptive item selection and uncertainty-wording interventions. Existing sources, chapters and historical evidence remain retained.
| Requirement | Final evidence and boundary |
|---|---|
| Concepts, examples and interview preparation | Source-family audits, numbered micro-katas, supplementary solved exercises and the capstone decision defense; source routes are not independent replications of every cited paper. |
| Dataset improvement and calibrated measurement | Executed sampling, annotation, judge qualification, artifact replay and incident-promotion controls; synthetic labels remain visibly separate from independent human evidence. |
| Productionization teaching | CI conformance, paired evidence, shadow/canary, exposure budgets, restriction, process interruption and rollback studies, with explicit application-integration prerequisites. |
| Software and navigation | Existing 679-test conformance baseline; subsequent hosted conformance passed through f7b9f65; the final local build has 44 pages and 6,907 checked local links with zero errors. |
| Delivery | Final coverage-map and wording-exercise changes still require their own successful hosted publication receipt. |
This is a book and worked-lab acceptance checkpoint, not an independently awarded A+ grade, a production safety certification or a claim that the learner's live-model, human-annotation and external-harness projects have been performed. The historical platform and empirical-study limitations below remain valid at their stated scopes; they do not authorize further scope expansion during this finishing pass.
Consolidated acceptance status — 7 September 2026¶
This table supersedes the current-status interpretation of older checkpoints below; their historical evidence and limitations remain intact. It does not certify an A+ grade or erase uncompleted experiments.
| Acceptance requirement | Inspected evidence | Verdict / remaining action |
|---|---|---|
| Preserve the supplied curriculum | Six Markdown inputs are fingerprinted; all 118 interview IDs, question variants and answer outlines are retained; companion PDFs preserve those IDs; six companion files have a reference-recovery index | Preservation checks pass at these stated scopes. Sentence/example-level equivalence and claim-linked recovery of companion references remain unverified. |
| Explain concepts through worked practice | Numbered kata sequence, retained studies, complete capstone solutions and indexed additional exercises for judge lineage, skills, voice timing and human workload | Substantial worked coverage exists. The remaining content audit must map each substantive source requirement to an actual explanation and solution, not count headings or demand invented field results. |
| Demonstrate evaluation-to-decision mechanics | Paired artifacts, grader qualification controls, uncertainty studies, durable effect/completion/window/controller stores and an executed incident-to-rollback packet | Local mechanics are demonstrated under their documented contracts. Generic interrupted-agent continuation and ownership fencing are not implemented. Do not advertise a fully recoverable production application. |
| Run the software and deliver the book | 679 local tests passed; hosted Publish book run 34090136429 succeeded for aed856e164528089946ffb63e5e0ed7dc9aac2f1, including conformance and Pages deployment |
Hosted delivery is verified for that revision. Subsequent documentation changes require publication, not another unchanged local full-suite run. |
| Make learning material reachable and readable | A built-site check inspected 44 HTML pages and 6,345 relative local links/anchors with zero errors; published judge/recovery pages had no desktop overflow and Capstone Solution G expanded | These checks pass for the inspected build/pages. They do not establish mobile, Safari or every-page visual quality. New links still need the normal strict build. |
| Distinguish teaching from empirical qualification | Chapters label synthetic inputs, fixed controls, source consistency, independent review requirements and deployment authority separately | No actual model-quality, independently human-calibrated or live application-release claim is earned by the local controls. Such claims require their own approved data, reviewers, budgets and deployment environment. |
Bounded remaining book work: complete the source-requirement-to-worked-solution audit, repair only the concrete gaps it identifies, and publish the resulting documentation with a successful hosted receipt. Do not add new topics or commission optional empirical studies merely to improve a completeness label. Preserve existing named experimental commitments as unfinished until their own acceptance evidence exists; explaining their protocol is not executing them. The objective remains active because the source-depth acceptance is not yet proved.
Publication preflight — 7 September 2026¶
The current local suite passes 679 tests in 63.089 seconds at acefbc0; this is software conformance, not model or deployment qualification. The publication-range whitespace check identified only a trailing blank line in the preserved interview appendix, now corrected. The remote main and latest successful Publish book run (34047693398) remain at 019ee70a0bb48930c24afc5704e7b27b6e729547. Therefore that hosted success does not verify the newer local content or recovery integration. The configured Pages workflow now depends on evaluation conformance before deployment. Final publication still requires review of the pending public change set, push, successful hosted execution for the exact revision and rendered-page verification; the local tests must not be reported as those receipts.
Supplied-input identity checkpoint — 7 September 2026¶
The source ledger now resolves the four numbered Markdown filenames to their actual titles and fingerprints all six Markdown inputs plus the later DOCX review. All six Markdown titles have corpus mappings. This closes input identification only: the ledger explicitly distinguishes its editorial section map from the still-required content-depth audit, and does not claim companion-format equivalence was reverified. No new research, runtime study or production qualification was performed.
Completion scope and capstone acceptance — 7 September 2026¶
The user's objective is a comprehensive teachable primer, not an assertion that its synthetic agent has earned production qualification. External reviewers, paid model runs and a real application deployment are prerequisites for those empirical claims, not automatic prerequisites for explaining and exercising the methods honestly. Existing experimental commitments remain listed below; this distinction does not mark an unfinished experiment complete or waive source coverage, worked examples, correctness or publication checks.
Capstone Variant G now assesses the durable components together using existing test evidence: effect inspection, retained completion, missing outcomes, attempt continuation and historical controller acknowledgement. Its solution rejects unproven end-to-end guarantees. Remaining local acceptance work is a retained integrated recovery/decision demonstration and its CI path, followed by the supplied-source coverage and final publication audit. Do not expand the topic inventory or commission external studies merely to close a teaching-status label. No overall completion is claimed here.
Integrated incident rollback checkpoint — 7 September 2026¶
The executed recovery-to-decision example now retains a registered 40-member window, an actual wrong-order tool effect, child-process termination before response persistence, the missing completion and a durable rollback receipt that replays unchanged. The CI workflow generates this packet before uploading artifacts. The study module has 86% combined statement/branch coverage from focused checks; its retained packet reports conformance true and deployment authorization false. This closes the narrow local incident-containment integration and CI configuration item above—not agent continuation, fencing, qualified full-window promotion or observed cloud execution. Supplied-source depth acceptance and final publication verification remain open; do not add optional studies merely to expand this checkpoint.
Durable controller acceptance checkpoint — 7 September 2026¶
Kata 105 adds atomic persistence around the existing simulation state machine: complete predecessor comparison, immutable decision-input binding, receipt replay and state/receipt commit in one transaction. RED 7daffc0 preceded implementation. Twenty-five focused controller/state-machine tests pass with 87% combined statement/branch coverage for the new module; lint and type checks pass. Real processes exercise contention and post-commit loss, and a SQL trigger verifies rollback after receipt insertion. No additional research or full-suite rerun was needed.
Remaining integration criteria are qualified observation ingestion tied to registered membership, action-boundary fencing and interrupted-attempt recovery, then a retained end-to-end process packet and CI wiring. The controller-local transaction does not close these requirements or the separately blocked human/model/deployment qualification requirements.
Joined window evidence checkpoint — 7 September 2026¶
Kata 104 joins registered membership to served-candidate completion evidence and observed order state. Integration 74be89a preserves all forty members, distinguishes missing records from missing completion, retains unknown labels and exposes wrong-order effects after a real process kill. Completion batches and world batches each use one transaction; there is no cross-database atomic snapshot or controller decision. Optional execution journals served candidates only; repeating it can reuse candidate evidence while rerunning baseline controls, not create independent trials.
The 48 focused integration/regression tests pass; lint, type checks and dependency audit pass. These existing checks are reused rather than repeating the full suite. Remaining acceptance work is durable controller acceptance and interrupted-attempt recovery, a retained integrated process packet with CI wiring, independently qualified human/model evidence, and demonstrated application deployment controls. This checkpoint does not close those requirements. Existing content and historical artifacts are preserved.
Durable window registration checkpoint — 7 September 2026¶
Kata 103 adds optional immutable window registration to the actual exposure driver. A shared pure builder creates forty full ordered cases, routes, namespaces and effect scopes. Registration binds the caller-supplied predecessor, timing, fault/maturity settings and manifest before any agent runs. A real spawned-process kill before the first agent leaves all forty members registered and zero effects charged.
RED 096c4fe and GREEN 0cb2450 preserve the implementation checkpoints. Forty-five focused tests pass; combined statement/branch coverage is 100% for the registry and 98% for the exposure driver. Lint, type checks and dependency audit pass. Main-agent comparison with the pre-refactor seven-window CLI confirms parity of non-timing/non-hash fields. The routing example and published process test pass; the catalog has 103 unique rendered kata links and the new solution opens without horizontal overflow at the checked desktop viewport.
The full suite passes 661 tests in 81.846 seconds with 91% combined statement/branch coverage. Independent review approved the registration-only scope; fresh source-checked release controls and budgeted generation/replay pass with deployment authorization false. Strict book build passes.
The registry validates saved plans against the current fixed-study builder; source drift can therefore refuse old-plan inspection rather than silently reinterpret it. Caller manifest equality and local campaign identity are not authenticated provenance or controller acceptance. Registration is not atomic with later effects, does not store per-request completion, and does not permit interrupted served candidates to rerun. Next, join registered membership to retained completion/missing-outcome accounting before adding durable controller acceptance. Existing legacy mode and historical packets remain unchanged; no production or live-model claim is added.
Completion binding hardening checkpoint — 7 September 2026¶
The two completion-review follow-ups below are repaired. RED 28f3ded reproduced cached Boolean/integer identity confusion and publication of an invalid first result. The journal now compares canonical JSON identity values and validates initial evidence before marking it completed. Invalid publication leaves the original reservation incomplete. Thirty focused completion/exposure tests pass with 97% combined statement/branch coverage for the completion module; lint, type checks, dependency audit and independent review pass. Kata 102 explains the counterexample. This is selected-field consistency, not a full artifact-schema or provenance verifier; controller/attempt recovery remains open.
Durable completion evidence checkpoint — 7 September 2026¶
Kata 102 adds a schema-1 completion journal around the existing durable run_case, without changing campaign schema 2. Reservation precedes execution; immutable completion evidence is stored afterward. A real spawned-process kill after refund commit preserves one charge and missing completion, while a finished record reopens without agent/tool execution. A worked ten-request interview example distinguishes original-attempt success, response coverage, eventual resolution and attempt counts.
RED 7631a2b established the missing API; RED ac2feda reproduced coherent artifact substitution against an unchanged request record. GREEN f50234a adds strict serialization, request/artifact binding, campaign-local identity and immutable publication. The 28 focused completion/exposure tests pass; completion-module combined statement/branch coverage is 97%. Targeted lint/type checks and dependency audit pass. Hashes and caller-supplied manifest equality do not establish authenticated provenance.
The stable-source full suite passes 647 tests in 78.646 seconds. The published process test, fresh source-checked release controls, fresh budget replay, strict book build and 102 catalog anchors pass. Desktop solution expansion has no horizontal overflow at the checked viewport. Final review found no blocker in the limited contract; type-exact cached identity comparison and validation before initial publication remain hardening follow-ups, not claims of a complete artifact-schema verifier.
This closes optional one-shot completion storage only. The exposure driver does not yet register durable window membership or use this journal, and no interrupted attempt can resume or transfer ownership. Per-invocation history, stale-worker fencing, controller checkpoints and a new retained process-level campaign packet remain open. No model-quality, customer-delivery, cloud-CI or deployment authority is claimed.
Recovery lifecycle contract checkpoint — 7 September 2026¶
The recovery implementation contract makes stage 3 implementable without pretending that a committed refund reconstructs an agent session. It maps the current volatile transcript, response and window-membership seams to durable records; separates operational identity from versioned grading identity; requires action-boundary ownership fencing; and specifies eight real-process crash/race controls. Original-attempt completion, eventual resolution and recovery burden remain distinct outcomes.
This is a source-grounded architecture and teaching addition, not a new runtime result. The first supported recovery procedure is explicitly separate deterministic reconciliation; arbitrary agent continuation remains unimplemented. Existing schema-2 behavior and historical packets are preserved. The next implementation starts with durable window/request registration and immutable attempt/completion evidence before controller recovery and a new process-level packet.
Durable exposure wiring checkpoint — 7 September 2026¶
Kata 101 completes stage 2: the optional durable backend traverses the existing exposure and multi-order runner into RefundWorld. Reopening between three windows preserves cumulative charges of 0, 2 and 2. Baseline and shadow remain isolated. All competing order evidence is read in one database transaction; before/after campaign balances are explicitly not per-request attribution.
Initial RED/GREEN 0b882f3 / 66e4559 established wiring. Review found that empty order events do not prove an attempt never started. RED 347d990 and GREEN eb827a0 add schema-2 unique durable request admission before agent execution, including zero-tool, clarification-only and competing-process cases. Final tests 2358dde add operational-input identity and schema-refusal boundaries. Independent review found no blockers; 77 focused tests pass with 97% combined statement/branch coverage and 92% branch coverage across the targeted modules. Targeted lint/type checks and dependency audit pass.
The full suite passes 634 tests in 78.181 seconds with 90% combined statement/branch coverage. The published three-window example, fresh source-checked release controls, fresh budgeted packet generation/replay, strict book build and 101 catalog anchors pass locally. The new solution opens without horizontal overflow at the checked desktop viewport. Existing artifacts remain unchanged. This supersedes the prior checkpoint's unwired-driver limitation only: durable responses, resumable attempts, controller checkpoints and a retained campaign-process recovery packet remain open. No paid-model, human-qualification, cloud-CI, Safari/mobile or production result is claimed.
Durable world backend checkpoint — 7 September 2026¶
Kata 100 completes the first storage/world stage of the integration contract. Existing RefundWorld tools can use an optional SQLite campaign that persists immutable seed bindings, current authorization, effects and events atomically. Reopened instances see prior evidence and cross-instance revocation; durable_state() supplies a paired snapshot/event projection. Reference-agent output, timeout handling and regrading match the memory path on the tested cases. Unsafe commits retain their specific unauthorized-action failure.
Worker RED/GREEN d3083c4 / aaa769b and independent grader RED/GREEN f04f21f / e71101d capture the initial implementation. Review then reproduced a partial-effect commit caused by treating every ValueError or PermissionError as an expected tool rejection. RED 7e6d5d0 and GREEN 62421a5 require matching state/event evidence before deferring a known failure; unexpected partial transitions roll back. Final cleanup 3b7bc5f changes formatting only. Independent re-review found no remaining boundary blockers.
The full regression run passes 619 tests in 76.179 seconds. Targeted lint/type checks, fresh source-checked release controls, fresh budget replay, strict build, published snippet and desktop solution expansion pass. The catalog now links 100 unique katas with valid rendered anchors. No historical artifact was regenerated or removed.
The exposure driver and multi-order runner do not yet select this backend. Durable customer-response storage, attempt/window checkpoints, whole-agent continuation and retained campaign-process evidence remain open. The tests reuse their retained output when regrading; they do not prove that a killed agent's response was recovered. No cloud execution, paid model study, human qualification or production authorization is claimed.
Durable campaign integration architecture checkpoint — 7 September 2026¶
Source inspection of execute_window, run_case, MultiOrderWorld, RefundWorld, grading and recovery storage established why the two budget implementations cannot simply be joined: the campaign's authoritative effect and authorization snapshot remain in memory. The integration contract now specifies one durable owner for seed/authority/effect/event state, mutually exclusive backends, served-only accounting, stable reopen identity, consistent grader projections and request/window recovery checkpoints. It preserves the existing public authorization-first tool behavior and separates historical reconciliation.
The five-step handoff starts with durable storage/world tests, then existing-path wiring, process recovery, retained evidence and CI. Agent continuation and grading across interrupted attempts remain explicit decisions before claiming full campaign recovery. This is an implementation-ready architecture contract, not a completed integration or another standalone payment demonstration. No new runtime performance, model quality or deployment result is claimed.
Durable payment-budget checkpoint — 7 September 2026¶
Kata 99 extends the existing SQLite recovery boundary with optional fixed database-owned policy. New effects automatically check count/currency capacity inside the same transaction as approval and payment insertion. Consumption comes from committed payments; exact historical replay remains free. Separate spawned processes compete for the last unit, and owned-worker kills before/after commit test rollback versus preserved consumption. The original recovery artifact and unbudgeted behavior remain intact.
RED 061e434, GREEN a5441fb and reviewed cleanup 9feeb94 preserve explicit checkpoints. Twenty durable/legacy recovery tests pass; targeted Ruff/Pyright checks are clean. The published snippet runs, the practice catalog contains 99 unique entries, and fresh source-checked release controls and fresh budgeted replay pass locally with deployment authorization false. Historical v1 budget bytes are pinned rather than regenerated when source inventory changes.
At final runtime revision 9feeb94, all 603 tests pass in 79.247 seconds with 90% overall branch-inclusive coverage. Strict build, dependency audit, desktop solution expansion and rendered 99-entry catalog checks pass. Independent code and teaching reviews found no blockers.
This establishes a bounded durable mock payment cap, not a durable integration of the exposure campaign, a retained process-budget study packet, authenticated policy ownership, hostile-process isolation or remote payment atomicity. Those remain required extensions, along with model/human qualification and actual deployment evidence. Nothing was pushed or deployed.
Advanced capstone integration checkpoint — 7 September 2026¶
The advanced defense round connects budgeted execution, replay/CI evidence and consequential-risk dependence to the final learner assessment. Three pressure variants have complete worked answers and a noncompensatory rubric. Eight exact pointers locate the newer budget evidence; the original three-artifact exercise, reference index and answer key remain unchanged. Learners must preserve separate study identities and defend unfinished work, current authority and causal assumptions.
All eight pointer values were checked against the retained packet. Independent review found no blockers and confirmed preservation of the original exercise. Twelve book-example/link tests, strict build, dependency audit and desktop expansion of all three new solutions pass. This closes an assessment-integration gap, not a human evaluation of learner readiness, a model study or a deployment qualification.
Complete practice navigation checkpoint — 7 September 2026¶
The homepage now links a complete topic-grouped practice index, preserving the earlier learning route and exercise explanations. A rendered-heading inventory found exactly 98 numbered katas, with no gaps or duplicate identifiers. The new seven-topic index links each once; all 98 anchors resolve in the built book. Numbers remain identifiers, not prerequisite or difficulty claims. Learners record predictions, results, changed-assumption explanations and unresolved evidence, then complete the existing capstone.
Twelve book-example/link tests, strict build and dependency audit pass. Desktop review confirms seven readable tables without page overflow and a working cross-chapter link. This verifies navigation, not the depth or empirical validity of every exercise, mobile rendering, full curriculum completion or publication. Update the catalog when adding or moving a numbered kata.
Frontier-risk dependence checkpoint — 7 September 2026¶
Kata 98 adds executable depth to the existing six-claim safety-case framework. Three synthetic populations share the same attempt and failure-condition marginals but yield 0, 1 and 100 consequential failures. The solution derives the conditional probability identity and logical bounds, distinguishes challenge-set effectiveness from field prevalence, and applies an explicit illustrative expansion policy without pretending to estimate real harm.
The published snippet executes with the stated results; independent review finds no mathematical or claim-scope blocker. Twelve book-example/link tests, strict build, dependency audit and desktop expanded-solution check pass. No runtime code, historical artifact or existing chapter content was removed. Actual joint-pathway evidence, independent human outcomes, live-model studies and deployment qualification remain open.
Budget replay CI wiring checkpoint — 7 September 2026¶
Kata 97 connects fresh budget packet generation and verification to the conformance workflow. RED aa50e75 detected the missing step; GREEN 2a4d40e added both commands with explicit Bash execution. A structural regression checks unconditional execution, absence of error suppression, matching packet paths and always-run artifact retention. Ten targeted workflow/study tests and local fresh generation/replay pass; independent review found no blockers or added privilege. This closes dedicated workflow wiring, not observed cloud execution or required branch protection. The earlier full-suite result remains historical; this change does not modify runtime source.
Retained budgeted campaign checkpoint — 7 September 2026¶
Kata 96 adds a versioned packet and fixed-builtin replay for the existing budget boundary. Three executed windows retain complete cases, tool events, ledgers, scope labels, namespaces, budget snapshots and decision inputs. Shadow/canary/restricted windows produce cumulative charges 0/2/2 and denials 0/2/6. Restriction does not mint new allowance. The original unbudgeted study remains unchanged.
Implementation through a4be341 includes RED/GREEN tests for rehashed evidence mutations, strict JSON types, timing validation, immutable output and interpreter portability. Independent review found no additional blockers. Retained replay, 25 targeted tests, strict book build, dependency audit and desktop solution expansion pass. A fresh source-checked current-release packet reports conformance true and deployment authorization false.
The final full suite passes all 593 tests in 77.575 seconds with 90% overall branch-inclusive coverage. This closes the earlier retained-packet follow-up. Replay checks consistency with fixed code in an exact source/interpreter environment, not authenticated provenance or real-world reproducibility. Dedicated cloud-CI wiring for this command, durable transactional budgets, live application integration, human calibration and model-backed qualification remain open. Nothing was pushed or deployed.
In-process hard-budget implementation checkpoint — 7 September 2026¶
Kata 95 now executes an optional campaign budget through execute_window, run_case, MultiOrderWorld and the authoritative refund commit boundary. A shared lock covers normal authorization changes/checks, replay, count/per-currency accounting and effects. Served-candidate scope excludes baseline/shadow counterfactuals. The worked window selects four candidate requests, commits two EUR 4,500 refunds, denies two and retains all four outcomes; forty baseline executions do not consume the allowance. Independent namespaces prevent false cross-world deduplication. Consumed worlds cannot reset; uncertain accounting/effect failures retain capacity and latch new effects.
Budget RED/GREEN checkpoints 530e942 / 85fc054, accounting-exception RED/GREEN 652156c / 7a5a7d2, cleanup 008d3a4. Main response-path RED/GREEN 5da0be4 / fd60407 repaired blind confirmation after returned denial; 809fac7 / 3877fe1 added ledger reconciliation for ambiguous returned status. A confirmed historical refund is distinct from a new effect caused by the latest attempt. Independent code/security review found no blocking defect within the trusted-harness, single-process mock scope.
At code revision 3877fe1, all 584 tests pass in 73.204 seconds: 90% overall branch-inclusive coverage, 94% across the five affected runtime modules and 97% for action_budget.py. Targeted Ruff/Pyright, the installed dependency audit, published Python example, strict build and desktop solution expansion pass. A fresh source-checked current-release run retains one HOLD and six BLOCK controls with deployment authorization false. Historical artifacts remain unchanged.
This closes the first in-process effect-cap integration, not the full hard-budget contract. The existing CLI and retained v1 exposure packet remain unbudgeted; the new snippet keeps its execution records in memory. A versioned retained budgeted study with independent replay, durable transactional enforcement, authenticated campaign ownership, additional request/customer/in-flight caps and real application rollout remain open. No paid model run, human qualification, cloud CI execution, mobile/Safari verification or production deployment is claimed.
Hard-budget integration design checkpoint — 7 September 2026¶
Read-only tracing of exposure_study, order_resolution, world and recovery_worker identified the actual effect hooks and three integration hazards: baseline/shadow controls can consume the wrong budget scope; repeated order/key strings in independent worlds can cause false deduplication; and commit-before-timeout accounting cannot wait for tool success. The hard-budget implementation contract now defines served-candidate scope, campaign ownership, authoritative per-currency amounts, replay identity, concurrent final-unit tests, reset/timeout behavior and outcome denominators. Historical artifacts remain unchanged.
This checkpoint is a code-grounded design, not an implemented cap. The strict book build, whitespace check and installed dependency audit pass; no runtime or financial-enforcement result is claimed. Next is a RED regression sequence at the shared commit boundary, followed by in-process integration and separately demonstrated durable transactional enforcement. A standalone budget counter or another disconnected payment demo would not close the existing exposure-path requirement. Live application deployment and the broader goal remain open.
Judge scope-evidence clarification checkpoint — 7 September 2026¶
Audit of the older groundedness example found wording that described synthetic cases as independently reviewed and listed English scopes as qualified without the required scoped evidence. The revision preserves the original counts, hypothetical decisions and companion artifact, but explicitly identifies the missing provenance, slice denominators, error bounds and authority. It concludes that no production qualification is established by the supplied table. Three false-pass estimands and the acceptance/review workload are now worked separately.
Independent review found no blocking correction. The numerical probe, all 35 book-example/scaffold tests, strict build, whitespace check and dependency audit pass. The rendered warning and qualification limit are visible in the desktop page without horizontal overflow. This is a substantive teaching-claim repair, not new human review, a runtime authority change, a full-suite rerun or publication. Broader empirical qualification and curriculum completeness remain open.
Probability-calibration decision checkpoint — 7 September 2026¶
Kata 94 deepens the existing probability-calibration example using the unchanged ten-row artifact. Its executable hindsight-constant comparison yields ECE zero but Brier 0.24000, log loss 0.673012 and zero accepted cases at threshold 0.8; selective risk is undefined. The worked solution explains why calibration, ranking, label independence and action utility are separate requirements, and links to the review-workload workshop. The constant uses the scored labels and is explicitly not a held-out model result.
The published Python example executes, all twelve book-example/link-validation tests pass, the strict book build and dependency audit pass, and the new worked solution expands without horizontal desktop overflow. These checks do not claim a new full runtime-suite run, actual recalibration experiment, mobile/Safari review, deployment or completion of the broader curriculum goal.
Book navigation enforcement checkpoint — 7 September 2026¶
A rendered local scan checked 43 HTML pages and 5,624 relative link occurrences, including target anchors, with no missing targets. The scan did not check external URLs, root-relative links, images or scripts. It exposed a prevention gap in the build configuration: a missing Markdown exercise anchor was informational and still passed --strict. RED checkpoint 6ad0df7 reproduces that behavior in a tiny isolated site; GREEN 942494c makes anchor warnings fail strict builds. The valid-anchor control and missing-page control distinguish the intended failure from a generally broken build.
All 29 scaffold and link-validation tests pass, as do the full strict book build, targeted Ruff/Pyright and installed dependency audit. No new full runtime-suite or coverage result is claimed for this configuration change. The gate checks target existence, not substantive coverage, working external references, accessibility, live browser layout or production authority. Cloud CI execution and publication remain unobserved, and the broader curriculum and empirical requirements remain open.
Shipping shortcut counterexample checkpoint — 7 September 2026¶
Kata 93 repairs a teaching mismatch without removing the original example. The short release_action function is now warned as intentionally unsafe: all-NaN measurements and unsupported finite reports both yield its misleading canary string, whereas missing cost represented by None blocks. The worked solution separates representation, measurement and authority, with an implementation handoff to the existing evidence, CI, exposure, containment and recovery labs. No production runtime was repaired or deployment authorized by this documentation change.
RED checkpoint 1567d45 records two failing chapter-contract tests; GREEN 4c9745b adds the warning and executable probes. Independent review found no blocking correction. The full suite passes 565 tests in 72.263 seconds with 90% branch-inclusive coverage across cx_eval_lab; targeted test-file Ruff/Pyright and the installed dependency audit pass. The strict build passes, and the warning plus expanded solution render without horizontal desktop overflow. The tests intentionally reproduce the defective shortcut; they do not certify it as a safe gate. No public deployment, mobile/Safari verification, paid model evaluation or completed production integration is claimed.
Learner entry-point checkpoint — 7 September 2026¶
The homepage portfolio route now names seven learner outputs, concrete checks and remedial chapter links. It connects numerical calibration to the existing seven-checkpoint route and requires a coding or research-report outcome for cross-domain practice. Independent review caught and corrected a distinction: knowledge-to-action and process recovery extend the CX domain; they are not second application domains. Homepage architecture cards now distinguish recorded local controls from still-unproven model and distributed-system transfer instead of describing all trace-producing work as pending.
The strict site build, whitespace check and installed-environment dependency audit pass. All 149 local homepage links and anchors resolve; the new domain-transfer link navigates to its rendered heading, and the desktop document has no horizontal overflow. These are curriculum-entry and local site checks, not a full-book usability audit, mobile/Safari verification, new runtime-test run or publication. Learning statuses and portfolio completion confer no application authority. The full empirical and operational requirements remain open.
Numerical calibration decision checkpoint — 7 September 2026¶
Katas 90–92 connect criterion-specific judge errors, abstention, dataset sampling, exact zero-event uncertainty, hypothetical prevalence shift and human-review capacity in one worked qualification HOLD. Three executable examples distinguish 95% agreement from a 25% unsafe false-pass rate, derive the 13.91% upper bound for zero events in twenty independent unsafe trials, and compute the selective judge's hypothetical 90.28% nominal review-capacity load. A must-hit interview rubric connects the calculations to owner-specific next evidence. Existing lessons and source mappings remain intact.
Independent numerical review found no blocking correction; its sampling-design clarification was incorporated. All three published Python blocks execute successfully, the strict book build and whitespace check pass, and all three solutions expand in the local desktop preview without horizontal document overflow (1585-pixel document width, 1595-pixel viewport). This documentation-only checkpoint does not claim a new full runtime-test run, Safari/mobile verification or publication. Counts and staffing are authored inputs, not independent annotation, model measurements or production evidence. Actual empirical qualification and the broader curriculum objective remain open.
Executed coding-patch evaluation checkpoint — 7 September 2026¶
Katas 87–89 add a second outcome domain with actual Python subprocess execution. Five authored patch controls distinguish visible-example overfitting, an unknown-outcome denominator bug, caller-state mutation and a protected-test edit. Four controls execute two suites each, totaling forty registered checks; the fifth is blocked before execution. Full code/diffs, input/output/exception records, test definitions and file inventories remain inspectable in the retained study. Replay runs current built-ins only, never source supplied by an artifact.
Review strengthened three boundaries: expected negative controls must complete with their intended failure pattern rather than time out; parent grading must derive results from retained outputs and input changes; semantic numeric equality must accept valid rates without accepting boolean counts. The book now also connects dated OpenAI coding-task validity findings and Anthropic resource-enforcement findings to the local lesson. Those external studies are not reproduced here.
RED checkpoints 71c984b and d7f20b1; GREEN implementation/artifact feab2cd. A fresh branch-inclusive run passes all 563 repository tests: 90% overall coverage and 91% across the two new modules. Targeted Ruff/Pyright and the installed-environment dependency audit pass. The documented fresh CLI, published replay snippet and retained-artifact replay pass; strict book build and three expanded worked solutions at a 1152-pixel desktop viewport pass. Current-release composition remains one HOLD and six BLOCK controls, with deployment authorization false. The conformance workflow includes the new local command; no cloud execution or publication is claimed.
This closes a bounded second-domain evaluation-mechanics gap, not actual coding-agent performance, protected acceptance access, independent test validity or secure execution of arbitrary patches. Actual model-backed transfer and the wider curriculum and operational qualification requirements remain open.
Failed-run CI evidence checkpoint — 7 September 2026¶
Katas 85–86 repair a concrete operating gap: timeouts previously bypassed command diagnostics, and failed conformance runs had no structured partial-run summary. The implementation now retains verified prior checks and a separate failure record, preserves the original failure, distinguishes launch/timeout/packet/replay errors, and never fabricates a candidate verdict or continues later controls. The lesson executes two ordinary local controls followed by a real harmless child-process timeout. It also explains why identical exit integers from different commands have different meanings, and why upload eligibility is not confirmed durable storage.
RED checkpoint b7261eb, GREEN a54990f, final validation/refactor 6bef7e2. At the final code revision, all 551 tests passed in a fresh branch-inclusive coverage run: 90% overall and 88% for ci_conformance.py. Targeted Ruff and Pyright passed; the isolated installed-environment dependency audit found no known vulnerabilities. The published capstone and timeout snippets execute successfully, the strict book build passes, and both new solutions open without page overflow at a 1152-pixel desktop viewport. Fresh current-release composition retains one HOLD and six BLOCK controls, with deployment authorization false.
These checks establish local failure-path behavior, not durability after disk failure, termination of the parent, host loss or arbitrary process-tree escape. No cloud CI run, branch-protection enforcement, actual model calibration or application rollout is claimed. The full curriculum objective and empirical transfer requirements remain open.
Integrated release-decision capstone — 7 September 2026¶
The capstone now requires a bounded decision memo and machine-readable evidence index across three existing retained studies. Its worked HOLD decision and three pressure variants connect measurement validity, missing telemetry, current qualification, incompatible evidence scopes and side effects completed before rollback. A noncompensatory rubric requires inspectable claims, correct denominators and owner-specific reopening evidence. Existing chapters and artifacts remain intact.
The chapter's published Python example checks three file-byte hashes and 29 exact values without running models or network calls. Independent review requested direct pointers to observed roots and each wrong-order commit, plus a worked evidence table; both improvements were incorporated. This is a learner integration and historical inspection exercise, not new execution attestation, human calibration or production authority. Actual model-backed transfer, independent annotation, durable authenticated operations and full curriculum completion remain open.
Verification: the published inspection block passes, and in-memory negative probes reject an incorrect hash, an incorrect expected value and a boolean/integer substitution. The strict book build and whitespace check pass. All four worked solutions open in the local desktop browser; the evidence table renders and page width equals the 1152-pixel viewport. These documentation checks do not claim a new full runtime-test run, mobile/Safari verification, cloud CI execution or publication.
Native telemetry delivery checkpoint — 7 September 2026¶
Katas 83–84 now retain actual OpenTelemetry SDK/OTLP HTTP traffic to a loopback test receiver. Four controls separate independent request accounting, business outcomes, HTTP attempts, unique spans and feedback joins. Capacity loss leaves only the two passing request roots: observed success becomes 100% while the complete business record remains two passes and two failures. A stored span followed by a 503 response causes one real repeated POST; deduplication retains eight unique spans from nine attempts without repeating the measured business action.
The receiver validates original protobuf bytes against a narrow metadata allowlist. Negative tests exercise private strings, exception content, ambient headers/proxies/resource settings, malformed and oversized requests, conflicting identities and bounded idle-connection shutdown. Review found and drove fixes for wrong-trace feedback joins and late socket timeout setup. The selected runtime manifest records actual package versions and two telemetry source hashes, not the full CX source closure or authenticated provenance. SQLite request inventory and receiver storage are in memory. Each study has sixteen instrumented mock business executions; each verification call separately reruns sixteen controls in fresh mock worlds without HTTP.
Verification at implementation/artifact revision bd87b6a: all 546 repository tests passed with 90% overall branch-inclusive coverage and 96% across the telemetry modules. Fresh native execution matches the retained stable projection; the published Python inspection example verifies without HTTP; both recorded source hashes match the local files. Targeted Ruff/Pyright, dependency audit and strict book build passed. Both worked solutions expanded in the local desktop preview without page overflow at 1152 pixels. Current-release composition still reports one hold and six blocks, with deployment authorization false. CI and make test now include the optional telemetry dependencies and a local delivery exercise; this is configuration, not observed cloud CI or application rollout.
Next evidence requirements remain: durable authenticated delivery, per-tool/distributed context propagation, asynchronous queue/backpressure and sampling-loss controls, feedback revisions and authenticated reviewers, and a real authorized exposure path. Actual model-backed architecture studies and independent semantic calibration also remain open. The cross-framework study is separately pending its dependency-security resolution. This checkpoint advances the full book objective without replacing it with a local transport test.
Template-to-gate repair checkpoint — 7 September 2026¶
Kata 82 follows a reproduced reporting gap through the legacy single-order grader, runner and hard gate. Failed template prerequisites previously left the resolution label at resolved and both message counters at zero. The repair makes unsupported template evidence unqualified, preserves a nonresolved diagnosis, and rejects newly registered templates without explicit prerequisite treatment. A passing semantic receipt cannot override this deterministic boundary. Missing evidence is not automatically counted as a proven false statement; a complete supported/refuted/unestablished assertion taxonomy remains future work.
The RED checkpoint is 603ba38, the GREEN repair is dce6f72, and import/type-narrowing cleanup is 3d33fdd. Eleven template regressions include valid ineligibility and all three committed templates, missing policy support, a new unmapped template, a bound passing receipt and runner-to-gate propagation. The published command runs successfully. Independent code and lesson review found no blocking issue; targeted Ruff/Pyright and the main installed-environment dependency audit passed. Both Katas 81 and 82 expanded in the local desktop preview with document width equal to the 1152-pixel viewport. Current-release source verification at 3d33fdd retained one hold and six blocks, with deployment authorization false. No public deployment, Safari/mobile verification, human calibration or paid model evaluation is claimed.
Final verification at 3d33fdd: all 529 repository tests passed, with 90% overall branch-inclusive coverage and 100% for evaluators.py. The strict book build passed. These are software and local evidence-path checks, not estimates of model capability or qualification of the complete application.
The cross-framework dependency review did not find an audit-clean released combination under the inspected requirements: LightEval pins the affected NLTK version, and LM Evaluation Harness requires sqlitedict with an unresolved deserialization advisory. These are dated compatibility findings, not evidence that exploitation occurred. See the LightEval dependency declaration, NLTK advisory, harness dependencies and sqlitedict advisory. The old isolated environment is not being executed or recommended. Next steps are a separately reviewed upstream/custom dependency repair for that study, or independent work on privacy-aware telemetry, actual agent transfer and the remaining qualification requirements. Do not suppress advisories or rename adapter-only scoring as full task-runner reproduction to close the requirement.
Cross-framework diagnostic checkpoint — 7 September 2026¶
Kata 81 adds an executable identity-join and paired-disagreement exercise: both invented harnesses score 3/4 while agreeing on only 2/4 predictions. The solution rejects duplicate and missing identities, preserves order-independent comparison and distinguishes equal aggregates from equivalent measurement. The chapter now explains likelihood normalization, token-boundary differences, replay versus model reruns, and dependency safety versus historical reproducibility.
The published Python solution executed successfully, independent editorial review found no required corrections, and the strict book build passed using python -m mkdocs. The environment's standalone mkdocs launcher contains a stale interpreter path; the module invocation works. These checks do not establish browser rendering, full repository test status or publication.
An isolated feasibility probe completed both native framework pipelines, but matching prompt strings did not reconcile continuation tokenization and likelihoods. Its dependency audit reported unresolved advisories, so execution in that environment stopped. No framework dependencies were added to the book's environment and no historical installation recipe is recommended. The probe is not yet a published, reviewed reproduction packet. Next: review patched compatible dependencies, rerun in a fresh isolated environment, retain native per-item requests and scores, and test an explicit boundary reconciliation without erasing the original mismatch. The broader goal and all outstanding empirical and deployment requirements remain active.
Long-report extraction checkpoint — 7 September 2026¶
Katas 78–80 and the complete casebook retain original/repaired 878/1,020-word memos, nine invented sources, six fixed questions and four registered synthesis obligations. Three controls extract directly from marked finding paragraphs without receiving gold IDs or support labels. Cited-only selection raises the unchanged original's observed support from 8/18 to 8/17 while omitting an unsupported claim. Sentence windows retain compound text without recovering its atoms; the repaired report's 16/16 conditional support still accompanies only 16/20 atomic recovery.
The inputs distinguish supported statements about uncertainty from unknown underlying answers. Independent review drove regressions for overlapping gold spans and disappearing synthesis obligations, plus a corrected temporal citation judgment. Both report texts, source passages and authored review rationales remain inspectable. This is format-bound extraction and review-record computation, not automatic semantic evaluation of all surrounding prose, independent human calibration or research-agent performance.
Verification at source/input revision 9ae4786: 523 repository tests passed with 90% overall and 93% long-report-module branch-inclusive coverage. Fresh CLI output and retained replay agree; the published Python example executes, and both full report texts plus all nine source bodies match the readable casebook. Targeted Ruff/Pyright, the installed-environment dependency audit and strict build passed. All three kata solutions expanded in the local desktop preview without page overflow at 1152 pixels. Current-release composition retained one hold and six blocks, with no deployment authorization. These checks do not claim cloud CI execution, mobile/Safari QA, publication or paid model evaluations.
Next locally actionable work: a pinned cross-framework reproduction with retained per-item outputs and an explicit mismatch reconciliation, or privacy-aware neutral telemetry with export-loss and feedback-join controls. Full-prose extraction validation, real research-agent comparisons, multi-agent/voice transfer and actual application release authority remain open; no narrower checkpoint replaces the full completion target.
Report-evidence checkpoint — 7 September 2026¶
Katas 76–77 add an inspectable compact report, four invented sources, five claim annotations, four citation attempts and a five-question task. Computation separates 3/4 resolving links, 2/4 entailing attempted links, 2/5 citation-supported claims and 1/5 claims with current cited support. A dated authored reference correction exposes a false pass: factual correctness changes from 4/5 to 3/5 without changing the report or corpus. Original and reassessed grades remain linked and inspectable.
The chapter now teaches exact span contracts, explicit citation-marker inventory, reference repair versus changed task dates, information-access cutoffs, and omission versus justified abstention. Independent review exposed a denominator failure when a broken citation's annotation was removed but its marker remained in the report; the implementation adds a regression for complete marker accounting. Semantic annotations and reviewer authority remain invented classroom inputs, not authenticated human evidence.
Verification at implementation revision 0bcb85b: 511 repository tests passed, with 89% overall and 95% report-module branch-inclusive coverage. Fresh CLI output, retained replay, the printed report and published Python example agree. Targeted Ruff/Pyright, dependency audit, strict book build, built anchors/artifact links and both expanded kata solutions in the local desktop browser passed; page width matched the 1152-pixel viewport. The composed current-release study retained one hold and six blocks with deployment authorization false. No live models, independent human review, mobile/Safari verification, cloud CI execution or publication is claimed by this checkpoint.
This completes a compact evidence-mechanics exercise, not the larger long-report validation requirement. Next: retain a substantial multi-section report and corpus with independently checkable extraction omissions, compound claims, conflicting sources and synthesis judgments; compare original and corrected reports without changing the evaluation contract. Actual research-agent runs, independent annotation validity, protected evidence provenance and production release authority remain separate requirements.
Verified judge-diagnostics checkpoint — 7 September 2026¶
Katas 74–75 now connect controlled transformations to retained judge requests, returned verdicts, identity-normalized comparisons and ensemble errors. Five deterministic controls execute 300 calls over 60 presentations of five authored cases. The first-slot control reverses answer identity in all 30 order contrasts; the length control changes in 20 of 40 padding contrasts; the keyword control changes in all 30 rubric contrasts. The three-member panel inherits the two cloned controls' same 48 erroneous views. Identical-content answer swaps diagnose arbitrary protocol selection, not necessarily a change in factual meaning.
Verification at implementation revision e426a8f: 500 repository tests passed with 89% branch-inclusive coverage; the retained study exactly matches fresh execution and replay; the chapter's Python examples execute successfully. Targeted Ruff and Pyright checks, the installed-environment dependency audit and strict book build passed. The current-release composition reports one diagnostic hold and six blocks, with deployment authorization false. The workflow now includes the study command; this is not evidence that cloud CI has run.
These controls teach how to investigate judge sensitivity; they do not measure live-model bias, validate human labels or qualify a production evaluator. Independent invariance review, repeated model calls and a metered pairwise adapter remain necessary for that experiment. The next locally executable depth gap is a complete frozen research report and source corpus with claim/citation spans, dated reference reassessment and preserved original/new grades. Existing content and earlier checkpoints remain intact.
Planning checkpoint — 7 September 2026¶
This checkpoint responds to the follow-up audit; it does not certify implementation, live-model qualification or deployment. Static inspection confirms that the local source already contains template factual prerequisites, context-bound semantic receipts, an optional semantic stage in the runner, paired artifact references and illustrative point-gain terminology. Reproduce their acceptance tests before closing the audit findings. Earlier delivery notes remain historical context, not an automatically current status register.
The next bounded delivery sequence is:
- Close correctness and trust findings. Reproduce both false-template probes, exercise the identity/eligibility/approval/escalation state matrix, and reject receipt reuse after any judged-context change. Verify that every accepted template has explicit factual prerequisites. Validate realistic resolution mode separately from the target-supplied argument-handling mode. Acceptance: wrong-but-authorized order actions fail independent grading, ambiguity can require clarification, and changing an evaluator-only target does not alter agent-visible inputs.
- Prove one complete evidence packet. From a clean checkout, reconstruct grades from retained requests, tool events, final states, prose, raw judge responses, receipts, usage and manifests. Preserve original grades; new grading produces linked reassessment artifacts. Check missing evidence, coherently rewritten records, current revocation and expiry. Acceptance: a reviewer can explain every grade and release restriction without trusting stored pass flags. Hash consistency alone is not authenticated provenance.
- Qualify semantic grading empirically. Freeze a criterion-specific judge and rubric; use independently reviewed, separated development/calibration/acceptance data. Recompute false-pass, false-block and abstention bounds from retained rows. Acceptance: runner integration works on actual approved model calls, qualification is restricted to supported slices, and unknown or out-of-scope evidence holds rather than passes. Model selection, budget and independent reviewers are dependencies, not assumed resources.
- Demonstrate transfer incrementally. Preserve deterministic knowledge-to-action and process-recovery controls, then run approved model-backed versions with registered repetitions and authoritative outcomes. Follow with matched-budget single/multi-agent traces and a compact coding or artifact domain. Voice requires real audio, synchronized events and independent annotations. Acceptance: observations are extracted from executions, not supplied failure labels; all arms and failed trials remain inspectable.
- Qualify decisions and risk controls. Extend known-population statistical studies to sparse discordance, unequal clusters, sequential looks, label error and holdout reuse. Recheck containment and cross-run isolation against independent phase, worker and event-order expectations. Organize frontier claims around capability, practical uplift, propensity, safeguard performance, human effects and residual uncertainty. Acceptance: state exactly which local mechanisms were demonstrated and which deployment or research claims remain unsupported.
- Integrate the book and verify publication separately. Each package contributes an example, counterexample, runnable exercise where appropriate, worked solution, interview prompt and scoped evidence label. Use one main navigation route: measurement → evidence → justified decision. Acceptance: requirement-by-requirement source coverage review, clean rerun of one end-to-end packet, strict build and desktop/mobile render checks. Push, publication and actual application rollout require their own verification; this planning checkpoint authorizes none of them.
First milestone: packages 1–2 produce one independently inspectable CX packet. Do not expand the subject inventory to compensate for missing empirical evidence. Local regression and replay work is ready for implementation; paid studies, independent calibration, human-participant research and real application exposure have separate approval and resource dependencies.
Every substantive subject needs: key points; assumptions; a small worked example; a failure or counterexample; a micro-kata where implementation is appropriate; a worked solution; an interview question with answer criteria; source provenance; and the evidence supporting its claims. Concepts involving humans or real deployments also identify the study or field evidence that cannot be replaced by synthetic fixtures.
For executable work, completion requires an inspectable chain from input through execution, grading, uncertainty, and the resulting decision. For open research, completion means an accurate account of methods, conflicting evidence, and unresolved uncertainty—not a fabricated solution.
Audit repair sequence¶
| Requirement | Current evidence | Next acceptance milestone |
|---|---|---|
| Factual template prerequisites | Kata 82 rejects verified-account, ineligible-approval and false execution-failure explanations, marks failed prerequisite evidence unqualified through the hard gate, and rejects unmapped template rules | Expand the state/wording cross-product, including revoked and denied approvals; distinguish refuted from unestablished assertions explicitly |
| Context-bound semantic judgments | Exact case/request/trace/state/policy binding; registry checks configuration, joint scope, expiry, revocation, error bounds and abstention; row-derived annotation compiler; separate anchored replay/current-calibration API and Katas 32–33 | Authenticated independently reviewed calibration data, authoritative registry retrieval and integration with release decisions |
| Honest scalar gain terminology | illustrative_point_gain:task_success regression |
Preserve distinction in every result and exercise |
| Complete trial evidence | Full paired artifacts and offline re-grading CLI; optional source check verifies full committed Python inventory, lock/declaration files, dataset/policy bytes and full retained case inputs; retained 20-trial example and Katas 34–35; separate current-calibration assessment | Authenticated execution provenance, installed-environment verification, outer decision recomputation, integrated source/current-calibration checks and immutable reassessment under new graders |
| Semantic stage in runner | Optional evaluator-owned stage and Responses adapter; calibration ingestion (Katas 18–20, 36–37); durable judge admission and cost accounting (Katas 38–40); pinned campaign checks now join trial evidence and restrict lab release receipts, with retained CI controls and Katas 41–42 | Authenticated/fresh snapshot and current-calibration composition, validated provider price bounds, agent/tool budgets, cancellation/reconciliation and independently reviewed live calibration; no invoice cap or live semantic qualification claimed |
| Realistic object resolution | Executed 16-trial multi-order study and Katas 21–23; native paired extension adds 32 agent executions, full-case/design registration, allowlisted mock-tool replay and blocked release receipts (Katas 43–44). Native semantic extension adds context-bound receipts, row-derived synthetic calibration and current revocation (Katas 45–46); native provider/campaign integration has installed-SDK/mock-HTTP controls (Katas 47–48) | Independently reviewed calibration; live model comparison, independent ambiguity cases, human/simulator validity and integrated current qualification/source checks |
| Statistical qualification | Executed sparse-discordance and sequential path enumeration; fixed-final/repeated-fixed/Bonferroni/likelihood-ratio comparisons, label/dependence counterexamples and Katas 28–31 | General clustered sequential methods, calibration-error propagation, registered live observation collection and enforced repeated-holdout controls |
| Process recovery experiment | Executed 21-trial child-process/SQLite study: commit/checkpoint crashes, revocation, uninterrupted control and duplicate-producing mutant | Transfer to actual agent tool execution, compaction, concurrency, full approval lifecycle and notification recovery |
| Architecture experiments | Hand-authored scorer fixtures exist | Executed retrieval/oracle/full-context arms; delegation/merge traces; matched budgets |
| Voice validity | Timeline scorer fixture exists | Recorded audio, synchronized action events and independent timing annotations |
| Second domain | Domain comparison prose exists | Coding or artifact task with independent authoritative outcome and reproducible failure |
| Frontier risk bridge | Source-verified worked safety case and Katas 13–15 separate capability, uplift, propensity, controls, internal deployment, human outcomes and useful autonomy | Execute local monitoring/isolation protocols; obtain relevant empirical transfer evidence; retain unresolved research limits |
| Operational release curriculum | Runnable offline CI conformance and book-publication dependency; Katas 16–17; executed 280-request exposure simulation with 349 agent runs, pending-cohort checks, expiry, restriction and rollback; Katas 24–27 | Observe GitHub workflow execution; connect qualified evidence to a real application router and independent stop path; implement hard action budgets and incident-to-dataset automation |
Whole-curriculum coverage requirements¶
| Subject | Teaching home | Required practice/evidence before completion |
|---|---|---|
| Foundations and seven surfaces | Foundations; Agent & System Evals | Turn a product promise into separate outcome, policy, tools, factuality, conversation, efficiency and operational checks |
| Dataset roles and continuous improvement | Dataset Design | Split development/calibration/sealed data; detect leakage; add reviewed incident regression; preserve versions and provenance |
| Metrics and statistics | Metrics; Evidence Spine | Recompute slices, uncertainty, reliability, probability calibration, missingness and rare-event limits |
| Human evaluation | Human Evaluation | Rubric pilot, blinded overlap, disagreement analysis, adjudication and annotation budget |
| Judge calibration | LLM as a Judge | Frozen labels, false-pass/false-block uncertainty, abstention, scope restriction, drift and requalification |
| Safety, fairness and robustness | Robustness, Safety & Fairness | Perturbations, adversarial tests, crossed slices and false refusal; justified limits of claims |
| Retrieval and research | RAG & Research Evals | Retrieve then act; verify atomic claims, citation support, freshness and completeness |
| Tools, agents and state | Agent & System Evals; Long-Running & Serving Failures | Object resolution, authorization, approval, idempotency, ambiguous commit, compaction and process restart |
| Benchmark validity and integrity | Benchmark Reproducibility; Eval Operations & Integrity | Contamination probes, scorer isolation, harness changes, resource noise and cross-run communication |
| Modern architectures | Modern Agent Architectures | Skill selection versus execution; actual multi-agent traces; voice timing; second-domain evaluation |
| Evaluation infrastructure | Evidence Spine; Eval Operations & Integrity | Scheduling, quotas, resumability, immutable evidence, trace joins, reruns, cost attribution and flaky-task handling |
| CI/CD and release gates | Build Lab; Release Gates & Shipping | Independent required rules, missing-evidence blocks, pinned candidate/baseline, exceptions, rollback target and deployment identity |
| Production learning and exposure | Production Evals | Shadow, canary, mature outcomes, guard bands, expand/restrict/rollback, human burden and incident promotion |
| Frontier OpenAI/Anthropic methodology | Research-to-Practice; References; System Studies | Source-grounded comparison of elicitation, auditing, safeguards, monitoring, internal deployment and unresolved risk science |
| Interview readiness | Interview & Design Drills; Micro-katas | Explain key points, solve unseen variants, defend evidence and diagnose a misleading release packet |
Study status and proof¶
The micro-katas include runnable grader regressions and complete trial replay. Remaining rows above are a delivery backlog, not a completion claim. No automatic score or chapter count establishes an A+ primer or interview readiness. Final review must inspect examples, solutions, source accuracy, execution artifacts, rendered pages, and transfer to unseen tasks.
Delivery order¶
- Repair factual grading and semantic context binding; teach those repairs as katas.
- Preserve full executions and implement packet validation/re-grading.
- Connect semantic evaluation and qualify its scope with reviewed evidence; improve object resolution.
- Run statistical method experiments and process recovery; execute retrieval comparisons.
- Add multi-agent, second-domain and audio evidence; implement CI/CD and exposure-control katas.
- Deepen frontier-risk methodology, finish subject-specific katas/solutions and interview prompts, then perform requirement-by-requirement technical and rendered-book review.
Paid model studies require an explicit model and metered run configuration. Human calibration and participant outcomes require actual independent review. Until those exist, publish methods and synthetic demonstrations with their limits clearly identified.
Follow-up audit remediation plan — 6 September 2026¶
Current execution handoff — current release checkpoint 96c76da¶
Rater-analysis checkpoint — 7 September: Katas 72–73 analyze the existing 24 synthetic annotations without replacing the source fixture or its four adjudication reasons. The retained study includes pairwise confusion/marginal counts, unknown/missing exclusions, the binary-decidable projection, an assumption-labelled agreement interval and exact empirical item-bootstrap sensitivity. R1/R2 agreement is 6/8 with κ = 25/41; excluding unknown judgments leaves 5/6 agreement on only six items. Bootstrap enumeration retains undefined κ probability 3409/8388608 and conditional finite-value percentiles 5/53 through 1; these are not a coverage-qualified κ interval.
A separate four-case assignment control holds baseline/candidate content identical. Assigning strict labels only to baseline and lenient labels only to candidate creates a +50-point apparent gain; crossing both reviewers over both arms yields zero, with each reviewer's paired difference also zero. This is an authored negative control, not an experiment on human reviewers. The analysis preserves absent individual rationales as absent and keeps authored adjudication distinct from independent validation. Next depth work should address judge order/verbosity/rubric robustness and complete report/citation evidence; real reviewer sampling, independently qualified labels and model-backed transfer remain outstanding.
Verification at source checkpoint 9d3d486: 487 tests passed with 89% branch-inclusive overall coverage and 95% for rater analysis. The CLI, retained artifact and complete re-execution agree; the published missing-label example passed, and the canonical source fixture is unchanged. Independent review enumerated item-index multisets separately from the implementation's cell-count algorithm and reproduced the complete bootstrap distribution, undefined mass and conditional percentiles. Targeted Ruff/Pyright, installed-environment dependency audit and strict book build passed. Both worked solutions rendered without horizontal overflow at the checked 1152-pixel desktop viewport. The current-release composition retained one diagnostic hold and six restrictive blocks. Cloud CI, mobile/Safari, actual human annotation, live model evaluation and deployment were not verified.
Sampling-estimation checkpoint — 7 September: Katas 70–71 now retain a fixed-frame sample-to-estimate comparison. The sampling artifact binds frame, allocations, selected worked-example rows, inclusion probabilities, exact interval endpoints, weighted count-cell distributions and both registered thresholds. The worked risk-enriched sample gives 37.5% raw failures, an 18.75% population-weighted estimate and a conservative interval of 18.75%–56.25%; the interval holds at the illustrative 35% limit. Exact enumeration over 5,544 proportional and 495 enriched sample-ID sets is compressed by outcome multiplicity, not described as new model executions.
Independent count-population qualification covers all 65 combinations of routine/risk failure counts for each allocation, with minimum coverage 131/132 and 98/99. Missing labels, unknown or inconsistent probabilities, incomplete support and duplicate/foreign IDs reject. Exact fraction comparisons prevent a rounded census endpoint from changing a threshold verdict. The estimator sees selected labels only; synthetic full truth belongs to the separate method study. This closes a fixed-design teaching pipeline, not adaptive sampling, actual production prevalence, label validity or deployment authority. Next depth work should address rater effects and judge-design robustness, followed by complete report/citation evidence and independently qualified model-backed transfer.
Verification at source checkpoint 7b96f8c: 474 tests passed with 89% branch-inclusive overall coverage and 95% for the sampling module. The published CLI exactly reproduced the retained artifact, and the published sample-estimation/missing-label block passed. Independent review separately derived all 65 coverage records per design and enumerated actual sample-ID combinations. Targeted Ruff/Pyright, installed-environment dependency audit and strict book build passed. Both worked solutions and mathematical notation rendered without horizontal overflow at the checked 1152-pixel desktop viewport. The integrated current-release study retained one diagnostic hold and six restrictive blocks. Cloud CI, mobile/Safari, live model evaluation and deployment were not verified.
Context-packing checkpoint — 7 September: Katas 68–69 now retain an executed five-passage, two-policy-unit comparison across four evidence arms, two downstream controls and three constructed cases. The 24-trial packet includes complete rankings, selection/exclusion records, proposed packing sizes, exact agent-visible JSON bodies, decisions, mock tool attempts and final state. With a 240-byte body budget, top-2 supplies one evidence unit in 141 bytes; top-5 retrieves both but recency-first packing supplies only the promotion in 191 bytes. Metadata filtering and exact deduplication retain both units in 136 bytes. The repaired correct control passes all three cases; the ignore-limit control still makes one backend-blocked invalid request. A fixed-ranking-order component counterfactual preserves earlier packed evidence, keeping the top-k attribution precise.
This closes local multi-evidence packing mechanics, not natural-language interpretation, actual model-token budgeting or source-authority authentication. The controls parse structured facts; the backend computes immutable mock state independently of the decision; no real payment or external service is involved. The filtering/deduplication arm combines two changes, so separate causal effects remain unmeasured. Next depth work should address representative/risk-enriched sampling and weighted estimates, followed by quantitative human/judge robustness studies and full-report claim/citation evidence. Independent labels, approved live-model budgets and real deployment authority remain separate prerequisites.
Verification at source checkpoint 1976501: 462 tests passed with 89% branch-inclusive overall coverage and 93% for context packing. The published CLI exactly reproduced the retained 24-trial artifact, and the published Python inspection block passed. Independent review checked the byte counts, prefix-order counterfactual, backend mutant and replay rejection controls. Targeted Ruff/Pyright, installed-environment dependency audit, strict book build and rendered study links passed. Both worked solutions opened without horizontal overflow at the checked 1152-pixel desktop viewport. The integrated current-release study reproduced one diagnostic hold and six restrictive blocks at the fixed committed revision; an earlier run spanning a source change correctly blocked source verification. Cloud CI, mobile/Safari, live model evaluation and deployment were not verified.
Mixed-trace analysis checkpoint — 7 September: Katas 66–67 now bridge raw trace inspection and reviewed regression creation. The retained workshop presents eight learner views and a separate instructor key, derived from the pinned native archive through full replay. Five traces fail, three pass, and six category assignments overlap across three cases and one customer. The lesson distinguishes C's correct-but-unclarified refund, H's safe-but-unfinished task, and the settlement claim's missing external support. A versioned authored taxonomy records inclusion/exclusion criteria, hypotheses, source-code references and discriminating checks; matched baseline references do not imply a newly executed or isolated causal intervention.
This closes a worked mixed-trace teaching method, not independent human error discovery or production prevalence estimation. Original source packets remain unchanged; no new agent executions or human annotations are claimed. Next depth work should address multi-evidence retrieval/context packing and representative sampling, while keeping independent judge/human qualification, model-backed transfer and real exposure control as separate requirements.
Verification at source checkpoint 5e02a16: 451 tests passed with 89% branch-inclusive overall coverage and 83% for the workshop module. The retained artifact rederives exactly; the seven focused tests check source/trust rejection, learner-view separation, overlapping counts and matched references. Independent review checked the lesson's observation-versus-cause distinction. Targeted Ruff/Pyright, dependency audit, the published CLI and Python block, and strict build passed. Both worked solutions opened without horizontal overflow at the checked 1152-pixel desktop viewport. The current-release composition rerun retained one diagnostic hold and six restrictive blocks. Cloud CI, mobile/Safari, live model evaluation and deployment were not verified.
Native slice checkpoint — 7 September: Katas 64–65 now derive structural and joint change reports from the original native semantic packets. The retained derivative references pinned source bytes and packet/receipt hashes rather than copying or rewriting the original executions. It replays 48 retained trials across three comparisons; it creates no new agent observations. The false-settlement control has eight known joint regressions despite eight structural passes. The first-record control has six structural/joint regressions, including every clarification-required pair. Tests distinguish qualified failures from abstention/missing qualification and prevent the structural projection from inheriting joint inference authority. Slice labels are retrospective; required status needs a separately registered plan and new experimental evidence.
The native adapter supports both structural-only v1 and semantic v2 packets. It preserves known structural defects when joint qualification is unknown. All four source cases share one customer; all derived diagnostics hold, while original release receipts remain blocked. The source/calibration pins are explicit pedagogical trust choices, not authenticated execution or current qualification. This closes local native slice-report integration, not empirical validation or deployment readiness. Next locally feasible depth work includes mixed-trace error analysis and context packing; independent human/model studies and real application exposure retain separate dependencies.
Verification at source checkpoint b2a36af: 444 tests passed with 89% branch-inclusive overall coverage, 98% for the native adapter and 97% for its retrospective driver. Independent review reproduced exact retained-report derivation, rejected coherently rewritten case evidence and wrong source pins, and checked the lesson's pair/slice counts. Targeted Ruff/Pyright, dependency audit, the published CLI and Python inspection block, and strict build passed. Both solutions opened without horizontal overflow at a 1152-pixel desktop viewport. The current-release composition rerun retained one diagnostic hold and six restrictive blocks. Original source packets were unchanged. Cloud CI, mobile/Safari, live model evaluation and deployment were not verified; CI is configured for local-equivalent conformance only.
Paired slice checkpoint — 7 September: Katas 62–63 turn complete legacy paired artifacts into fixed/regressed case-repetition lists, required/exploratory slice rows and distinct trial/customer-weighted comparisons. The retained study executes 16 deterministic mock-tool runs: six pairs improve, two regress, and the overall +50-point trial-weighted change differs from the +33.3-point equal-customer change. The one-customer required slice regresses in both repetitions; it has no inferential interval. Missing support and unqualified prose hold the diagnostic without erasing known failed checks. All measurements are synthetic; this is not live-model or deployment evidence.
This implements the Code-First report requirement for the legacy schema. Native multi-order support, authenticated registration/provenance, representative observations, qualified inference and multiplicity control remain open. The next locally feasible source-depth gaps are mixed-trace error analysis and multi-evidence context packing; independent human/judge validation and external-framework reproduction retain their own prerequisites. The dated source audit records the six Markdown reports inspected and does not claim a new reread of every PDF/DOCX.
Verification at source checkpoint c50ca89: 429 tests passed with 88% branch-inclusive overall coverage and 89% for the slice module. Review reproduced and repaired coherent full-case/agent-input substitutions against an unchanged registered dataset, boolean trial-index acceptance and single-customer interval construction. The retained packet still matches exact re-execution. Targeted Ruff/Pyright, dependency audit, both published Python blocks and strict book build passed. The current-release composition rerun retained one diagnostic hold and six restrictive blocks. Both solutions opened without horizontal overflow at the checked 1152-pixel desktop viewport. Cloud CI, mobile/Safari rendering, live-model evaluation and deployment were not verified; the new CI command is configured only.
Incident-to-regression checkpoint — 7 September: Katas 60–61 now connect an executed synthetic failure, a fixed causal-preserving minimization, exact review binding, regression-only version publication and reruns materialized from the released case. The retained packet reproduces the reference's one commit, the mutant's two commits and seven rejected promotion controls. Dedicated tests preserve populated parent entries, recompute their case identities and check current protected inventories against both new proposals and carried-forward parents. This closes local incident-promotion mechanics, not actual human adjudication, authenticated reviewer/parent authority, automatic de-identification, complete lineage discovery or representative production sampling.
Verification: 412 tests passed with 88% overall branch-inclusive coverage and 87% for the incident module; 11 focused incident/artifact tests passed again after cosmetic lint changes. The retained replay snippet, strict book build, targeted Ruff/Pyright and dependency audit passed. Both worked solutions opened at the checked desktop viewport without horizontal overflow. CI execution is configured, not observed in the cloud; mobile/Safari rendering and deployment were not verified in this checkpoint.
Next delivery work should audit the supplied-source and learner requirements across the whole curriculum, then address the highest-value missing executable transfer: real delegation/merge traces under matched total budgets, a coding/artifact domain, and skill selection separated from execution. Independent calibration, approved model-call budgets, synchronized audio and human annotations remain prerequisites for the corresponding empirical claims. Continue to preserve existing content and distinguish local demonstrations from qualified production evidence.
Integrated walkthrough checkpoint — 7 September: Katas 58–59 now trace one ambiguous refund from agent-visible request through clarification, tool calls, every order ledger, semantic evidence, paired artifact join and current decision. Both inspection blocks executed successfully against retained v2 evidence. A clean local clone of 4ed96a7 with a fresh locked environment reproduced the current-release study: 16 replayed passing trials, one inconclusive customer cluster, one diagnostic hold and six restrictive blocks. The original five-case route remains available; navigation now separates it from the native joined study and from optional paid inference. Legacy/native counters and file/value source mappings are explicitly distinguished. This closes the local learner-path and clean-checkout reproduction milestone, not independent calibration or live-model qualification.
Remaining priorities: replace synthetic qualification inputs with independently reviewed evidence when reviewers and an approved model budget are available; extend actual agent execution into the existing retrieval/recovery studies; complete multi-agent, coding/artifact and voice transfer studies; and connect approved evidence to an independently controlled real application deployment. Local work can also advance reviewed incident-to-regression ingestion, sealed-data controls and a full source-to-curriculum depth audit while those external dependencies remain open. Do not manufacture human review or inference results to close these rows.
Cross-run isolation checkpoint — 7 September: Katas 56–57 now retain six executed cache comparisons with 18 responses and 55 operations. Query-only keys leak across all three worker lifecycles; scoped keys preserve the registered private boundary and allowed sharing. Regrading independently checks phase/run/worker assignments, concurrent thread joins and A-write-before-B-read ordering. Dedicated regressions reject coherently rehashed lifecycle and ordering substitutions, and returned reports cannot mutate the canonical policy through a shared reference. This supersedes the later paragraph's pending status for the cooperative local cache study only. Hostile-worker sandboxing, external evidence authority, other communication channels, real-agent behavior and deployment qualification remain open.
The next implementation priority is to consolidate the current end-to-end CX evidence path and its learner walkthrough, rather than add another isolated scorer fixture. Validate a clean-checkout packet against the correctness, evidence-ownership and realistic-resolution requirements above, identifying the exact boundary at which independently reviewed calibration and approved live model calls are still needed. Then extend actual agent execution into the retrieval/recovery studies and the remaining multi-agent, second-domain and voice protocols. Existing local control results remain regression evidence, not substitutes for that transfer work.
Local verification for this checkpoint: 401 tests passed with 88% branch-inclusive aggregate coverage (82% for the isolation module); 11 focused isolation/artifact tests passed again after import normalization. Strict MkDocs build, targeted Ruff/Pyright and dependency audit passed. Both new worked solutions opened in the local browser without horizontal overflow at a 1152-pixel desktop viewport. Mobile/Safari rendering, cloud CI and production publication were not verified in this checkpoint.
Katas 49–50 now join real committed-source verification, independently selected file/value inputs, historical replay, current calibration, campaign accounting and the base release receipt in a recomputed point-in-time decision. The retained seven-control study holds the current synthetic diagnostic and blocks all six restrictive variants. Native judgment/campaign timestamps cannot postdate the proposed decision, and coherently hashed non-finite values are rejected. This supersedes earlier rows that list local source/current-calibration composition as wholly pending. Independent human calibration, authenticated registry and prerequisite provenance, loaded-environment attestation, new-grader reassessment and atomic handoff to an exposure controller remain open.
The next content priority is an executed knowledge-to-action comparison: retrieval, oracle-document and full-context arms with authoritative final-state outcomes and extracted execution observations. Keep deterministic controls distinct from a later budgeted model study. In parallel with that learning path, prepare representative, independently reviewed calibration and untouched acceptance cases; synthetic calibration must not be promoted merely because integration tests pass. Frontier containment/isolation experiments and a second coding/artifact domain remain required parts of the full objective.
Knowledge-to-action checkpoint: Katas 51–53 now execute that comparison as 18 deterministic local trials, with preregistered input/control hashes, genuine lexical rankings, fresh mock ledgers and re-extracted outcomes. Oracle/full-context interventions repair access failures but not the wrong-amount mutant. The control reads structured policy fields and receives the active version; it does not test natural-language interpretation or version discovery. The next empirical step for this path is representative held-out cases and a separately budgeted model adapter with qualified semantic observations, while preserving these deterministic regression controls. This supersedes the prior paragraph's pending status for the local control comparison only—not for the model study or full architecture coverage.
Containment checkpoint: Katas 54–55 now execute six barrier-ordered local process controls. They demonstrate that a coordinator cancellation acknowledgement does not revoke a queued child's write authority, that late revocation preserves earlier effects, and that a correctly enforced false stop leaves legitimate work undone. Actual initial/final SQLite state and worker/coordinator identities are retained; timeout tests verify child cleanup. This closes a local cooperative-enforcement example, not threat-detector qualification, distributed propagation, hostile-worker isolation, human recovery burden or live deployment. The next frontier-control study should test cross-run isolation with synthetic markers and explicit allowed-sharing policies; keep it distinct from planned multi-agent cooperation.
This handoff supersedes the work ordering, not the retained historical baseline and phase descriptions below. It responds to the follow-up audit against the current local source. Local implementation is not evidence that the public site has received these changes. No new live evaluation, deployment, or human qualification is claimed by this plan.
The audit's template prerequisites, context binding, full paired artifacts, optional semantic runner stage and illustrative superiority naming now have local implementations. A separate unresolved-order runtime path and an executed deterministic multi-order study also exist. The original explicit-order runtime remains useful for argument-handling tests; it must not be scored as an object-resolution experiment. Calibration ingestion, sequential counterexamples, process recovery and exposure control have local worked examples, with the limitations recorded in the status table above. Recheck these implementations rather than rebuilding them from the audit's older snapshot.
Since the earlier 10249c0 handoff, native multi-order paired execution and semantic grading have been joined in an offline control study. Katas 45–46 retain four row-derived synthetic calibration executions and three sixteen-execution comparisons, with joint candidate passes 0/8, 2/8 and 8/8 (checkpoint e02be32). Full artifacts replay with explicitly supplied historical fixture trust; revocation is assessed separately, and every release remains blocked. Katas 47–48 additionally connect the native criterion to the installed provider SDK and direct campaign wrapper using mocked HTTP, with clear/hold/block accounting controls and retained provider evidence. Next, prepare independently reviewed calibration and untouched cases, and join current qualification/source checks to the release path. Synthetic reviewer IDs, overlapping controls and loose diagnostic thresholds do not meet the empirical acceptance gates below; agent budgets, provider cancellation and actual deployment remain separate work.
Next deliverable: one independently inspectable CX experiment, connecting unresolved customer intent, real execution, reviewed grading, immutable evidence and a bounded decision. Separate lab results must not be presented as this integrated experiment until they actually share a verified evidence chain.
| Work package / dependency | Implementation and book change | Exit test / retained proof |
|---|---|---|
| A. Close the trust boundary — first | Extend models.py / evaluators.py prerequisite matrices; harden artifacts.py replay with source/grader identity and an externally owned current qualification registry. Define historical reproduction separately from present-day qualification and new-grader reassessment. |
Both reported false passes and cross-context receipt reuse fail. Changed source, unknown grader, revoked qualification and missing evidence cannot establish current authority. Historical evidence remains inspectable after revocation. A new grade links the unchanged original artifact, its grader/configuration and the reason for reassessment. |
| B. Complete the semantic evaluator — builds on A; annotation preparation can start independently | Extend semantic.py and calibration_data.py with a metered provider adapter, deadline/cancellation, bounded retries and operator-owned registry admission. Prepare a blinded annotation protocol and versioned development/calibration/sealed splits. Preserve reviewer provenance and disagreements. |
Tests first use fake provider responses to exercise timeout, malformed output, budget exhaustion and abstention. Then independently reviewed labels reproduce registered error bounds; frozen configuration is assessed on untouched cases in scope. An allowlisted reviewer name or a dataset hash alone is not proof of independent human review. |
| C. Join the complete CX packet — A + B | Connect the unresolved-order path to paired runner artifacts and qualified semantic grading. Register cases, repeated trials, baseline/candidate versions, ordering randomization, stopping rule and decision method before execution. Include wrong-order, unnecessary escalation, contradictory prose and clarification-cost outcomes. | A clean rerun command reconstructs grades and aggregate decisions from retained evidence. Each trial preserves model-visible input, full tool results, final state, customer prose, judge evidence, response/usage IDs, timing, costs and manifest. A held-out ambiguity set tests transfer. Small or unqualified evidence yields a hold, not a deployment claim. |
| D. Demonstrate execution-to-observation extraction — reuse C | First run knowledge-to-action arms: ordinary retrieval, oracle documents and full context. Then extend process recovery to actual agent tools with commit-before-checkpoint interruption, compaction and revoked approvals. Follow with matched-budget single/multi-agent execution and a second coding or artifact domain. | Scorers derive observations from actual traces and authoritative state, not supplied failure flags. Publish failures, repeated trials, paired differences, extraction validation and resource accounting. Recovery checks constraints, duplicate side effects and unfinished work after resumption. Architecture and model changes are isolated or explicitly confounded. |
| E. Qualify decision methods and controls — can run offline alongside A–C | Extend known-population studies to unequal clusters, sparse/all-tie outcomes, label error and sequential decisions. Enforce holdout-access history and campaign budgets. Connect accepted decision receipts to simulated exposure control, then separately to a specified application environment. | Register numerical acceptance criteria before simulations; retain coverage, false promotion, power, Monte Carlo uncertainty where applicable and invalid-method counterexamples. A sample-count threshold alone never qualifies a method. Integration rejects stale or mismatched packets and demonstrates restriction, hard-stop latency, in-flight exposure and rollback identity. |
| F. Deepen frontier-risk evidence — independent source review; reuse artifact integrity from A | Verify first-party sources at implementation time. Extend the risk chapter with benign cross-run isolation and monitor-interruption experiments. Keep capability, practical uplift, harmful propensity, safeguards, exposure and residual uncertainty as separate claims. Include internal AI-development deployments, useful-reliability calculations and human-study protocols. | External authoritative logs expose unintended communication or tampering attempts. Measure completed actions before containment, false stops, recovery and behavior after denial. Human effects require ethically appropriate participant evidence; simulated transcript ratings cannot substitute. Mark unresolved field science separately from missing local engineering. |
| G. Integrate teaching and publication — accompanies each package; final review after C–F | Update existing chapters rather than replacing their breadth. For each package add a worked trace, failure case, micro-kata, solution, interview answer criteria and evidence-status label. Provide a single learner route linking measurement → evidence → justified decision. | Audit every supplied-source coverage row and substantive subject. Independently rerun a complete packet, check links and desktop/mobile rendering, then verify the published commit and visible page content match the reviewed source. Publishing remains a separately authorized action. |
Experimental design details that must not be deferred¶
- Evidence ownership: the evaluation service controls the manifest and artifact store; a separate authorized operator controls registry admission and reviewer identities. Hash consistency is not authenticated provenance. Test a coherently rewritten packet and candidate attempts to modify evaluator files, not only accidental single-field corruption. Reproduce the integrated packet in a clean checkout with documented prerequisites.
- Object resolution: candidate records may legitimately expose order IDs. Hide the evaluator's target designation, expected answer and ordering hints, not every occurrence of the correct ID. Changing only the evaluator label must leave model-visible inputs unchanged. Server authorization blocks foreign objects; independent grading must catch a wrong-but-authorized order action.
- Registration: freeze primary outcomes, independent experimental units, repetitions and sample rationale, budget, missingness policy, uncertainty method and decision rule before inspecting study results. Retain protocol deviations explicitly.
- Calibration: define the unit of independence, reviewer assignment, overlap, adjudication, error tolerances and scope before collecting labels. Report uncertainty and disagreement; keep tuning data outside final qualification. Revocation and expiry affect present authority without erasing the original record.
- Architecture comparisons: declare what stays fixed, what changes and total resource budgets, including coordinator, subagent, retrieval and judge costs. Oracle documents test an intervention, not a production retrieval strategy. Retain all registered arms, including unsuccessful ones.
- Recovery: specify interruption points and an external side-effect oracle. Restarting a process successfully is not proof of preserved approvals, constraints or exactly-once effects.
- Voice: schedule a distinct study only once audio, timestamp synchronization and independent annotations are available. Define timing tolerance, interruption boundaries and premature-action ground truth; text-only fixtures remain scorer tests.
- Release: separate application authorization from book CI. Define an operator, rollback target, immutable receipt schema, evidence freshness, hard action budgets and an independent stop mechanism before connecting to real traffic. A green local exercise grants none of these powers.
Decisions needed only for externally dependent work¶
Local trust fixes, fake-provider integration, statistical studies, teaching improvements and protocol design can proceed without paid calls. Before live studies, obtain agent/judge model selections, a total spend cap and an approved dataset; before human qualification, obtain independent reviewers and the annotation protocol. Real exposure control needs a named application environment and release authority. Audio and participant studies need suitable data permissions and study review. These dependencies do not block the remaining local work.
The first acceptance checkpoint is A complete plus B's offline integration, followed by the reviewed calibration and integrated packet in C. Do not widen the architecture inventory merely to compensate for an incomplete packet. Report completion per claim: explained, implemented, experimentally demonstrated, or qualified for a specifically named use.
Capability¶
Enable a reader to independently inspect an execution, reconstruct its grade, assess the evaluator's qualification, and explain the bounded release or risk decision it supports. The next revision prioritizes complete experiments over more topic inventory. This is a plan, not a claim that the remaining work has been implemented or deployed.
The local baseline at e53c53d already contains factual-template regressions, context-bound receipts, complete paired artifacts and replay, an optional semantic stage, statistical counterexamples, process-recovery evidence, a frontier-risk chapter, and offline CI conformance. On 6 September, the full local unittest discovery passed 166 tests. That verifies software regressions, not live model quality, human-label validity, or production deployment.
Constraints¶
- Preserve existing chapters, examples, references, and source context. Add links and integrated explanations; correct erroneous claims explicitly rather than silently discarding material.
- Keep explained, implemented, experimentally demonstrated, and deployment-qualified status separate, scoped to the exact capability being claimed.
- Candidate agents must not issue semantic qualifications, alter trusted labels, or write authoritative scoring state. Judges propose verdicts; evaluator-owned code verifies scope and creates receipts.
- Preserve original runs and grades. Reassessment creates a new linked artifact rather than overwriting historical evidence.
- Missing, stale, revoked, mismatched, or unverifiable evidence cannot authorize expansion. A
lab_onlyreceipt never grants production authority. - No paid calls without a model configuration and spend cap; no claim of human calibration without independently reviewed labels.
Implementation contract and acceptance gates¶
| Phase | Changes and teaching deliverables | Acceptance evidence |
|---|---|---|
| 1. Close grading and replay trust gaps | Expand template tests across identity, eligibility, approval required/denied/revoked, escalation reason and transaction state. Harden evidence binding against replay across contexts. Enforce grader/source provenance and current calibration authority during replay; add immutable re-grading under a new grader. | Both reported false explanations fail. A receipt copied to different case evidence fails. Missing or changed artifacts and revoked qualification fail closed. Old and new grades can be inspected side by side without modifying the original packet. |
| 2. Complete semantic qualification | Ingest versioned annotation rows and recompute registry statistics rather than trusting supplied aggregate counts. Separate development, calibration and sealed evaluation data. Add a metered judge adapter, enforced timeout, disagreement review and usage accounting. Teach a blinded rubric pilot with worked false-pass, false-block and abstention examples. | A reviewer can reproduce each registry bound from retained labels. A frozen judge is tested on independent cases within its registered scope. The runner preserves its inputs, raw response, verdict, usage and receipt. Out-of-scope or uncertain judgments remain unqualified. |
| 3. Remove object-resolution answer leakage | Separate evaluator-only target identity from customer-visible requests. Introduce competing orders, ambiguous descriptions, corrections, clarification and counterbalanced record order. Keep server-side authorization independent of agent selection. | Tests show the evaluator's target designation and expected answer are absent from model-visible inputs; legitimate candidate order IDs remain available. Backend controls reject unauthorized attempts, while independent grading detects wrong-but-authorized actions. Necessary clarification and unnecessary clarification are scored separately. Report performance by ambiguity and distractor condition. |
| 4. Publish complete execution studies | First execute retrieval, oracle-document and full-context arms with a fixed model, tool contract and comparable budgets. Extend the existing crash study into actual agent tool execution and context reconstruction. Next add actual multi-agent delegation/merge traces and one coding or artifact domain. Voice follows when audio and independent timing annotations are available. | Each study retains inputs, execution traces, authoritative outcomes, grading evidence, costs, repeated trials, uncertainty and a defensible conclusion. Recovery verifies both durable side effects and restored constraints. A fixture that declares a failure is never described as detecting that failure from execution. |
| 5. Qualify statistical and operating decisions | Extend existing sparse-failure studies to unequal clusters, dependence, sequential looks, label error and holdout reuse. Pre-register estimands and method acceptance criteria. Add simulation-only shadow/canary/expand/restrict/rollback exercises and incident-to-reviewed-dataset promotion. | Known-population experiments report interval coverage and false-promotion behavior, including adverse cases. Routing tests cover mature outcomes, stale/duplicate windows, hard-stop violations and rollback identity. Observe the actual CI workflow separately; book publication and application deployment remain different authorities. |
| 6. Turn the frontier-risk bridge into inspectable evidence | Build on the existing chapter: execute benign synthetic cross-run isolation and monitor-interruption studies, with authoritative logs outside worker write scope. Compare capability, uplift, propensity, safeguards and residual uncertainty. Add useful-autonomy calculations and ethically reviewed human-study designs. | Report missed/late containment, actions completed before interruption, false stops, recovery burden and unintended sharing. Each risk claim links evidence, assumptions, controls and unresolved uncertainty. Local protocols do not claim frontier-risk thresholds or human behavioral effects they have not established. |
| 7. Integrate and independently review the book | Organize navigation around “what must I measure?”, “what evidence do I have?” and “what does it justify?”. Retain existing conceptual maps as supporting views. Audit every source-coverage row and major subject against the learning contract above. | Each substantive subject has a worked example, failure case, solution, interview prompt and explicit evidence status. A fresh reader can rerun at least one complete study and reconstruct its decision. Strict site build, links and rendered desktop/mobile pages pass review; published commit identity matches the reviewed source. |
The core evidence interface is: a versioned run manifest identifies full trial artifacts; trial records reference their hashes; judgments bind exact evidence and evaluator configuration; qualifications bind reviewed calibration evidence and scope; decision receipts reference that chain and state their authority ceiling. Human reviewers own adjudication, the evaluation service owns grading and registry checks, and an authorized release operator owns exposure decisions.
Qualification has a lifecycle: draft → independently reviewed → qualified for a declared scope → expired or revoked. Failure to qualify produces abstention/hold, not implicit approval. Exposure exercises have a separate lifecycle: shadow → bounded canary → expansion, with restriction or rollback on adverse evidence. Initially these transitions are local simulations only.
Non-goals¶
This revision does not establish universal frontier safety, infer human effects from simulated users, certify all clustered statistical methods from a sample-count rule, or turn a green book-build workflow into permission to deploy an AI application. It does not remove the existing breadth of the primer.
Open decisions and handoff¶
Phases 1 and 3, annotation-ingestion plumbing, offline statistical studies and local control simulations are ready for implementation with regression-first tests. Live qualification and model comparisons require the chosen agent/judge models and a spend cap; human qualification requires independent reviewers and an annotation protocol. Voice requires an audio source and annotation plan. Actual application rollout requires a specified deployment target, authority model and rollback mechanism.
Recommended first deliverable: one complete CX evidence packet, combining unambiguous evidence ownership, realistic object resolution, reviewed semantic grading and independent replay. Only after that packet is inspectable should the same harness be generalized to additional architectures. Each phase adds its worked example and limitations to the book alongside the implementation, rather than postponing teaching material to a final documentation pass.