Skip to content

Source Coverage Ledger

This ledger prevents the book from becoming a selective summary. It maps every substantive theme in the supplied research corpus to its teaching home, supporting example, and inspectable artifact.

What coverage means

A theme is not covered because its name appears once. Coverage requires an explanation, assumptions or limitations, a worked or counterexample, and an artifact or exercise where the topic permits one. Executable claims additionally require running-system evidence.

Source inventory

Course-report project sequence checked against the book

Teaching acceptance check — 7 September 2026: report 15's eight-row project table was compared directly with the course page's eight-lab map. Every project retains its learning purpose, worked preparation and learner receipt, including the original human-annotation and cross-framework execution requirements. No additional course-topic chapter is needed. This accepts the learning route, not completion of the learner's experiments or independent human qualification.

Repository-report teaching check: report 9's comparison, licensing, specialized-tool and architecture requirements map to System Studies' inventory, portability contract, same-refund-case comparison, judge-adapter example and solved stack-selection exercise. Vendor quick-start snippets remain in the supplied source rather than being represented as tested current commands. Tool adoption still requires the dated version/license/data-boundary checks in the study checklist. This accepts the architecture teaching coverage, not execution of every vendor integration or validation of snapshot rankings.

The source report's eight practical projects now have an explicit lab-to-exercise and receipt map: failure analysis, golden datasets, human evaluation, judge calibration, RAG, agents, red teaming and benchmark reproducibility. The map preserves the source's practice sizes while rejecting their use as qualification thresholds. It identifies independent annotation and the named two-framework execution as learner evidence still required, rather than substituting authored controls. This verifies the eight-project section's teaching route, not the entire supplied corpus's depth or completion of those empirical labs.

Durable window membership before execution

Kata 103 preserves all forty planned requests through a real worker kill before the first agent call. The existing execution path consumes registered cases/routes/scopes; conflicting registrations fail before work. A worked routing counterexample changes nine assignments when a canary predecessor is incorrectly replaced with expanded state. This closes immutable registration, not controller acceptance, attempt resumption or safety-stop enforcement. The registry binds a local campaign and caller-supplied manifest; hashes are consistency checks, not authenticated provenance.

Persisted completion versus committed effect

Kata 102 executes a real process kill after the existing mock refund commits but before its full case artifact returns. The separate completion journal retains an incomplete reservation, the campaign retains one charged effect, and replay refuses to invent a completed attempt. Completed records return identical evidence without rerunning the agent. Tests reject identity conflicts and inconsistent artifact-to-request bindings. This is one-shot completion storage, not recovery of an interrupted agent, durable window membership, controller recovery or customer-delivery proof. Existing content and historical packets remain unchanged.

Durable served-candidate routing and admission

Kata 101 connects the persisted backend to the existing multi-order/exposure path. Reopened shadow, canary and restricted windows retain cumulative charges of 0, 2 and 2; candidate completions are 40 shadow controls, 2/4 served and 0/4 served. Baseline and shadow do not consume served-candidate allowance. Grading reads all competing orders in one transaction. A unique durable request-start marker blocks reused and concurrent namespaces, including zero-tool and clarification-only attempts. This is admission fencing and persistent accounting, not resumable attempts, response recovery, controller checkpointing or deployment qualification. Historical artifacts remain unchanged.

Durable existing-world evidence and authorization

Kata 100 puts optional persisted campaign state behind the existing RefundWorld methods. Seed bindings, current authorization, refund effects and events are stored transactionally; paired state/event reads preserve grading evidence after reopen. Reference/timeout parity and specific unauthorized-action failures are tested. Review reproduced and repaired partial commits caused by broad exception handling. This is the first integration stage, not wiring through multi-order exposure, durable response storage, a process-recovery packet or production authority. Existing artifacts remain historical and unchanged.

Durable local payment allowance

Kata 99 adds optional database-owned policy to the existing recovery payment boundary. Approval, capacity and effect insertion share a write transaction; consumption derives from committed payments. Separate processes test competition for the final unit and worker loss before/after commit. This is a bounded local database implementation, not a durable integration of the multi-window exposure campaign or a remote payment system. Historical unbudgeted artifacts remain unchanged. The historical budgeted-packet test now pins its original file bytes and expects exact replay refusal after source/interpreter drift; fresh-packet replay remains required.

Dependence between attempts and safeguard failures

Kata 98 turns the existing warning about multiplied risk scores into an executable finite-population counterexample. Identical 1% marginals permit zero, one or one hundred joint failures per 10,000 opportunities. The solution separates conditional effectiveness, compatible populations, logical bounds, sampling uncertainty and consequence assumptions. These are synthetic labels and exact arithmetic, not measured harmful propensity, frontier-company risk estimates or release qualification.

Retained budgeted campaign replay

Kata 96 retains complete three-window execution evidence in a separately versioned budgeted packet. Fixed-builtin replay compares tool events, ledgers, namespaces, scope, budget accounting and controller decisions rather than trusting rehashed summaries. Only finite nonnegative trial timing is normalized. Exact source/interpreter matching deliberately limits portability; hashes and replay do not authenticate original execution. This closes the retained-packet follow-up below, not durable enforcement, actual model evaluation or deployment authority.

In-process campaign action budgets

Kata 95 now injects a shared count/per-currency allowance through the existing exposure/case/world path to the actual mock commit boundary. It separates served-candidate effects from baseline/shadow controls, enforces replay/namespace/reset rules, tests concurrent final-unit attempts and conservatively latches unknown accounting failures. The reference response path now reconciles unexpected tool status instead of blindly claiming a refund. This is single-process in-memory enforcement; the historical unbudgeted packet and CLI remain unchanged. Kata 96 supplies the retained budgeted replay packet; durable enforcement and real application integration remain open.

Historical judge example: scope and denominator clarification

The 120-case groundedness example now separates an imagined independent-review assumption from actual annotation evidence. Its original English qualification labels and companion JSON remain preserved as hypothetical values, explicitly not established by the aggregate table or valid for production-registry ingestion. The worked denominator comparison distinguishes 4/32, 4/40 and 4/74, plus classification coverage versus automatic acceptance and review load. No new calibration evidence is claimed.

Probability calibration versus useful decisions

Kata 94 extends the existing ten-row probability artifact without replacing it. A hindsight constant illustrates zero empirical ECE with no ranking, worse Brier/log loss and undefined selective risk at zero coverage. The solution distinguishes label reuse, held-out recalibration, threshold changes and review workload. Guo et al.'s primary calibration paper supplies methodological context, not local model evidence or a universal recipe for verbal confidence.

Shipping-example validity and authority

Kata 93 closes a teaching mismatch in the release chapter: its original short function is preserved as an explicitly unsafe counterexample, with executable NaN, unsupported-finite and missing-cost probes. The worked answer connects input validation, evidence qualification and current release authority to CI, routing, containment and recovery labs. This is a documentation/test repair, not a replacement production gate or deployed integration.

Numerical qualification and interview decision workshop

Katas 90–92 connect the question bank's calibration, uncertainty, data-curation and operational themes through supplied confusion matrices and executable calculations. The solutions distinguish conditional error from agreement, classification coverage from automatic acceptance, a fixed-sample zero-event bound from sequential stopping, and enriched-class sampling from representative within-class sampling. A hypothetical prevalence/queue calculation culminates in a qualification HOLD. This is an authored reasoning exercise, not new human labels, empirical model qualification or observed production staffing performance.

Executed coding-patch outcome controls

Katas 87–89 add a second executable domain: five authored patch controls, two visible and eight public acceptance cases, actual subprocess results, candidate diffs, protected-file checks and parent-verified outputs/input changes. The original and overfit controls both pass visible cases but differ on acceptance; the test-edit control is blocked before execution. The retained study binds local built-ins, harness/test definitions and interpreter; replay never executes its stored source. This establishes local patch-evaluation mechanics, not model-generated coding performance, sealed acceptance, independent test validity or an OS sandbox. The benchmark chapter also incorporates the dated OpenAI coding-task audit and Anthropic resource-enforcement study.

Incomplete CI evidence and integrated decision defense

Katas 85–86 exercise failed-run diagnostics, including a real harmless child-process timeout after two completed local controls. They separate expected candidate rejection from missing or invalid experiment evidence, and local writes from confirmed artifact storage. The capstone joins the learner's decision argument to three retained studies through 29 exact fields, with worked revocation, missingness and completed-side-effect variants. These additions deepen the operating and interview paths; they do not establish cloud workflow execution, crash-durable storage, independent human calibration or real application exposure.

Executed incident-to-regression controls

Katas 60–61 now reproduce a synthetic timeout-after-commit incident, apply a fixed minimization, bind a scripted review to exact proposal/source/policy identities, and publish a new regression-only version. The retained packet shows a passing reference and failing duplicate-producing mutant reconstructed from the released case, plus seven rejected promotion controls. Tests cover populated-parent preservation, recomputed content identity, exact duplicate rejection and known protected-role conflicts. This implements local promotion mechanics—not real incident ingestion, independent reviewer authentication, automatic de-identification, semantic deduplication, complete exposure history or sealed acceptance qualification.

Integrated CX evidence walkthrough

Katas 58–59 connect the existing native multi-order, semantic, campaign and release implementations through one retained v2 packet. The Python inspection blocks were executed; a separate clean local clone at 4ed96a7 with a newly installed locked environment reproduced the current study's diagnostic hold and six blocks. This strengthens the reproducible learner path, not model-quality evidence: the agent, judge responses, calibration, clocks, prices and prerequisite assertions remain synthetic. No independent human qualification or application exposure was performed.

Executed cross-run cache controls

The Cross-Run Isolation Study and Katas 56–57 execute six SQLite-backed comparisons: query-only versus run-scoped keys, each with fresh worker objects, a reused object and two concurrent threads. The retained packet contains 18 responses and 55 operations. The buggy key exposes one foreign marker in each worker mode; scoped keys expose none while preserving same-run hits, both shared-policy reads and two overwrite refusals. Replay binds phases, worker identities, concurrent thread identities and causal order to the registered contract. This is application-level cooperative-code evidence—not hostile-worker isolation, authenticated logs, model behavior or a population leakage rate.

Executed knowledge-to-action controls

The Knowledge-to-Action Study and Katas 51–53 add 18 local executions across lexical retrieval, oracle documents and full context, crossed with a correct structured-policy control and a wrong-amount mutant. The retained packet supports re-execution of rankings, context, decisions, rejected/accepted mock writes and final-state grades. Correct-control completion is 1/3, 3/3 and 3/3; mutant completion is 0/3, 1/3 and 1/3. All cases share one customer and two invented policy documents. Natural-language policy interpretation, active-version discovery, realistic search scale, representative sampling, model runs and deployment qualification remain open; these results do not establish a production improvement.

Current release evidence composition

Katas 49–50 add explicit file/value source verification and a composed current-release assessment. The retained study verifies actual committed repository bytes, replays one sixteen-trial native packet and applies seven current-state controls: current diagnostic holds; revoked, expired, synthetic-disabled, wrong-revision, changed-case and unqualified-prerequisite controls block. Four separate calibration executions and all provider responses are synthetic. Historical grades survive revocation unchanged. This closes local source/current-calibration composition—not authoritative registry retrieval, human calibration, execution attestation, atomic exposure authorization or observed cloud CI.

Executed statistical-method evidence

The statistical method study now supplies full count-level results for known-population interval coverage, false promotion, power, and biased-label counterexamples, plus an unequal-cluster estimand example. Its retained artifact is checked against the generator by the test suite. The alternative follows NIST's binomial-tail inversion equations, verified on 2026-09-06. This supports the narrow fixed-sample binomial lesson; general clustered and sequential method qualification remains incomplete.

Supplied research corpus

The Semantic Grading Lab adds tested runner integration and Katas 08–10 for scoped qualification, class-error bounds, abstention, configuration drift, expiry/revocation, and retained judge evidence. Katas 18–20 add an executed row-derived annotation compiler, a retained four-row synthetic packet, independent-adjudicator checks and declared split-leakage controls. Fixture judges, synthetic labels and reviewer identifiers are control tests, not human calibration or live-model accuracy evidence. Authenticated human annotation provenance, representative data collection and live semantic studies remain open.

Katas 36–37 add the optional Responses judge adapter: strict verdict parsing, incomplete/refusal abstention, effective-endpoint checks, defensive usage validation and retained raw response evidence. The installed SDK is exercised through an in-memory HTTP transport; four synthetic paired judgments are retained and replayed. No live semantic accuracy is established. Judge costs remain separate audit overhead, and hard campaign budgets/cancellation and reviewed live calibration are still open.

Katas 38–40 extend that implementation with a durable local estimated-cost judge-admission ledger, stable paired invocation IDs, unknown-cost reservations and selected agent/judge cost aggregation. A retained four-trial mock-world campaign contains two scripted judge calls and two admission denials; subprocess crash and multiprocess contention tests exercise persistence and atomic admission. This is not a provider invoice ceiling, whole-experiment resume, live semantic study or production release prerequisite. Agent/tool budgets, authenticated reconciliation, cancellation and full operational costs remain open.

Katas 41–42 add pinned campaign-snapshot validation against packet execution, qualification, verdicts and cost records, plus integration with the existing lab release receipt. Four retained offline controls and a CI command exercise clear/hold/block composition. Operator anchors and prerequisite/calibration data in these controls are explicit fixtures. Snapshot freshness, issuer authentication, current-calibration composition and actual canary authority remain unestablished.

Research source Primary contribution Canonical treatment
Best GitHub Repositories for LLM Evaluation in 2026 Tool categories, quick starts, licensing, portability, model/application/production layers System Studies, Benchmark Reproducibility, Courses
LLM and AI Evaluation Interview Question Bank Thirteen-topic evaluation taxonomy, evidence hierarchy, benchmark map, senior whiteboards All Part I chapters, Interview & Design Drills, and exercises
LLM Evaluation Around OpenAI Deep Research Research-agent capability, browsing benchmarks, safety, source authority, temporal fragility, validity RAG & Research Evals, Robustness, Benchmark Reproducibility
Best Courses and YouTube Resources for Learning LLM Evaluation Learning routes, eight practical labs, tool practice, human/statistics gaps Courses, Part II lab, Source Coverage Ledger
Eval-Driven AI Programming Error-analysis-first discipline, evaluator patterns, calibration, experiments, CI/CD, capstone Dataset Design, Metrics, Judge Calibration, Build Lab
Evals, Observability and Release Gates for Production AI Systems Control loop, telemetry, release policy, canaries, incidents, governance, ownership, economics Production Evals and Release Gates
Evaluating LLM-Based AI Systems: Research Review and Production Blueprint 2025–26 frontier methods, deployment evidence, extreme-tail reliability, stateful personalization, evaluator meta-evaluation, dynamic audits, monitorability, and realtime modalities Research-to-Practice Evidence plus linked Part I chapters

The supplied PDF and DOCX files have corresponding research titles, but title agreement does not establish content equivalence. The editable Markdown counterparts remain the canonical teaching inputs; retain the companion files as provenance and check for additional source links or substantive differences before treating them as redundant.

Companion-PDF preservation check — 7 September 2026: text extraction from all four supplied PDFs found substantial shared wording with the corresponding Markdown, but not identical text. The interview PDF contains all 118 canonical question IDs preserved in the appendix; an additional regex match, P02, is part of an ACL source URL, not a missing question. The PDFs also expose explicit reference URLs where some Markdown uses generated citation markers. Those links are potentially useful provenance, not duplicate curriculum topics. This check establishes question-ID preservation only: token overlap cannot establish sentence, table, answer-outline or source-link equivalence, and no PDF layout assessment or DOCX-equivalence claim is made.

Section-level completeness audit

Companion source-recovery receipt

Ranked frontier-review method map

The numbers below are the 46 entries in section 5 of the supplied later DOCX, not independent validations of its paper findings. Several papers motivate the same learning method; the book need not reproduce each benchmark to teach that method.

Source entries Worked teaching route Acceptance boundary
1, 2, 3, 6, 25, 26, 27, 28, 40, 42, 43, 44, 45 Judge qualification, ranking reversal, lineage and sensitivity; annotation and denominator controls; calibration decision workshop Covers local validity, reference dependence, protocol effects, correlated errors and qualified decisions; authored labels are not expert studies.
4, 13, 20 Metrics and uncertainty; statistical-method study; frontier risk decisions Paired/clustered uncertainty and rare-event assumptions require separate qualification; no universal tail guarantee.
5, 10, 11, 29, 30, 31 Benchmark reproduction; coding-patch study; long-report study Worked freshness, harness, scoring and construct-validity distinctions, not reproduction of all named benchmarks.
7, 8, 9, 12, 34, 35, 36, 37 Agent evaluation; recovery; modern architectures Outcome, trajectory, subgoals, budgets and persistence; some architecture observations remain supplied fixtures.
14, 15, 16, 18, 19, 32, 33 Safety and reward optimization; frontier risk/containment; cross-run isolation Separates capability, harmful effects, monitoring and enforcement; local controls do not establish field prevalence.
17 Assigned-cohort human-workload comparison and deployment-simulation protocol Illustrates offline/field and user-outcome boundaries; not a personalized-product field replication.
21, 22, 23, 24, 38, 39 Existing research evaluation, knowledge-to-action and report studies Existing decomposition and stateful evidence routes retained; further RAG expansion deferred at the user's request.
41 Uncertainty-wording paired exercise Explicit criterion, paired wording intervention, flip-rate and class-error solution; authored counts do not establish actual model bias.
46 Voice and timing examples Audio-condition and action-timing protocol with authored timing arithmetic, not recorded voice-model qualification.

This closes the routing inventory and its identified entry-41 teaching gap. It does not certify every source claim, empirical transfer, or complete reproduction of the ranked studies.

Ranked-review depth check — 7 September 2026: the later DOCX's ranked evidence capsules include an explicit instance-versus-system-ranking validity distinction (entry 25). The ranking-reversal kata supplies the missing numerical counterexample and solution. The fourteen-headline map remains useful but is not a substitute for inspecting the detailed methods. This check reads the supplied claims; it does not independently verify the reported findings or reproduce all 46 studies.

The extracted companion reference index preserves 198 within-document distinct URL strings across the four PDFs and two matching DOCX files, with source filenames, SHA-256 fingerprints and extraction method. URLs can recur across documents. This is a recovery index, not a new bibliography endorsement: extraction may stop at a wrapped line, and each recovered URL still needs resolution and claim-level attribution before it supports a new factual statement. Existing references remain unchanged.

The two DOCX/Markdown comparisons also show why both formats should be retained. Markdown contains Mermaid diagrams and equation source omitted by plain-text DOCX extraction; DOCX text contains numbered reference URLs in addition to the prose. Normalized word-sequence similarity was approximately 92.5% for Code-First and 93.0% for Evals/Observability, but those values are diagnostic only—not completeness scores. These checks do not establish that every short wording change is immaterial or that diagram images are absent from the original documents. Do not discard either source on this evidence.

Later DOCX project-depth check: section 13's Judge Lineage Auditor now has a worked anonymous-output/name-intervention comparison. Its synthetic counts distinguish identity intervention, non-causal generator-group association, shared panel errors, abstention and conditional error. This closes a worked-example gap, not empirical lineage attribution or a full reread/verification of the DOCX's 46 ranked research claims.

This second map checks the structure of each source, not only the combined topic taxonomy.

Interview-bank preservation check: the complete appendix now retains all 118 unique canonical IDs across thirteen themes, with every supplied question/variant, level and answer outline checked against deep-research-report-10.md. The condensed drills remain available. This proves question/outline preservation, not that every outline is itself a full worked solution; deeper teaching remains in the mapped chapters, exercises and capstone.

Reward-model depth check: source-bank L4–L8 and L11–L12 now have a worked optimization-pressure case, with separate preference/task/safety denominators, paired counts, executable arithmetic, a block/hold solution and an experiment protocol. This closes the identified checklist-only treatment in that section; the data are authored, not empirical RLHF results.

Input identity check — 7 September 2026

The six supplied Markdown files were found locally; their actual first-line titles match the six corresponding corpus rows above. The numbered filenames resolve as follows. Fingerprints identify the inspected input version, not the quality or completeness of its treatment.

Input filename Corpus title / role SHA-256
deep-research-report-9.md GitHub repository landscape 00322670378566b7283ae8958bda1e74de5ecf7211e7d97d36fe8af4510a564e
deep-research-report-10.md Interview question bank b384e42874ca12487d3854877ee7f6736d7979cb749d68efe5dd9eec83021839
deep-research-report-13.md OpenAI Deep Research evaluation report 6832298d6e379f661d708d3fa58c67f16ddd7bbdbf2d64b3f5822dcfe1a50f1c
deep-research-report-15.md Courses and YouTube report 726f05c6ac9e99e49ed3d1726c8cddc0c1f3a0fe8960cedfbd94cac92bfd50df
Eval-Driven AI Programming A Code-First Path to Production AI Engineering.md Eval-driven programming 15514e073b0b3f3d56c20e429eca462b0f8f942f5dc25f8875044984abf3da31
x Evals, Observability and Release Gates for Production AI Systems.md Production evaluation operations 21abcf7640bc529af0dc6db41f7e3671416fa2dae14a9a9267ac726026c895d0
llm-evaluation-research-review-2026.docx Later research-review input; file presence/fingerprint checked, contents not re-audited in this check 7156e7d910fe0cb1e29b8c262ca0ee65ed95450f1575ef2d33dc058192b637f6

Audit limit: the following section map is an editorial coverage assertion, not a completed independent section-by-section acceptance test. Title matches, fingerprints and broad chapter links do not prove that every source example has an adequate worked solution. Companion PDF/DOCX content equivalence was not reverified in this input check. Final content acceptance must inspect the mapped teaching sections and examples; it must not infer completeness from this inventory alone.

Source Source sections retained Portal treatment
GitHub repository landscape Ranked shortlist; comparison dimensions; licence implications; quick starts; specialised tools; layered architecture and next steps System Studies keeps the complete tool inventory, selection/portability/licensing contracts, worked adapter comparison, and a cross-tool reproduction exercise. Quick-start commands remain in the source because the book teaches contracts rather than duplicating change-prone vendor syntax.
Interview question bank Research taxonomy/counts; ranked core questions; 13-theme bank; evaluation architecture; whiteboards; benchmark/tool map Thirteen Part I homes plus Interview & Design Drills, answer contract, counterprompts, whiteboards, senior scenarios, and source/reference maps.
Deep Research evaluation report Identification method; first-party corpus; HLE/GAIA/BrowseComp constructs; validity trajectory; live-web fragility; safety; external research network; timeline/source map RAG & Research Evals, Benchmark Reproducibility, Robustness, primary reference map, temporal manifests, browsing budgets, authority/citation/completeness and indirect-injection controls.
Courses and YouTube report Assessment method; ranked resources; learning paths; eight practical projects; tool stack; requested-framework map; paper spine; human/statistical gaps and next steps Courses & Learning Paths, Part II project table, System Studies inventory, References, Human Evaluation, Metrics, and judge-calibration exercises. Commercial rankings and stale star counts are not repeated as architectural evidence.
Eval-Driven AI Programming EDD discipline; resource routes; typed cases/results; deterministic/reference/model/human graders; calibration; failure analysis/intervention ladder; observability/CI; curriculum/capstone; first-session 80/20 Foundations through Release Gates, executable CX lab and mutants, prompt experiment, living dataset, evidence pyramid, courses, and staged roadmap.
Evals, Observability and Release Gates Definitions/KPIs; architecture/tooling; methods/risk profiles/gate template; observability/incidents/governance/ownership/economics; operational cases; roadmap/resources/risks Foundations, seven surfaces, Production, Release Gates, governance roles, programme economics/roadmap, System Studies case patterns, and synthetic gate/canary artifacts.
Evaluating LLM-Based AI Systems Fourteen 2025–26 developments; five-layer/six-family taxonomy; ranked research map; production blueprint; demo ideas; proposed book structure Research-to-Practice Evidence supplies the dated maturity rubric and primary-source proof matrix; Metrics covers rare events; Judge covers evaluator qualification and correction; RAG covers stateful personalization; Robustness covers auditing/monitorability; System Studies covers dynamic generation and realtime modalities. Research-only claims remain explicitly non-gating.

Excluding a volatile ranking, price, star count, or copied quick-start is intentional: the durable concept and verification method are retained, while the live claim must be rechecked at use time.

Recent capability addendum

The Sequential Decisions Lab adds exact finite-horizon path enumeration and Katas 28–31. It compares fixed-final, repeated-fixed, Bonferroni and likelihood-ratio tests, with power, false-promotion and stopping-cost results plus label-bias and dependence counterexamples. The model is a narrow iid Bernoulli loss problem; it does not qualify general clustered traffic, model accuracy or release authority.

The Exposure Control Lab adds an executed 280-request routing simulation and Katas 24–27 for stable cohorts, maturity, missing labels, expiry, quality restriction, recovery hysteresis and rollback. It retains 349 deterministic mock-agent executions. It does not confer statistical or deployment qualification, enforce monetary/request caps, interrupt active work or configure a cloud application router.

The Order Resolution Study adds 16 executed deterministic trials with independent mock order ledgers, label-free unresolved inputs, customer corrections, scripted clarification and counterbalanced ordering. Katas 21–23 distinguish authorized wrong-object actions, retrospective clarification and rejected attempts hidden by later success. The live SDK path is contract-tested with a fake runner; live-model accuracy, question quality and qualified semantic grading remain unmeasured.

Katas 43–44 extend that study with two native paired comparisons (32 new deterministic agent executions). Each packet preserves all order ledgers, replays allowlisted mock-tool transitions, binds full case/design identities and recomputes structural grades. The release builder consumes the paired results with explicitly unqualified prerequisites: even the 8/8 structural control blocks. All cases share one customer; semantic qualification, live-model transfer and deployment authority remain unestablished.

Katas 45–46 add a native multi-order truth criterion, context-bound receipts and a row-derived synthetic calibration study. Three additional paired controls join structural and semantic outcomes, and a current-registry revocation assessment preserves historical replay. The deliberately literal judge and four synthetic annotation rows test interfaces; permissive diagnostic thresholds and overlapping scenarios do not establish semantic validity. Original structural-only packets retain their original evidence status.

Katas 47–48 connect the native criterion to the optional provider adapter and direct budget wrapper. The retained installed-SDK/in-memory-HTTP study executes 64 paired mock-agent trials plus four calibration executions, with 35 paired SDK requests plus four calibration requests. Known usage, 429, admission limits and reservation overrun produce campaign clear/hold/hold/block; every application release remains blocked. Request/criterion/ledger joins and separate retained-provider re-decoding tests establish local consistency, not human calibration, model accuracy, billing authenticity, current release authority or live deployment. Independent labels, untouched transfer cases, agent budgets, cancellation and incident-driven dataset operations remain open.

The CI Gate Lab adds locally executed pass/hold/block controls with full replayable packets and a configured read-only GitHub Actions workflow. Katas 16–17 separate conformance from candidate qualification and deployment. Cloud execution, branch-protection enforcement, and application canary control are not yet demonstrated.

The Frontier Risk Decisions chapter addresses the audit's risk-decision gaps using six rechecked primary sources: Anthropic's RSP v3 discussion, DeepMind's Frontier Safety Framework update and human-participant research, OpenAI's safeguards disclosure, and METR's incident investigation and task-horizon limitations. Katas 13–15 are worked reasoning/calculation exercises. Katas 54–55 add six executed local process/SQLite enforcement controls with retained effects and causal event order: early revocation, late revocation, coordinator-only cancellation, delegated revocation, false stop and benign completion. The packet records eight process lifetimes. Scripted alerts do not measure detector accuracy; hostile-code isolation, distributed revocation, participant studies and deployment qualification remain open.

The supplied reports remain the coverage baseline. This dated addendum captures capabilities and failure modes that became more prominent after parts of that corpus were written.

Emerging capability or risk Portal treatment Qualification boundary
Scorer and evaluator isolation Agent & System Evals separates model-action sandboxing from trusted task/scorer code and fresh scoring state A sandbox flag does not establish host isolation; review the actual runtime and dependency boundary
Evaluation integrity and awareness Benchmark Reproducibility threat-models answer keys, task identity, harness tampering, denominators, traces, and release identity Suspicious retrieval can be a security finding without remaining a valid capability score
Generated dynamic behavioral evals System Studies compares Petri, Bloom, and adaptive Giskard multi-turn scenarios Generated cases amplify discovery; reviewed fixed cases and calibrated graders retain gate authority
Adaptive simulated users Agent & System Evals and System Studies cover changing user turns, shared trace, termination, and non-collusion Simulator behavior must be checked against held-out human sessions and is not a traffic distribution
Tool ownership and lifecycle System Studies records the 2026 Petri transfer and the AgentKit lifecycle change Reverify current owners, releases, licenses, data path, and exportability at adoption time
Research-to-practice maturity Research-to-Practice Evidence separates production controls, field evidence, operational tools, and research/benchmarks A repository or paper is never promoted to “production-proven” without a named operational use and first-party receipt
Extreme-tail reliability Research-to-Practice Evidence and Metrics qualify Five-Nines/CEM importance sampling Reproduce proposal support, weight stability, and offline-to-live validity before any local gate authority
Stateful personalization and memory Research-to-Practice Evidence and RAG & Research Evals define reset/persist, counter-user, and preference-update arms Field evidence proves an offline gap, not a universal production evaluation recipe
Evaluator-of-evaluators and judge correction Research-to-Practice Evidence and LLM as a Judge cover AgentRewardBench and bias-corrected reporting Meta-benchmarks screen candidates; representative local human evidence qualifies them
Monitorability and realtime voice Research-to-Practice Evidence and Robustness/System Studies record disclosed operational use and its limits Operational use is scoped to the named system, control, and source; it does not establish a complete safety or quality stack
Long-running and resumable agents Long-Running & Serving Failures retains four-arm persistence, approval, exactly-once goals, and restart scorer contracts; Process Recovery Study adds 21 executed local trials and Katas 11–12 Real process interruption now produces durable mock-payment evidence; compaction, distributed exactly-once claims, persistent-model behavior, and production restart evidence remain unqualified
Knowledge-to-action systems Modern Agent Architectures compares retrieval, oracle documents, full context, and alternative search interfaces Supplying the right document removes one failure source but does not qualify reasoning, action, or final state
Skill selection Modern Agent Architectures separates activation, instruction freshness, and execution Local protocol logic is tested; no live skill-loader trial has been run
Multi-agent coordination Modern Agent Architectures adds delegation, coverage, duplication, handoff, merge, contention, effect, and matched-budget criteria Buildable scorer fixtures do not establish execution coverage or superiority of a multi-agent design
Serving-failure behavior Long-Running & Serving Failures covers throttling, partial streams, ambiguous commits, fallback, and degraded tools Fallback must satisfy the same product contract; no live provider fault experiment has been run
Evaluation-service integrity Eval Operations & Integrity adds fragment-merge, conflict, contamination, factorial, and simulator-scoring contracts These are scorer/merge fixtures, not an operated distributed evaluation service

This addendum does not replace the source-by-source audit. It extends the book while keeping the same rule: an emerging technique becomes operational evidence only after its construct, data, evaluator, environment, and authority are qualified.

Interview-taxonomy coverage

Item-level teaching audit: RAG and agents

Source IDs Inspected teaching evidence Finding
A1, A3–A6 Knowledge-to-action study, context-packing Katas 68–69, citation Katas 76–77 Executed local decomposition, oracle controls and claim/citation outcomes with worked solutions.
A2 Ranked-evidence calculation Adds concrete hit/precision/recall, reciprocal-rank and nDCG calculation with judged-collection assumptions.
A7–A11 Timeout-after-commit example, two valid paths exercise, process recovery study, incident rollback State, trajectory, recovery, irreversible effects, efficiency and noncompensating release decisions remain separate. Generic agent continuation is not implemented.
A12 Model/harness factorial, deployed/eliciting comparison, order-resolution study Worked attribution and answer-leakage distinctions; fixed-control results are not evidence of emergent model capability.

A1–A12 now have concrete teaching routes. This completes the question-family mapping, not verification of every claim in the other research reports or execution of every proposed architecture study.

Item-level teaching audit: tooling and continuous evaluation

Source IDs Inspected teaching evidence Finding
M1–M2 Controlled prompt development, test pyramid, CI Katas 16–17 Concrete iteration and CI contracts; expected candidate rejection can be a successful software test.
M3 Telemetry Delivery Study, sensitive payload controls Executed loopback payload/receiver evidence and feedback joins; raw prompts are not indiscriminately logged.
M4, M7 Matured outcome contract, Exposure Control Lab, assigned-population example Worked missing-label, exposure, quality and human-work outcomes. Proxy drift is a trigger, not proof of changed true outcome rates.
M5 Gate dependency order, unsafe shortcut counterexample, release capstone Concrete pass/hold/block reasoning with independent constraints, evidence identity and authority.
M6 Layered production sampling, weighted sampling study Executed sampling arithmetic and decision limits rather than grading every event semantically.
M8 Tool-selection matrix, stack-selection exercise Worked contract-based choice; vendor popularity does not establish suitability or lifecycle support.
M9 Incident-to-regression controls, interrupted-effect rollback Executed local incident reproduction, reviewed-fixture promotion and bounded containment decision.
M10 Service architecture, budget campaign Katas 38–40, partial-run exercise Concrete component contracts, accounting and incomplete-pair solution. Distributed multi-tenant operation is not demonstrated by local fragments or SQLite stores.

M1–M10 have worked teaching routes. Hosted book conformance and Pages publication succeeded at aed856e; that is not deployment of the evaluated customer-support application. The source outline's binary pass/fail is deliberately expanded to pass/hold/block so insufficient evidence is not confused with a proven candidate failure.

Item-level teaching audit: safety and alignment

Source IDs Inspected teaching evidence Finding
L1, L4–L8, L11–L12 Preference/reward workflow, optimization-pressure example and solution, reviewer-effects exercise Separates supervision quality, reward prediction and post-optimization behavior; worked proxy/preference/task/safety contrasts and independent-evaluation protocol. No actual RLHF/RLAIF training comparison is claimed.
L2, L10 Threat manifest, adaptive-budget calculation Explicit attacker capability, budget, stopping, denominators and validity limitations.
L3, L9 Crossed harmful/benign evaluation, optimization decision, noncompensable constraints Concrete harmful-compliance/false-refusal separation and block/hold solution; neither the refusal count nor proxy reward alone decides release.

The L1–L12 questions have explicit worked teaching routes using existing examples. Population safety, feedback validity and generalization beyond the supplied controls remain empirical claims requiring separate evidence, not properties granted by an alignment-method name.

Item-level teaching audit: interpretability

Source IDs Inspected teaching evidence Finding
I1–I2 Definitions and explanation limits, policy-passage intervention Concrete distinction between understandable output and causal evidence; coherent rationales are not privileged access to computation.
I3–I5 Method-evaluation table, known refund mechanism Explicit intervention/control design and worked feature precision/recall; necessity, sufficiency, coverage and false discoveries remain distinct. Corrected the blanket weight-randomization claim: a control must establish what mechanism it changed before expecting an attribution change.
I6 False discoveries and complementary safety evidence, frontier risk decisions Bounded incorporation of mechanism evidence rather than substitution for behavioral and operational controls.

The existing I1–I6 toy example and design satisfy this scoped teaching check. No neural-model interpretability experiment was run or qualified by it. Feature-set precision and recall are only proxies in the known toy mechanism, not universal measures of causal completeness.

Item-level teaching audit: fairness

Source IDs Inspected teaching evidence Finding
BF1, BF5–BF8 Fairness contract, raw/standardized case-mix solution and intersection follow-up Worked outcome/denominator comparison, group-support limits, harm and retained-slice reasoning; no universal fairness score.
BF2–BF4 Matched identity-prompt and toxicity-grader exercise Adds a concrete comparison protocol, differential false-positive calculation and BBQ/StereoSet construct distinction. Reference labels and groups are authored, not a collected fairness study.

BF1–BF8 now have scoped teaching routes. Actual fairness assessment still requires the affected population, justified harm criteria, reliable labels, uncertainty and appropriate governance; these examples do not establish equitable deployment.

Item-level teaching audit: robustness and adversaries

Source IDs Inspected teaching evidence Finding
R1–R3, R6 Transformation relation table, metamorphic pseudocode, conditional robustness Concrete invariant/directional examples, execution relation and consistently-wrong caveat.
R4–R5 Threat manifest, crossed benign/disallowed cases Attacker control, prohibited effect, normal-edge variation and false-refusal checks are distinguished.
R7, R9–R10 Adaptive attack-budget exercise Adds worked scenario/query denominators, stopping, held-out families and observed-worst-group limitations. Mock results do not establish realistic attacker coverage.
R8 Noncompensable invariants, incident rollback solution, capstone Concrete containment decision and preserved committed-effect evidence rather than averaging away severe failure.

The R1–R10 teaching routes are explicit. The new adaptive example uses authored counts; it does not execute an attack optimizer or claim empirical security qualification.

Item-level teaching audit: benchmarks

Source IDs Inspected teaching evidence Finding
B1–B4 Benchmark construct map, five validity threats, benchmark decision exercise Benchmark definitions and concrete product-transfer counterexamples; a model score is not a refund-system release gate.
B5, B8 Integrity threat and receipt, exposed-case rejection Worked provenance/access and contamination decisions; no single black-box probe establishes absence of memorization.
B6 Saturation and contamination, dataset lifecycle, weak-set redesign Concrete refresh, role separation and preserved-regression design. No promise that any finite benchmark resists saturation indefinitely.
B7 Finite-candidate pass@k derivation Added missing subset estimator, ten-candidate numerical example, assumptions and selection-oracle limitation.

This B1–B8 audit checks teaching coverage. It does not claim new benchmark model runs or close the separately listed cross-framework reproduction experiment.

Item-level teaching audit: fundamentals

Source IDs Inspected teaching evidence Finding
F1, F10 Evaluation contract, validity, seven-surface repair exercise Concrete contract, evidence distinctions and solved repair decisions; training loss is not the application objective.
F2 Perplexity calculation Added missing token-loss calculation and hard-invariant counterexample.
F3 Classification metrics, joint error table Formulas, class naming, imbalance counterexample and concrete conditional-error denominators.
F4–F5 Negation/paraphrase overlap exercise Worked diagnostic and solution separate lexical overlap, contextual similarity and factual truth. No fabricated library scores.
F6 Long-report study, especially Katas 78–80 Inspectable documents, extraction/omission/compound-claim and synthesis decisions with solutions. Generic summary quality is not reduced to overlap.
F7 Exact-match contract, identity-joined scorer exercise Concrete executable classification comparison, duplicate/missing-ID rejection and normalization limits.
F8–F9 Independent release rules, human-workload counterexample Solved noncompensating safety constraints and misleading proxy-improvement scenario.

This F1–F10 review closes the missing perplexity and concrete overlap-counterexample gaps. It establishes scoped teaching coverage, not empirical validity of every named metric on the reader's application.

Item-level teaching audit: prompts and dataset curation

Source IDs Inspected teaching evidence Finding
P1, P5 Seven dataset roles, weak-set redesign Concrete dataset composition, access boundaries and solved release-design task.
P2–P3 Controlled prompt experiment, reproducibility manifest Versioned intervention and paired quality/safety/resource comparison; not a claim that an actual prompt was optimized successfully.
P4 Crossed demonstration-set/order example Worked averages and interaction, selection-bias defense, per-item and prompt-length requirements.
P6 Synthetic manifest and review, incident promotion controls Concrete generation metadata and executed rejection/promotion mechanics; human validity is not inferred from a manifest.
P7–P8 Mixed-trace workshop, sampling study Worked clustering/repair and weighted-sampling calculations; local controls do not demonstrate processing millions of real traces.
P9–P10 Twenty-prompt selection, exposed-case rejection, assigned-population impact Concrete selection, leakage and offline-to-impact counterexamples with solutions. No universal numeric correction for arbitrary adaptive reuse is claimed.

This scoped P1–P10 audit adds the missing P4 worked comparison and reuses existing dataset mechanics. It establishes teaching routes, not live prompt improvement or production-data qualification.

Item-level teaching audit: LLM judges

Source IDs Inspected teaching evidence Finding
J1, J3 Judging designs, rubric schema, authority exercise Concrete criterion/schema and decision examples distinguish relative preference from absolute acceptability.
J2, J5 Katas 74–75, lineage/name intervention Retained order/rubric/verbosity controls and worked panel/name calculations. These are not live-model bias estimates.
J4, J6, J8 Katas 08–10, calibration decision workshop, 120-case denominator explanation Worked qualification, disagreements, authority restriction, abstention and requalification; independent labels remain necessary for real qualification.
J7 Generator/judge joint-error calculation Adds the generator-conditioned escape rate, covariance, phi correlation and prevalence-dependent acceptance error missing from judge-panel comparisons alone.

All eight questions now have explicit teaching routes with concrete examples or solved decisions. This scoped audit does not certify the rest of the source corpus or claim that the synthetic reference labels were independently collected.

Item-level teaching audit: human evaluation

The H1–H8 source questions were checked against the following concrete teaching sections. These are educational coverage findings, not independent-human study results.

Source IDs Worked teaching evidence Scope
H1–H2 Decision/task contrasts, annotation specification, weak-study redesign and solution Defines construct, evidence, population, blinding, reviewer overlap and decision authority.
H3 Anchored ordinal scale and pairwise alternatives Concrete scale and both-unacceptable counterexample; no universal superiority claim for either format.
H4–H5 Eight-case pilot, Kata 72, adjudication record Item-level disagreement, exact nominal kappa calculation, missingness, rubric repair and preserved original labels.
H6 Ratings-versus-cases planning calculation Worked approximate precision and annotation-cost calculation; explicit independence, label-error and slice limitations. Not a general power calculator.
H7 Reviewer qualification, pilot H03/H06/H07 Distinguishes authoritative state, policy expertise, language ambiguity and subjective tone rather than declaring one reviewer universally correct.
H8 Assignment-confounding control, rater-effects model and solution Worked offset/treatment decomposition, disconnected-design non-identification and model-family cautions. No fitted empirical mixed-effects model is claimed.

This closes the H6/H8 worked-method gaps identified in this pass. It does not establish coverage of unrelated question families or replace actual reviewer qualification.

Item-level teaching audit: calibration and statistics

This pass checked the retained C1–C8 and S1–S8 questions against the teaching sections below. “Worked” means the cited section supplies concrete inputs and an explained result; it does not mean empirical qualification. The findings are limited to these sixteen questions.

Source IDs Exact teaching evidence Finding
C1–C3 Ten probability/outcome rows and score calculations; zero-ECE counterexample Worked definitions, probability-versus-correctness distinction, Brier/NLL/ECE calculations and failure case.
C4 Threshold decisions; qualification/workload workshop Worked coverage, selective error and review-load decisions; independent validation required.
C5 Fixed-bin support, mean-confidence and accuracy table Reliability-diagram coordinates and binning method are supplied. A plotted diagram is not supplied by this section; readers can draw the listed coordinates.
C6–C8 Meaning-frequency entropy exercise; sequence likelihood normalization Worked string/meaning distinction, length-normalization reversal and consistently-wrong counterexample. No claim to reproduce the source's theoretical hallucination theorem.
S1–S2, S7 First divergent boundary, Kata 81, manifest Concrete disagreement/join exercise, solution and retained-field contract. External-framework execution is still not demonstrated.
S3–S4 Paired bootstrap code; zero-discordance counterexample Code preserves pairing; worked counterexample explains why a narrow interval can fail. The short bootstrap function is instructional, not a validated general-purpose API.
S5 Benchmark decision exercise; capstone decision defense Worked evidence-demand and hold/block reasoning, including quality, risk, costs and scope.
S6 Twenty-prompt selection exercise; sequential decision study Worked family threshold, adjusted p-value, independence counterexample, union-bound explanation and hold/validation solution. Simultaneous selection and repeated looks are kept distinct; supplied p-values are assumed valid, not derived from model trials.
S8 Repeated-trial command and five-cluster hold; unequal-cluster exercise Worked distinction between repeats, independent customers and weighting. General hierarchical resampling with stochastic generation remains a methodological extension, not implemented by the short bootstrap snippet.

Do not infer acceptance of the other 102 questions from this subset or treat the source answer outlines as full worked solutions.

Fundamentals & evaluation metrics

Concepts Book home Example/artifact
Product objective, failure cost, evaluation contract, unit, task/case/trial/trace/outcome Foundations CX refund contract and vocabulary table
Deterministic, reference, semantic, executable, human, and online evidence Foundations; Metrics Evidence hierarchy and metric selection examples
BLEU, ROUGE, semantic similarity, task and operational metrics Metrics Text-metric limitations and executable-state counterexample
Seven independent evaluation surfaces Agent & System Evals; Build Lab Timeout-after-commit surface table

Benchmarks & contamination

Concepts Book home Example/artifact
MMLU, BIG-bench, GPQA, GSM8K, HumanEval, SWE-bench, HELM Benchmark Reproducibility Benchmark construct map
Saturation, prompt sensitivity, answer extraction, harness effects Benchmark Reproducibility Four-item two-harness worked example
Leakage, semantic overlap, evaluation awareness, private and temporal sets Dataset Design; Benchmark Reproducibility Contamination controls and task-detail probes
Reproducible manifests and comparable budgets Benchmark Reproducibility Run manifest and cross-harness exercise

Human evaluation & annotation

Concepts Book home Example/artifact
Reviewer selection, training, qualification, blinding Human Evaluation Annotation protocol
Categorical, ordinal, pairwise, ranking, tie, both-unacceptable Human Evaluation Response-format design table
Inter-rater reliability, disagreement, label noise, ambiguity Human Evaluation Eight-conversation pilot
Adjudication, protocol drift, sampling bias, power Human Evaluation; Metrics Adjudication record and repair exercise

LLM-as-a-judge

Concepts Book home Example/artifact
Pointwise, pairwise, ordinal, reference-guided, claim-level judging LLM as a Judge Grounded-status rubric
Position, verbosity, identity, self-preference, anchoring, correlated failure LLM as a Judge Bias-control suite
Human qualification, confusion matrix, false pass/block, abstention LLM as a Judge 120-case synthetic calibration report
Versioning, shadow comparison, authority and requalification LLM as a Judge Criterion/slice authority decision

Robustness & adversarial evaluation

Concepts Book home Example/artifact
Behavioral testing, CheckList-style capabilities, perturbations Robustness, Safety & Fairness Six-case transformation suite
Metamorphic relations and semantic-preserving validation Robustness, Safety & Fairness Paraphrase-invariance code
Prompt injection, indirect injection, jailbreaks, adaptive threats Robustness, Safety & Fairness Threat manifest and crossed attack suite
Long context, missing/stale evidence, tool faults Robustness; RAG; Agents Perturbation and timeout examples

Bias & fairness

Concepts Book home Example/artifact
Language, segment, dialect, accessibility, and workflow slices Robustness, Safety & Fairness English/Turkish crossed design
BBQ, StereoSet, RealToxicityPrompts and benchmark limitations Robustness; Benchmark Reproducibility Safety/fairness benchmark map
Disparity, support, uncertainty, case mix, intersectionality Robustness; Metrics Crossed-slice reporting checklist
Privacy and governance of attributes Robustness; Production Governed sampling guidance

Calibration & uncertainty

Concepts Book home Example/artifact
Accuracy versus confidence, reliability curves, ECE, Brier, NLL Metrics Probability-calibration section
Semantic uncertainty and repeated generation Metrics; Agent Evals Repeatability and pass@k/pass^k
Judge-to-human calibration and threshold sensitivity LLM as a Judge Calibration report
Selective prediction, abstention, escalation, risk–coverage LLM as a Judge Cascade and authority example
Simulator and offline-to-production calibration Agent Evals; Production Simulator checklist and predictive-validity metrics

Prompt evaluation & dataset curation

Concepts Book home Example/artifact
Error analysis before generic grader creation Dataset Design Trace clustering workflow
Golden, optimization, regression, adversarial, calibration, sealed, shadow roles Dataset Design Canonical seven-role table
Synthetic generation, provenance, review, deduplication Dataset Design Generation manifest
Failure mining, promotion, versioning, supersession, retirement Dataset Design; Production Incident-to-regression worked example
Dataset health, coverage, freshness, disagreement, contamination Dataset Design Health report and coverage ledger

Reproducibility & statistics

Concepts Book home Example/artifact
Estimands, denominators, units, pairing, clustered repeats Metrics Experiment estimand and paired bootstrap
Confidence intervals, rare events, power, MDE Metrics Rule-of-three explanation and design checklist
Non-inferiority, superiority, equivalence, inconclusive Metrics; Release Gates Candidate/baseline decision example
Multiplicity and exploratory versus confirmatory slices Metrics Reporting rules
Manifests, prompts, extractors, inference settings, exclusions Benchmark Reproducibility Reproducibility manifest

Tooling, MLOps & continuous evaluation

Concepts Book home Example/artifact
Promptfoo, DeepEval, Inspect, LM Evaluation Harness, LightEval System Studies; Benchmark Reproducibility Tool-selection and cross-harness exercises
Phoenix, Langfuse, Opik, MLflow, Weave, OpenTelemetry System Studies; Production Portable trace/adapter contract
Trace → dataset → evaluator → experiment → gate Production; Build Lab Production sample and release packet
CI test pyramid and cost-aware cascade Release Gates; Judge Calibration Gate policy and course transfer tasks
Licensing, self-hosting, maintenance, export, vendor lifecycle System Studies Portability checklist

RAG, agents & system evaluation

Concepts Book home Example/artifact
Retrieval recall/precision/ranking and evidence-set completeness RAG & Research Evals Versioned retrieval case
Answer correctness, groundedness, atomic facts, citations RAG & Research Evals Claim-to-evidence table
Tool selection/arguments, trajectory, partial orders, state Agent & System Evals Timeout trace and trajectory policy
Multi-turn memory, commitments, simulator, termination Agent & System Evals Session/simulator specification
Environment, harness, scaffolding, budgets, recovery Agent & System Evals; Benchmark Reproducibility System manifest and mutant validation

Interpretability

Concepts Book home Example/artifact
Explanation plausibility versus causal faithfulness Robustness, Safety & Fairness Rationale counterexample
Evidence removal, tool-result intervention, counterfactual testing Robustness; RAG Causal intervention checklist
Limits of free-form rationales and chain-of-thought as evidence Robustness Explicit non-privileged-output rule
Mechanistic method evaluation: toy ground truth, necessity/sufficiency, causal completeness, predictive intervention, stability, coverage, false discoveries Robustness, Safety & Fairness; References Known-mechanism refund example and method-evaluation table

Safety, RLHF & alignment

Concepts Book home Example/artifact
Harmful compliance, false refusal, unsafe partial compliance Robustness, Safety & Fairness Crossed benign/adversarial suite
HarmBench, JailbreakBench, StrongREJECT, XSTest Robustness; Benchmark Reproducibility Benchmark map and caveats
Preference data, reward models, Constitutional AI concepts Robustness, Safety & Fairness Preference/reward evaluation section
Reward overoptimization, Goodhart, specification gaming Robustness; Agent Evals Proxy-failure examples
Red teaming, permission boundaries, irreversible actions Robustness; Release Gates Threat manifest and hard invariants

Deep Research evaluation

Concepts Book home Example/artifact
HLE, GAIA, BrowseComp and different capability constructs RAG & Research Evals; Benchmark Reproducibility Research benchmark map
HLE dataset revision changes and GAIA answer leakage RAG & Research Evals Version-bound comparison and retrieval-integrity controls
PersonQA stale references and evaluator correction RAG & Research Evals Dated adjudication and scorer-version replay
Long-output adaptation; deployed versus capability-eliciting configurations RAG & Research Evals Separate report protocol and configuration claims
Short-answer reliability versus open-report validity RAG & Research Evals Completeness map
Source authority, diversity, contradiction, and synthesis RAG & Research Evals Source metadata and coverage table
Citation entailment, completeness, correctness, and quality RAG & Research Evals Atomic-claim grading example
Live-web fragility, stale truth, timestamps, content fingerprints RAG & Research Evals Research-run manifest
Search/test-time compute, pass@k, cost, latency, abstention RAG; Metrics Budget and risk–coverage sections
Browsing prompt injection and safety Robustness; RAG Indirect-injection controls

Observability and release gates

Operational-depth check — 7 September 2026: the source's ownership, traceability, incident, cost and resource-planning requirements have concrete teaching routes: the release chapter's requirement/owner/evidence table and expiring exception record; the production chapter's hard-rollback canary example; the telemetry and incident-promotion execution studies; and the build chapter's capacity worksheet. The evaluation-budget micro-kata closes the formula-only cost treatment with a \(960/\)1,120 worked comparison and explicit omitted-cost assumptions. This checks those operational requirements, not the source's legal dates or every external case-study claim. Hosted book conformance and deployment passed for 8c79990 in run 34092166560; that is book delivery evidence, not a deployed customer application's rollout.

Concepts Book home Example/artifact
Evals, observability, and gates as one closed loop Foundations; Production; Release Gates Four planes and incident loop
Leading/lagging signals, SLOs, cost per success Production; Metrics Signal contract
Trace lineage across data, retrieval, model, tools, policy, outcome Production Portable trace
Shadow, canary, A/B, progressive rollout, hysteresis, rollback Production; Release Gates Release manifest
Delayed outcomes, drift, sampling, asynchronous judging Production Production sample and sampling exercise
Incident response and production-to-regression promotion Production; Dataset Design Canary regression story
Governance, owner roles, privacy, retention, exception handling Release Gates; Production Exception record and release packet
Economics and evaluator cascade Metrics; Judge; Release Gates Cost-aware cascade

Course and project practice

Code-first exercise check — 7 September 2026: the source's progressive sequence maps to the typed CX runner, full trial artifacts, mixed-trace analysis (Katas 66–67), semantic qualification lab, paired slice diagnostics, incident promotion and CI gate controls. Its explicit strict/subset/superset trajectory exercise now has an executable comparison and solution. The book adapts account/invoice operations to the existing customer/order refund domain; it does not claim to have executed the source's independent twenty-trace human pilot or every suggested model/tool-description intervention. Those learner experiments remain distinct from the retained synthetic controls.

The research corpus recommends one evolving lab rather than unrelated notebooks. The CX Eval Lab is the implemented core; the projects below have different maturity levels. A required learning artifact is not automatically completed proof.

Project Book home Required learning artifact Current evidence status
Failure-analysis lab Dataset Design Actionable taxonomy and promoted cases Executed promotion plus a replayed eight-trace workshop with versioned teaching categories, overlapping counts and proposed cause tests; independent analyst-derived taxonomy and field validation remain open
Golden-dataset lab Dataset Design Version, roles, access, health report Versioned local promotion and known protected-inventory checks; complete organizational lineage and access control unproven
Human-evaluation lab Human Evaluation Protocol, labels, agreement, adjudication Katas 72–73 analyze the existing synthetic labels with confusion tables, missing/unknown support, resampling sensitivity and a reviewer-assignment negative control; actual independent human pilot and validity remain unproven
Judge-calibration lab LLM as a Judge; Semantic Grading Lab Confusion, slices, abstention, authority Executed synthetic compilation, adapter and qualification controls; Katas 74–75 add deterministic order/length/rubric diagnostics and shared-error voting controls. Actual model sensitivity and independent empirical qualification remain open
RAG evaluation lab RAG & Research Evals; Knowledge-to-Action Study Retrieval and claim attribution Executed two-document retrieval/action controls and a same-byte-budget multi-evidence packing study; natural-language interpretation and complete long-report claim/citation evaluation remain open
Agent evaluation lab Agent Evals; Build Lab Seven-surface trace and state report Executed mock CX, recovery and joined evidence paths; not demonstrated across all seven surfaces in real model workflows
Red-team lab Robustness, Safety & Fairness Threat-linked crossed suite Threat models, fixtures and local integrity controls; adaptive model red-teaming and external risk transfer remain open
Benchmark reproduction lab Benchmark Reproducibility; System Studies Cross-harness manifest and differences Worked extractor counterexample and reproduction protocol; actual pinned LM Evaluation Harness/LightEval comparison not retained
Production gate lab Production; Release Gates Shadow/canary simulation and incident receipt Executed local exposure simulation, current decision composition and incident promotion; real traffic controller and observed cloud CI remain unverified

The Courses & Learning Paths page adds the requested DeepLearning.AI courses and requires an artifact receipt for each.

The Interview & Design Drills page retains the full thirteen-theme taxonomy and converts the source bank into answer contracts, counterprompts, all eight whiteboard patterns, senior scenarios, and evidence receipts. The source report remains the exhaustive 118-question collection; the portal supplies the structured practice and full teaching chapters rather than duplicating every wording variant.

Machine-readable synthetic examples

The prose examples also have repository-local JSON companions. They are synthetic teaching data, never production evidence:

Artifact Path What its test proves
Human annotation pilot evals/cx-support/examples/human-annotations-v1.json Eight unique items, planned reviewer overlap, and explicit adjudication
Judge qualification summary evals/cx-support/examples/judge-calibration-v1.json The six confusion cells total 120 and the false-pass calculation matches the report
Evolving dataset change evals/cx-support/examples/dataset-change-v1.1.json Version lineage, regression role, incident source, and approval are present
RAG atomic claims evals/cx-support/examples/rag-claims-v1.json Correctness, evidence support, and unverifiable current truth remain separate fields
Production canary evals/cx-support/examples/production-canary-v1.json A duplicate-refund hard invariant forces rollback despite better averages
Probability calibration evals/cx-support/examples/probability-calibration-v1.json Row-level predictions recompute Brier, NLL, binned ECE, and selective thresholds
Research-to-practice evidence evals/cx-support/examples/practice-evidence-v1.json Every maturity claim retains primary sources, an explicit non-claim, local adoption rule, and bounded authority
Paired evidence receipt evals/cx-support/examples/paired-evidence-receipt-v1.json Point difference, lower-bound decision, independent-cluster count, and synthetic authority ceiling remain consistent
Advanced agent scorer fixtures evals/cx-support/examples/advanced-protocols-v1.json Factorial effects and simulator error recompute from hand-authored inputs; declared persistence, serving, skills, knowledge/action, voice, multi-agent, and integrity failures remain explicit

tests/test_book_examples.py validates these internal relations. The artifacts make the examples inspectable; they do not turn a synthetic study into an empirical result.

Coverage status and remaining proof

Layer Status What remains before the complete system claim
Concept explanation Broad treatment across Part I and later studies Maintain source-by-source depth review; a chapter mapping alone is not completion
Bite-sized synthetic examples Examples and katas across the deep chapters Close subject-specific worked-method gaps below; maintain visual/readability checks
Templates and inspectable artifacts Original JSON companions plus retained execution packets Distinguish schema/fixture consistency from actual execution and empirical transfer
Existing deterministic CX slice Running Preserve and reverify
Replay source and input consistency Source/input verification and historical replay now compose with current qualification, campaign accounting and outer decisions; Katas 49–50 and 58–59 retain the joined path Authenticate provenance and installed environment, obtain authoritative current registry inputs and bind a decision atomically to real exposure
Human annotation and judge calibration execution Full design plus consistency-tested synthetic artifacts Real reviewers/model runs are required before an empirical claim
RAG and multi-turn agent execution Executed lexical retrieval/oracle/full-context controls, scripted clarification and process-level mock-payment recovery; atomic-claim examples remain synthetic Extend to multi-evidence packing, full report evidence, compaction, agent-mediated restart and multi-turn model trials before claiming system-level qualification
Statistical release evidence Manifest, repeated paired runner, clustered teaching interval, minimum-evidence hold, and immutable receipt are executable Replace teaching samples with a registered independent measured population and method appropriate to the release estimand
Modern agent architectures Skills, voice, multi-agent, serving, factorial and simulator scorer contracts have hand-authored fixtures; containment and cross-run cache studies now execute local mechanisms Actual skill selection, delegation/merge and voice observations; matched-budget comparisons and field validation remain required
Shadow/canary production control Executed local cohort/exposure simulation plus configured CI conformance workflow Connect approved evidence to a real application controller; observe cloud gates and rollback through the deployed path

Dated depth audit — 7 September 2026

This pass reread the four supplied editable reports (deep-research-report-9, -10, -13, -15) and the complete Code-First and Evals/Observability/Release-Gates Markdown reports. It compared their substantive learning requirements with current chapters and relevant code. It does not claim a new reread of every PDF/DOCX rendering or a fresh verification of every external claim. Earlier corpus mappings remain available above.

Source requirement Current gap after the recent implementations Next completion evidence
Code-First: reporting fixed/regressed cases and comparisons by slice Katas 62–63 join legacy artifacts into slice-change lists; Katas 64–65 add native structural/joint diagnostics, explicit fixture trust and retrospective feature slicing without rewriting original release receipts Authenticate preregistration; use representative observations and qualified uncertainty/multiplicity methods before a deployment claim
Code-First: human-first error analysis Katas 66–67 replay eight mixed traces into learner views and a distinct teaching key; observations, proposed causes, overlap-aware counts and matched baseline contrasts are explicit Independent human analysis of representative traces, iterative taxonomy refinement, new targeted interventions and field validation; the authored taxonomy is not discovered human evidence
Reports 9/10: human annotation and judge validation Katas 72–73 compute synthetic item-level agreement, unknown/missing denominators, exact empirical resampling sensitivity and pass-propensity/assignment controls while preserving original adjudications Independent human sampling/annotation and validity checks; reviewer-pool uncertainty, connected incomplete-panel designs, judge order/rubric/verbosity sensitivity and actual model runs for qualification claims
Reports 9/10: judge bias and correlated errors Katas 74–75 retain 300 deterministic control calls over 60 matched presentations, identity-normalized contrasts and a panel inheriting cloned errors Independently reviewed invariance cases, repeated live-judge calls, metered pairwise transport, real cross-judge error dependence and pointwise release-grader qualification
Reports 9/10: representative and adaptive sampling Katas 70–71 implement fixed-frame proportional/risk-enriched estimation, exact inclusion weights, retained sample-to-estimate joins, false-clear counterexamples and interval qualification across all 65 binary count populations at the registered sizes Actual probability-sampled traffic and label provenance; adaptive/overlapping selection, nonresponse and noisy-label corrections, customer-level sampling, temporal transfer and real exposure decisions
Code-First: retrieval interventions Katas 68–69 execute 24 ranking/packing/action controls with duplicates, stale sources, same-body-byte budgets and an oracle diagnostic; retrieved-unit recall rises while recency-first packing loses evidence Model-backed interpretation, measured complete-request token budgets, independently qualified source authority, component ablations and representative field validation
Report 13: research-agent construct validity and temporal fragility Katas 76–77 retain compact citation/reassessment evidence; Katas 78–80 add complete 878/1,020-word authored memos, nine frozen sources, format-bound extraction controls, compound/omission diagnostics and fixed-obligation synthesis review Independent full-prose annotation and extraction qualification beyond marked findings; semantic contradiction/synthesis validity, actual research-agent comparisons, authenticated correction authority and corpus provenance
Report 15: cross-harness reproduction Kata 81 executes invented per-item disagreement and identity-join checks; boundary/likelihood/replay guidance is developed, but no reviewed external-framework run is retained in the book Audited compatible environment, pinned model/task/prompts, native per-item requests and scores across two harnesses, observed boundary differences and a tested reconciliation
Observability report: semantic telemetry through rollout Katas 83–84 execute real SDK/OTLP HTTP delivery to a bounded loopback receiver, raw-payload privacy checks, independent inventory accounting, retry/deduplication and pending feedback joins Durable authenticated collector and feedback, per-tool/distributed instrumentation, asynchronous loss/backpressure controls and a verified application rollout path

These are depth gaps in already-covered subjects, not reasons to discard the existing chapters. Paired slice reporting, mixed-trace error analysis, evidence packing, fixed-design sampling, synthetic rater analysis, judge-sensitivity controls and report/citation reassessment now have local worked artifacts. The substantial report workshop executes finding-parser diagnostics against authored reviews; it does not validate arbitrary-prose extraction or independent semantic judgment. The telemetry study now demonstrates local export loss and feedback controls, not durable distributed operation or live rollout. External framework reproduction remains pending dependency review; actual reviewer studies and model-backed transfer retain their separate prerequisites. No single study closes the entire curriculum.

This ledger distinguishes a comprehensive book from a fully implemented production platform. The book can explain and demonstrate synthetic artifacts before every platform feature is executable, but it must state that boundary honestly.

Maintenance rule

When a new research document is added:

  1. extract its substantive topics;
  2. map each topic to an existing chapter or create an explicit new home;
  3. add an example, counterexample, artifact, or exercise where useful;
  4. add primary sources to the claim ledger;
  5. record temporal or vendor-lifecycle caveats;
  6. update this page before claiming complete coverage.