Source Coverage Ledger¶
This ledger prevents the book from becoming a selective summary. It maps every substantive theme in the supplied research corpus to its teaching home, supporting example, and inspectable artifact.
What coverage means
A theme is not covered because its name appears once. Coverage requires an explanation, assumptions or limitations, a worked or counterexample, and an artifact or exercise where the topic permits one. Executable claims additionally require running-system evidence.
Source inventory¶
Course-report project sequence checked against the book¶
Teaching acceptance check — 7 September 2026: report 15's eight-row project table was compared directly with the course page's eight-lab map. Every project retains its learning purpose, worked preparation and learner receipt, including the original human-annotation and cross-framework execution requirements. No additional course-topic chapter is needed. This accepts the learning route, not completion of the learner's experiments or independent human qualification.
Repository-report teaching check: report 9's comparison, licensing, specialized-tool and architecture requirements map to System Studies' inventory, portability contract, same-refund-case comparison, judge-adapter example and solved stack-selection exercise. Vendor quick-start snippets remain in the supplied source rather than being represented as tested current commands. Tool adoption still requires the dated version/license/data-boundary checks in the study checklist. This accepts the architecture teaching coverage, not execution of every vendor integration or validation of snapshot rankings.
The source report's eight practical projects now have an explicit lab-to-exercise and receipt map: failure analysis, golden datasets, human evaluation, judge calibration, RAG, agents, red teaming and benchmark reproducibility. The map preserves the source's practice sizes while rejecting their use as qualification thresholds. It identifies independent annotation and the named two-framework execution as learner evidence still required, rather than substituting authored controls. This verifies the eight-project section's teaching route, not the entire supplied corpus's depth or completion of those empirical labs.
Durable window membership before execution¶
Kata 103 preserves all forty planned requests through a real worker kill before the first agent call. The existing execution path consumes registered cases/routes/scopes; conflicting registrations fail before work. A worked routing counterexample changes nine assignments when a canary predecessor is incorrectly replaced with expanded state. This closes immutable registration, not controller acceptance, attempt resumption or safety-stop enforcement. The registry binds a local campaign and caller-supplied manifest; hashes are consistency checks, not authenticated provenance.
Persisted completion versus committed effect¶
Kata 102 executes a real process kill after the existing mock refund commits but before its full case artifact returns. The separate completion journal retains an incomplete reservation, the campaign retains one charged effect, and replay refuses to invent a completed attempt. Completed records return identical evidence without rerunning the agent. Tests reject identity conflicts and inconsistent artifact-to-request bindings. This is one-shot completion storage, not recovery of an interrupted agent, durable window membership, controller recovery or customer-delivery proof. Existing content and historical packets remain unchanged.
Durable served-candidate routing and admission¶
Kata 101 connects the persisted backend to the existing multi-order/exposure path. Reopened shadow, canary and restricted windows retain cumulative charges of 0, 2 and 2; candidate completions are 40 shadow controls, 2/4 served and 0/4 served. Baseline and shadow do not consume served-candidate allowance. Grading reads all competing orders in one transaction. A unique durable request-start marker blocks reused and concurrent namespaces, including zero-tool and clarification-only attempts. This is admission fencing and persistent accounting, not resumable attempts, response recovery, controller checkpointing or deployment qualification. Historical artifacts remain unchanged.
Durable existing-world evidence and authorization¶
Kata 100 puts optional persisted campaign state behind the existing RefundWorld methods. Seed bindings, current authorization, refund effects and events are stored transactionally; paired state/event reads preserve grading evidence after reopen. Reference/timeout parity and specific unauthorized-action failures are tested. Review reproduced and repaired partial commits caused by broad exception handling. This is the first integration stage, not wiring through multi-order exposure, durable response storage, a process-recovery packet or production authority. Existing artifacts remain historical and unchanged.
Durable local payment allowance¶
Kata 99 adds optional database-owned policy to the existing recovery payment boundary. Approval, capacity and effect insertion share a write transaction; consumption derives from committed payments. Separate processes test competition for the final unit and worker loss before/after commit. This is a bounded local database implementation, not a durable integration of the multi-window exposure campaign or a remote payment system. Historical unbudgeted artifacts remain unchanged. The historical budgeted-packet test now pins its original file bytes and expects exact replay refusal after source/interpreter drift; fresh-packet replay remains required.
Dependence between attempts and safeguard failures¶
Kata 98 turns the existing warning about multiplied risk scores into an executable finite-population counterexample. Identical 1% marginals permit zero, one or one hundred joint failures per 10,000 opportunities. The solution separates conditional effectiveness, compatible populations, logical bounds, sampling uncertainty and consequence assumptions. These are synthetic labels and exact arithmetic, not measured harmful propensity, frontier-company risk estimates or release qualification.
Retained budgeted campaign replay¶
Kata 96 retains complete three-window execution evidence in a separately versioned budgeted packet. Fixed-builtin replay compares tool events, ledgers, namespaces, scope, budget accounting and controller decisions rather than trusting rehashed summaries. Only finite nonnegative trial timing is normalized. Exact source/interpreter matching deliberately limits portability; hashes and replay do not authenticate original execution. This closes the retained-packet follow-up below, not durable enforcement, actual model evaluation or deployment authority.
In-process campaign action budgets¶
Kata 95 now injects a shared count/per-currency allowance through the existing exposure/case/world path to the actual mock commit boundary. It separates served-candidate effects from baseline/shadow controls, enforces replay/namespace/reset rules, tests concurrent final-unit attempts and conservatively latches unknown accounting failures. The reference response path now reconciles unexpected tool status instead of blindly claiming a refund. This is single-process in-memory enforcement; the historical unbudgeted packet and CLI remain unchanged. Kata 96 supplies the retained budgeted replay packet; durable enforcement and real application integration remain open.
Historical judge example: scope and denominator clarification¶
The 120-case groundedness example now separates an imagined independent-review assumption from actual annotation evidence. Its original English qualification labels and companion JSON remain preserved as hypothetical values, explicitly not established by the aggregate table or valid for production-registry ingestion. The worked denominator comparison distinguishes 4/32, 4/40 and 4/74, plus classification coverage versus automatic acceptance and review load. No new calibration evidence is claimed.
Probability calibration versus useful decisions¶
Kata 94 extends the existing ten-row probability artifact without replacing it. A hindsight constant illustrates zero empirical ECE with no ranking, worse Brier/log loss and undefined selective risk at zero coverage. The solution distinguishes label reuse, held-out recalibration, threshold changes and review workload. Guo et al.'s primary calibration paper supplies methodological context, not local model evidence or a universal recipe for verbal confidence.
Shipping-example validity and authority¶
Kata 93 closes a teaching mismatch in the release chapter: its original short function is preserved as an explicitly unsafe counterexample, with executable NaN, unsupported-finite and missing-cost probes. The worked answer connects input validation, evidence qualification and current release authority to CI, routing, containment and recovery labs. This is a documentation/test repair, not a replacement production gate or deployed integration.
Numerical qualification and interview decision workshop¶
Katas 90–92 connect the question bank's calibration, uncertainty, data-curation and operational themes through supplied confusion matrices and executable calculations. The solutions distinguish conditional error from agreement, classification coverage from automatic acceptance, a fixed-sample zero-event bound from sequential stopping, and enriched-class sampling from representative within-class sampling. A hypothetical prevalence/queue calculation culminates in a qualification HOLD. This is an authored reasoning exercise, not new human labels, empirical model qualification or observed production staffing performance.
Executed coding-patch outcome controls¶
Katas 87–89 add a second executable domain: five authored patch controls, two visible and eight public acceptance cases, actual subprocess results, candidate diffs, protected-file checks and parent-verified outputs/input changes. The original and overfit controls both pass visible cases but differ on acceptance; the test-edit control is blocked before execution. The retained study binds local built-ins, harness/test definitions and interpreter; replay never executes its stored source. This establishes local patch-evaluation mechanics, not model-generated coding performance, sealed acceptance, independent test validity or an OS sandbox. The benchmark chapter also incorporates the dated OpenAI coding-task audit and Anthropic resource-enforcement study.
Incomplete CI evidence and integrated decision defense¶
Katas 85–86 exercise failed-run diagnostics, including a real harmless child-process timeout after two completed local controls. They separate expected candidate rejection from missing or invalid experiment evidence, and local writes from confirmed artifact storage. The capstone joins the learner's decision argument to three retained studies through 29 exact fields, with worked revocation, missingness and completed-side-effect variants. These additions deepen the operating and interview paths; they do not establish cloud workflow execution, crash-durable storage, independent human calibration or real application exposure.
Executed incident-to-regression controls¶
Katas 60–61 now reproduce a synthetic timeout-after-commit incident, apply a fixed minimization, bind a scripted review to exact proposal/source/policy identities, and publish a new regression-only version. The retained packet shows a passing reference and failing duplicate-producing mutant reconstructed from the released case, plus seven rejected promotion controls. Tests cover populated-parent preservation, recomputed content identity, exact duplicate rejection and known protected-role conflicts. This implements local promotion mechanics—not real incident ingestion, independent reviewer authentication, automatic de-identification, semantic deduplication, complete exposure history or sealed acceptance qualification.
Integrated CX evidence walkthrough¶
Katas 58–59 connect the existing native multi-order, semantic, campaign and release implementations through one retained v2 packet. The Python inspection blocks were executed; a separate clean local clone at 4ed96a7 with a newly installed locked environment reproduced the current study's diagnostic hold and six blocks. This strengthens the reproducible learner path, not model-quality evidence: the agent, judge responses, calibration, clocks, prices and prerequisite assertions remain synthetic. No independent human qualification or application exposure was performed.
Executed cross-run cache controls¶
The Cross-Run Isolation Study and Katas 56–57 execute six SQLite-backed comparisons: query-only versus run-scoped keys, each with fresh worker objects, a reused object and two concurrent threads. The retained packet contains 18 responses and 55 operations. The buggy key exposes one foreign marker in each worker mode; scoped keys expose none while preserving same-run hits, both shared-policy reads and two overwrite refusals. Replay binds phases, worker identities, concurrent thread identities and causal order to the registered contract. This is application-level cooperative-code evidence—not hostile-worker isolation, authenticated logs, model behavior or a population leakage rate.
Executed knowledge-to-action controls¶
The Knowledge-to-Action Study and Katas 51–53 add 18 local executions across lexical retrieval, oracle documents and full context, crossed with a correct structured-policy control and a wrong-amount mutant. The retained packet supports re-execution of rankings, context, decisions, rejected/accepted mock writes and final-state grades. Correct-control completion is 1/3, 3/3 and 3/3; mutant completion is 0/3, 1/3 and 1/3. All cases share one customer and two invented policy documents. Natural-language policy interpretation, active-version discovery, realistic search scale, representative sampling, model runs and deployment qualification remain open; these results do not establish a production improvement.
Current release evidence composition¶
Katas 49–50 add explicit file/value source verification and a composed current-release assessment. The retained study verifies actual committed repository bytes, replays one sixteen-trial native packet and applies seven current-state controls: current diagnostic holds; revoked, expired, synthetic-disabled, wrong-revision, changed-case and unqualified-prerequisite controls block. Four separate calibration executions and all provider responses are synthetic. Historical grades survive revocation unchanged. This closes local source/current-calibration composition—not authoritative registry retrieval, human calibration, execution attestation, atomic exposure authorization or observed cloud CI.
Executed statistical-method evidence¶
The statistical method study now supplies full count-level results for known-population interval coverage, false promotion, power, and biased-label counterexamples, plus an unequal-cluster estimand example. Its retained artifact is checked against the generator by the test suite. The alternative follows NIST's binomial-tail inversion equations, verified on 2026-09-06. This supports the narrow fixed-sample binomial lesson; general clustered and sequential method qualification remains incomplete.
Supplied research corpus¶
The Semantic Grading Lab adds tested runner integration and Katas 08–10 for scoped qualification, class-error bounds, abstention, configuration drift, expiry/revocation, and retained judge evidence. Katas 18–20 add an executed row-derived annotation compiler, a retained four-row synthetic packet, independent-adjudicator checks and declared split-leakage controls. Fixture judges, synthetic labels and reviewer identifiers are control tests, not human calibration or live-model accuracy evidence. Authenticated human annotation provenance, representative data collection and live semantic studies remain open.
Katas 36–37 add the optional Responses judge adapter: strict verdict parsing, incomplete/refusal abstention, effective-endpoint checks, defensive usage validation and retained raw response evidence. The installed SDK is exercised through an in-memory HTTP transport; four synthetic paired judgments are retained and replayed. No live semantic accuracy is established. Judge costs remain separate audit overhead, and hard campaign budgets/cancellation and reviewed live calibration are still open.
Katas 38–40 extend that implementation with a durable local estimated-cost judge-admission ledger, stable paired invocation IDs, unknown-cost reservations and selected agent/judge cost aggregation. A retained four-trial mock-world campaign contains two scripted judge calls and two admission denials; subprocess crash and multiprocess contention tests exercise persistence and atomic admission. This is not a provider invoice ceiling, whole-experiment resume, live semantic study or production release prerequisite. Agent/tool budgets, authenticated reconciliation, cancellation and full operational costs remain open.
Katas 41–42 add pinned campaign-snapshot validation against packet execution, qualification, verdicts and cost records, plus integration with the existing lab release receipt. Four retained offline controls and a CI command exercise clear/hold/block composition. Operator anchors and prerequisite/calibration data in these controls are explicit fixtures. Snapshot freshness, issuer authentication, current-calibration composition and actual canary authority remain unestablished.
| Research source | Primary contribution | Canonical treatment |
|---|---|---|
| Best GitHub Repositories for LLM Evaluation in 2026 | Tool categories, quick starts, licensing, portability, model/application/production layers | System Studies, Benchmark Reproducibility, Courses |
| LLM and AI Evaluation Interview Question Bank | Thirteen-topic evaluation taxonomy, evidence hierarchy, benchmark map, senior whiteboards | All Part I chapters, Interview & Design Drills, and exercises |
| LLM Evaluation Around OpenAI Deep Research | Research-agent capability, browsing benchmarks, safety, source authority, temporal fragility, validity | RAG & Research Evals, Robustness, Benchmark Reproducibility |
| Best Courses and YouTube Resources for Learning LLM Evaluation | Learning routes, eight practical labs, tool practice, human/statistics gaps | Courses, Part II lab, Source Coverage Ledger |
| Eval-Driven AI Programming | Error-analysis-first discipline, evaluator patterns, calibration, experiments, CI/CD, capstone | Dataset Design, Metrics, Judge Calibration, Build Lab |
| Evals, Observability and Release Gates for Production AI Systems | Control loop, telemetry, release policy, canaries, incidents, governance, ownership, economics | Production Evals and Release Gates |
| Evaluating LLM-Based AI Systems: Research Review and Production Blueprint | 2025–26 frontier methods, deployment evidence, extreme-tail reliability, stateful personalization, evaluator meta-evaluation, dynamic audits, monitorability, and realtime modalities | Research-to-Practice Evidence plus linked Part I chapters |
The supplied PDF and DOCX files have corresponding research titles, but title agreement does not establish content equivalence. The editable Markdown counterparts remain the canonical teaching inputs; retain the companion files as provenance and check for additional source links or substantive differences before treating them as redundant.
Companion-PDF preservation check — 7 September 2026: text extraction from all four supplied PDFs found substantial shared wording with the corresponding Markdown, but not identical text. The interview PDF contains all 118 canonical question IDs preserved in the appendix; an additional regex match, P02, is part of an ACL source URL, not a missing question. The PDFs also expose explicit reference URLs where some Markdown uses generated citation markers. Those links are potentially useful provenance, not duplicate curriculum topics. This check establishes question-ID preservation only: token overlap cannot establish sentence, table, answer-outline or source-link equivalence, and no PDF layout assessment or DOCX-equivalence claim is made.
Section-level completeness audit¶
Companion source-recovery receipt¶
Ranked frontier-review method map¶
The numbers below are the 46 entries in section 5 of the supplied later DOCX, not independent validations of its paper findings. Several papers motivate the same learning method; the book need not reproduce each benchmark to teach that method.
| Source entries | Worked teaching route | Acceptance boundary |
|---|---|---|
| 1, 2, 3, 6, 25, 26, 27, 28, 40, 42, 43, 44, 45 | Judge qualification, ranking reversal, lineage and sensitivity; annotation and denominator controls; calibration decision workshop | Covers local validity, reference dependence, protocol effects, correlated errors and qualified decisions; authored labels are not expert studies. |
| 4, 13, 20 | Metrics and uncertainty; statistical-method study; frontier risk decisions | Paired/clustered uncertainty and rare-event assumptions require separate qualification; no universal tail guarantee. |
| 5, 10, 11, 29, 30, 31 | Benchmark reproduction; coding-patch study; long-report study | Worked freshness, harness, scoring and construct-validity distinctions, not reproduction of all named benchmarks. |
| 7, 8, 9, 12, 34, 35, 36, 37 | Agent evaluation; recovery; modern architectures | Outcome, trajectory, subgoals, budgets and persistence; some architecture observations remain supplied fixtures. |
| 14, 15, 16, 18, 19, 32, 33 | Safety and reward optimization; frontier risk/containment; cross-run isolation | Separates capability, harmful effects, monitoring and enforcement; local controls do not establish field prevalence. |
| 17 | Assigned-cohort human-workload comparison and deployment-simulation protocol | Illustrates offline/field and user-outcome boundaries; not a personalized-product field replication. |
| 21, 22, 23, 24, 38, 39 | Existing research evaluation, knowledge-to-action and report studies | Existing decomposition and stateful evidence routes retained; further RAG expansion deferred at the user's request. |
| 41 | Uncertainty-wording paired exercise | Explicit criterion, paired wording intervention, flip-rate and class-error solution; authored counts do not establish actual model bias. |
| 46 | Voice and timing examples | Audio-condition and action-timing protocol with authored timing arithmetic, not recorded voice-model qualification. |
This closes the routing inventory and its identified entry-41 teaching gap. It does not certify every source claim, empirical transfer, or complete reproduction of the ranked studies.
Ranked-review depth check — 7 September 2026: the later DOCX's ranked evidence capsules include an explicit instance-versus-system-ranking validity distinction (entry 25). The ranking-reversal kata supplies the missing numerical counterexample and solution. The fourteen-headline map remains useful but is not a substitute for inspecting the detailed methods. This check reads the supplied claims; it does not independently verify the reported findings or reproduce all 46 studies.
The extracted companion reference index preserves 198 within-document distinct URL strings across the four PDFs and two matching DOCX files, with source filenames, SHA-256 fingerprints and extraction method. URLs can recur across documents. This is a recovery index, not a new bibliography endorsement: extraction may stop at a wrapped line, and each recovered URL still needs resolution and claim-level attribution before it supports a new factual statement. Existing references remain unchanged.
The two DOCX/Markdown comparisons also show why both formats should be retained. Markdown contains Mermaid diagrams and equation source omitted by plain-text DOCX extraction; DOCX text contains numbered reference URLs in addition to the prose. Normalized word-sequence similarity was approximately 92.5% for Code-First and 93.0% for Evals/Observability, but those values are diagnostic only—not completeness scores. These checks do not establish that every short wording change is immaterial or that diagram images are absent from the original documents. Do not discard either source on this evidence.
Later DOCX project-depth check: section 13's Judge Lineage Auditor now has a worked anonymous-output/name-intervention comparison. Its synthetic counts distinguish identity intervention, non-causal generator-group association, shared panel errors, abstention and conditional error. This closes a worked-example gap, not empirical lineage attribution or a full reread/verification of the DOCX's 46 ranked research claims.
This second map checks the structure of each source, not only the combined topic taxonomy.
Interview-bank preservation check: the complete appendix now retains all 118 unique canonical IDs across thirteen themes, with every supplied question/variant, level and answer outline checked against deep-research-report-10.md. The condensed drills remain available. This proves question/outline preservation, not that every outline is itself a full worked solution; deeper teaching remains in the mapped chapters, exercises and capstone.
Reward-model depth check: source-bank L4–L8 and L11–L12 now have a worked optimization-pressure case, with separate preference/task/safety denominators, paired counts, executable arithmetic, a block/hold solution and an experiment protocol. This closes the identified checklist-only treatment in that section; the data are authored, not empirical RLHF results.
Input identity check — 7 September 2026¶
The six supplied Markdown files were found locally; their actual first-line titles match the six corresponding corpus rows above. The numbered filenames resolve as follows. Fingerprints identify the inspected input version, not the quality or completeness of its treatment.
| Input filename | Corpus title / role | SHA-256 |
|---|---|---|
deep-research-report-9.md |
GitHub repository landscape | 00322670378566b7283ae8958bda1e74de5ecf7211e7d97d36fe8af4510a564e |
deep-research-report-10.md |
Interview question bank | b384e42874ca12487d3854877ee7f6736d7979cb749d68efe5dd9eec83021839 |
deep-research-report-13.md |
OpenAI Deep Research evaluation report | 6832298d6e379f661d708d3fa58c67f16ddd7bbdbf2d64b3f5822dcfe1a50f1c |
deep-research-report-15.md |
Courses and YouTube report | 726f05c6ac9e99e49ed3d1726c8cddc0c1f3a0fe8960cedfbd94cac92bfd50df |
Eval-Driven AI Programming A Code-First Path to Production AI Engineering.md |
Eval-driven programming | 15514e073b0b3f3d56c20e429eca462b0f8f942f5dc25f8875044984abf3da31 |
x Evals, Observability and Release Gates for Production AI Systems.md |
Production evaluation operations | 21abcf7640bc529af0dc6db41f7e3671416fa2dae14a9a9267ac726026c895d0 |
llm-evaluation-research-review-2026.docx |
Later research-review input; file presence/fingerprint checked, contents not re-audited in this check | 7156e7d910fe0cb1e29b8c262ca0ee65ed95450f1575ef2d33dc058192b637f6 |
Audit limit: the following section map is an editorial coverage assertion, not a completed independent section-by-section acceptance test. Title matches, fingerprints and broad chapter links do not prove that every source example has an adequate worked solution. Companion PDF/DOCX content equivalence was not reverified in this input check. Final content acceptance must inspect the mapped teaching sections and examples; it must not infer completeness from this inventory alone.
| Source | Source sections retained | Portal treatment |
|---|---|---|
| GitHub repository landscape | Ranked shortlist; comparison dimensions; licence implications; quick starts; specialised tools; layered architecture and next steps | System Studies keeps the complete tool inventory, selection/portability/licensing contracts, worked adapter comparison, and a cross-tool reproduction exercise. Quick-start commands remain in the source because the book teaches contracts rather than duplicating change-prone vendor syntax. |
| Interview question bank | Research taxonomy/counts; ranked core questions; 13-theme bank; evaluation architecture; whiteboards; benchmark/tool map | Thirteen Part I homes plus Interview & Design Drills, answer contract, counterprompts, whiteboards, senior scenarios, and source/reference maps. |
| Deep Research evaluation report | Identification method; first-party corpus; HLE/GAIA/BrowseComp constructs; validity trajectory; live-web fragility; safety; external research network; timeline/source map | RAG & Research Evals, Benchmark Reproducibility, Robustness, primary reference map, temporal manifests, browsing budgets, authority/citation/completeness and indirect-injection controls. |
| Courses and YouTube report | Assessment method; ranked resources; learning paths; eight practical projects; tool stack; requested-framework map; paper spine; human/statistical gaps and next steps | Courses & Learning Paths, Part II project table, System Studies inventory, References, Human Evaluation, Metrics, and judge-calibration exercises. Commercial rankings and stale star counts are not repeated as architectural evidence. |
| Eval-Driven AI Programming | EDD discipline; resource routes; typed cases/results; deterministic/reference/model/human graders; calibration; failure analysis/intervention ladder; observability/CI; curriculum/capstone; first-session 80/20 | Foundations through Release Gates, executable CX lab and mutants, prompt experiment, living dataset, evidence pyramid, courses, and staged roadmap. |
| Evals, Observability and Release Gates | Definitions/KPIs; architecture/tooling; methods/risk profiles/gate template; observability/incidents/governance/ownership/economics; operational cases; roadmap/resources/risks | Foundations, seven surfaces, Production, Release Gates, governance roles, programme economics/roadmap, System Studies case patterns, and synthetic gate/canary artifacts. |
| Evaluating LLM-Based AI Systems | Fourteen 2025–26 developments; five-layer/six-family taxonomy; ranked research map; production blueprint; demo ideas; proposed book structure | Research-to-Practice Evidence supplies the dated maturity rubric and primary-source proof matrix; Metrics covers rare events; Judge covers evaluator qualification and correction; RAG covers stateful personalization; Robustness covers auditing/monitorability; System Studies covers dynamic generation and realtime modalities. Research-only claims remain explicitly non-gating. |
Excluding a volatile ranking, price, star count, or copied quick-start is intentional: the durable concept and verification method are retained, while the live claim must be rechecked at use time.
Recent capability addendum¶
The Sequential Decisions Lab adds exact finite-horizon path enumeration and Katas 28–31. It compares fixed-final, repeated-fixed, Bonferroni and likelihood-ratio tests, with power, false-promotion and stopping-cost results plus label-bias and dependence counterexamples. The model is a narrow iid Bernoulli loss problem; it does not qualify general clustered traffic, model accuracy or release authority.
The Exposure Control Lab adds an executed 280-request routing simulation and Katas 24–27 for stable cohorts, maturity, missing labels, expiry, quality restriction, recovery hysteresis and rollback. It retains 349 deterministic mock-agent executions. It does not confer statistical or deployment qualification, enforce monetary/request caps, interrupt active work or configure a cloud application router.
The Order Resolution Study adds 16 executed deterministic trials with independent mock order ledgers, label-free unresolved inputs, customer corrections, scripted clarification and counterbalanced ordering. Katas 21–23 distinguish authorized wrong-object actions, retrospective clarification and rejected attempts hidden by later success. The live SDK path is contract-tested with a fake runner; live-model accuracy, question quality and qualified semantic grading remain unmeasured.
Katas 43–44 extend that study with two native paired comparisons (32 new deterministic agent executions). Each packet preserves all order ledgers, replays allowlisted mock-tool transitions, binds full case/design identities and recomputes structural grades. The release builder consumes the paired results with explicitly unqualified prerequisites: even the 8/8 structural control blocks. All cases share one customer; semantic qualification, live-model transfer and deployment authority remain unestablished.
Katas 45–46 add a native multi-order truth criterion, context-bound receipts and a row-derived synthetic calibration study. Three additional paired controls join structural and semantic outcomes, and a current-registry revocation assessment preserves historical replay. The deliberately literal judge and four synthetic annotation rows test interfaces; permissive diagnostic thresholds and overlapping scenarios do not establish semantic validity. Original structural-only packets retain their original evidence status.
Katas 47–48 connect the native criterion to the optional provider adapter and direct budget wrapper. The retained installed-SDK/in-memory-HTTP study executes 64 paired mock-agent trials plus four calibration executions, with 35 paired SDK requests plus four calibration requests. Known usage, 429, admission limits and reservation overrun produce campaign clear/hold/hold/block; every application release remains blocked. Request/criterion/ledger joins and separate retained-provider re-decoding tests establish local consistency, not human calibration, model accuracy, billing authenticity, current release authority or live deployment. Independent labels, untouched transfer cases, agent budgets, cancellation and incident-driven dataset operations remain open.
The CI Gate Lab adds locally executed pass/hold/block controls with full replayable packets and a configured read-only GitHub Actions workflow. Katas 16–17 separate conformance from candidate qualification and deployment. Cloud execution, branch-protection enforcement, and application canary control are not yet demonstrated.
The Frontier Risk Decisions chapter addresses the audit's risk-decision gaps using six rechecked primary sources: Anthropic's RSP v3 discussion, DeepMind's Frontier Safety Framework update and human-participant research, OpenAI's safeguards disclosure, and METR's incident investigation and task-horizon limitations. Katas 13–15 are worked reasoning/calculation exercises. Katas 54–55 add six executed local process/SQLite enforcement controls with retained effects and causal event order: early revocation, late revocation, coordinator-only cancellation, delegated revocation, false stop and benign completion. The packet records eight process lifetimes. Scripted alerts do not measure detector accuracy; hostile-code isolation, distributed revocation, participant studies and deployment qualification remain open.
The supplied reports remain the coverage baseline. This dated addendum captures capabilities and failure modes that became more prominent after parts of that corpus were written.
| Emerging capability or risk | Portal treatment | Qualification boundary |
|---|---|---|
| Scorer and evaluator isolation | Agent & System Evals separates model-action sandboxing from trusted task/scorer code and fresh scoring state | A sandbox flag does not establish host isolation; review the actual runtime and dependency boundary |
| Evaluation integrity and awareness | Benchmark Reproducibility threat-models answer keys, task identity, harness tampering, denominators, traces, and release identity | Suspicious retrieval can be a security finding without remaining a valid capability score |
| Generated dynamic behavioral evals | System Studies compares Petri, Bloom, and adaptive Giskard multi-turn scenarios | Generated cases amplify discovery; reviewed fixed cases and calibrated graders retain gate authority |
| Adaptive simulated users | Agent & System Evals and System Studies cover changing user turns, shared trace, termination, and non-collusion | Simulator behavior must be checked against held-out human sessions and is not a traffic distribution |
| Tool ownership and lifecycle | System Studies records the 2026 Petri transfer and the AgentKit lifecycle change | Reverify current owners, releases, licenses, data path, and exportability at adoption time |
| Research-to-practice maturity | Research-to-Practice Evidence separates production controls, field evidence, operational tools, and research/benchmarks | A repository or paper is never promoted to “production-proven” without a named operational use and first-party receipt |
| Extreme-tail reliability | Research-to-Practice Evidence and Metrics qualify Five-Nines/CEM importance sampling | Reproduce proposal support, weight stability, and offline-to-live validity before any local gate authority |
| Stateful personalization and memory | Research-to-Practice Evidence and RAG & Research Evals define reset/persist, counter-user, and preference-update arms | Field evidence proves an offline gap, not a universal production evaluation recipe |
| Evaluator-of-evaluators and judge correction | Research-to-Practice Evidence and LLM as a Judge cover AgentRewardBench and bias-corrected reporting | Meta-benchmarks screen candidates; representative local human evidence qualifies them |
| Monitorability and realtime voice | Research-to-Practice Evidence and Robustness/System Studies record disclosed operational use and its limits | Operational use is scoped to the named system, control, and source; it does not establish a complete safety or quality stack |
| Long-running and resumable agents | Long-Running & Serving Failures retains four-arm persistence, approval, exactly-once goals, and restart scorer contracts; Process Recovery Study adds 21 executed local trials and Katas 11–12 | Real process interruption now produces durable mock-payment evidence; compaction, distributed exactly-once claims, persistent-model behavior, and production restart evidence remain unqualified |
| Knowledge-to-action systems | Modern Agent Architectures compares retrieval, oracle documents, full context, and alternative search interfaces | Supplying the right document removes one failure source but does not qualify reasoning, action, or final state |
| Skill selection | Modern Agent Architectures separates activation, instruction freshness, and execution | Local protocol logic is tested; no live skill-loader trial has been run |
| Multi-agent coordination | Modern Agent Architectures adds delegation, coverage, duplication, handoff, merge, contention, effect, and matched-budget criteria | Buildable scorer fixtures do not establish execution coverage or superiority of a multi-agent design |
| Serving-failure behavior | Long-Running & Serving Failures covers throttling, partial streams, ambiguous commits, fallback, and degraded tools | Fallback must satisfy the same product contract; no live provider fault experiment has been run |
| Evaluation-service integrity | Eval Operations & Integrity adds fragment-merge, conflict, contamination, factorial, and simulator-scoring contracts | These are scorer/merge fixtures, not an operated distributed evaluation service |
This addendum does not replace the source-by-source audit. It extends the book while keeping the same rule: an emerging technique becomes operational evidence only after its construct, data, evaluator, environment, and authority are qualified.
Interview-taxonomy coverage¶
Item-level teaching audit: RAG and agents¶
| Source IDs | Inspected teaching evidence | Finding |
|---|---|---|
| A1, A3–A6 | Knowledge-to-action study, context-packing Katas 68–69, citation Katas 76–77 | Executed local decomposition, oracle controls and claim/citation outcomes with worked solutions. |
| A2 | Ranked-evidence calculation | Adds concrete hit/precision/recall, reciprocal-rank and nDCG calculation with judged-collection assumptions. |
| A7–A11 | Timeout-after-commit example, two valid paths exercise, process recovery study, incident rollback | State, trajectory, recovery, irreversible effects, efficiency and noncompensating release decisions remain separate. Generic agent continuation is not implemented. |
| A12 | Model/harness factorial, deployed/eliciting comparison, order-resolution study | Worked attribution and answer-leakage distinctions; fixed-control results are not evidence of emergent model capability. |
A1–A12 now have concrete teaching routes. This completes the question-family mapping, not verification of every claim in the other research reports or execution of every proposed architecture study.
Item-level teaching audit: tooling and continuous evaluation¶
| Source IDs | Inspected teaching evidence | Finding |
|---|---|---|
| M1–M2 | Controlled prompt development, test pyramid, CI Katas 16–17 | Concrete iteration and CI contracts; expected candidate rejection can be a successful software test. |
| M3 | Telemetry Delivery Study, sensitive payload controls | Executed loopback payload/receiver evidence and feedback joins; raw prompts are not indiscriminately logged. |
| M4, M7 | Matured outcome contract, Exposure Control Lab, assigned-population example | Worked missing-label, exposure, quality and human-work outcomes. Proxy drift is a trigger, not proof of changed true outcome rates. |
| M5 | Gate dependency order, unsafe shortcut counterexample, release capstone | Concrete pass/hold/block reasoning with independent constraints, evidence identity and authority. |
| M6 | Layered production sampling, weighted sampling study | Executed sampling arithmetic and decision limits rather than grading every event semantically. |
| M8 | Tool-selection matrix, stack-selection exercise | Worked contract-based choice; vendor popularity does not establish suitability or lifecycle support. |
| M9 | Incident-to-regression controls, interrupted-effect rollback | Executed local incident reproduction, reviewed-fixture promotion and bounded containment decision. |
| M10 | Service architecture, budget campaign Katas 38–40, partial-run exercise | Concrete component contracts, accounting and incomplete-pair solution. Distributed multi-tenant operation is not demonstrated by local fragments or SQLite stores. |
M1–M10 have worked teaching routes. Hosted book conformance and Pages publication succeeded at aed856e; that is not deployment of the evaluated customer-support application. The source outline's binary pass/fail is deliberately expanded to pass/hold/block so insufficient evidence is not confused with a proven candidate failure.
Item-level teaching audit: safety and alignment¶
| Source IDs | Inspected teaching evidence | Finding |
|---|---|---|
| L1, L4–L8, L11–L12 | Preference/reward workflow, optimization-pressure example and solution, reviewer-effects exercise | Separates supervision quality, reward prediction and post-optimization behavior; worked proxy/preference/task/safety contrasts and independent-evaluation protocol. No actual RLHF/RLAIF training comparison is claimed. |
| L2, L10 | Threat manifest, adaptive-budget calculation | Explicit attacker capability, budget, stopping, denominators and validity limitations. |
| L3, L9 | Crossed harmful/benign evaluation, optimization decision, noncompensable constraints | Concrete harmful-compliance/false-refusal separation and block/hold solution; neither the refusal count nor proxy reward alone decides release. |
The L1–L12 questions have explicit worked teaching routes using existing examples. Population safety, feedback validity and generalization beyond the supplied controls remain empirical claims requiring separate evidence, not properties granted by an alignment-method name.
Item-level teaching audit: interpretability¶
| Source IDs | Inspected teaching evidence | Finding |
|---|---|---|
| I1–I2 | Definitions and explanation limits, policy-passage intervention | Concrete distinction between understandable output and causal evidence; coherent rationales are not privileged access to computation. |
| I3–I5 | Method-evaluation table, known refund mechanism | Explicit intervention/control design and worked feature precision/recall; necessity, sufficiency, coverage and false discoveries remain distinct. Corrected the blanket weight-randomization claim: a control must establish what mechanism it changed before expecting an attribution change. |
| I6 | False discoveries and complementary safety evidence, frontier risk decisions | Bounded incorporation of mechanism evidence rather than substitution for behavioral and operational controls. |
The existing I1–I6 toy example and design satisfy this scoped teaching check. No neural-model interpretability experiment was run or qualified by it. Feature-set precision and recall are only proxies in the known toy mechanism, not universal measures of causal completeness.
Item-level teaching audit: fairness¶
| Source IDs | Inspected teaching evidence | Finding |
|---|---|---|
| BF1, BF5–BF8 | Fairness contract, raw/standardized case-mix solution and intersection follow-up | Worked outcome/denominator comparison, group-support limits, harm and retained-slice reasoning; no universal fairness score. |
| BF2–BF4 | Matched identity-prompt and toxicity-grader exercise | Adds a concrete comparison protocol, differential false-positive calculation and BBQ/StereoSet construct distinction. Reference labels and groups are authored, not a collected fairness study. |
BF1–BF8 now have scoped teaching routes. Actual fairness assessment still requires the affected population, justified harm criteria, reliable labels, uncertainty and appropriate governance; these examples do not establish equitable deployment.
Item-level teaching audit: robustness and adversaries¶
| Source IDs | Inspected teaching evidence | Finding |
|---|---|---|
| R1–R3, R6 | Transformation relation table, metamorphic pseudocode, conditional robustness | Concrete invariant/directional examples, execution relation and consistently-wrong caveat. |
| R4–R5 | Threat manifest, crossed benign/disallowed cases | Attacker control, prohibited effect, normal-edge variation and false-refusal checks are distinguished. |
| R7, R9–R10 | Adaptive attack-budget exercise | Adds worked scenario/query denominators, stopping, held-out families and observed-worst-group limitations. Mock results do not establish realistic attacker coverage. |
| R8 | Noncompensable invariants, incident rollback solution, capstone | Concrete containment decision and preserved committed-effect evidence rather than averaging away severe failure. |
The R1–R10 teaching routes are explicit. The new adaptive example uses authored counts; it does not execute an attack optimizer or claim empirical security qualification.
Item-level teaching audit: benchmarks¶
| Source IDs | Inspected teaching evidence | Finding |
|---|---|---|
| B1–B4 | Benchmark construct map, five validity threats, benchmark decision exercise | Benchmark definitions and concrete product-transfer counterexamples; a model score is not a refund-system release gate. |
| B5, B8 | Integrity threat and receipt, exposed-case rejection | Worked provenance/access and contamination decisions; no single black-box probe establishes absence of memorization. |
| B6 | Saturation and contamination, dataset lifecycle, weak-set redesign | Concrete refresh, role separation and preserved-regression design. No promise that any finite benchmark resists saturation indefinitely. |
| B7 | Finite-candidate pass@k derivation | Added missing subset estimator, ten-candidate numerical example, assumptions and selection-oracle limitation. |
This B1–B8 audit checks teaching coverage. It does not claim new benchmark model runs or close the separately listed cross-framework reproduction experiment.
Item-level teaching audit: fundamentals¶
| Source IDs | Inspected teaching evidence | Finding |
|---|---|---|
| F1, F10 | Evaluation contract, validity, seven-surface repair exercise | Concrete contract, evidence distinctions and solved repair decisions; training loss is not the application objective. |
| F2 | Perplexity calculation | Added missing token-loss calculation and hard-invariant counterexample. |
| F3 | Classification metrics, joint error table | Formulas, class naming, imbalance counterexample and concrete conditional-error denominators. |
| F4–F5 | Negation/paraphrase overlap exercise | Worked diagnostic and solution separate lexical overlap, contextual similarity and factual truth. No fabricated library scores. |
| F6 | Long-report study, especially Katas 78–80 | Inspectable documents, extraction/omission/compound-claim and synthesis decisions with solutions. Generic summary quality is not reduced to overlap. |
| F7 | Exact-match contract, identity-joined scorer exercise | Concrete executable classification comparison, duplicate/missing-ID rejection and normalization limits. |
| F8–F9 | Independent release rules, human-workload counterexample | Solved noncompensating safety constraints and misleading proxy-improvement scenario. |
This F1–F10 review closes the missing perplexity and concrete overlap-counterexample gaps. It establishes scoped teaching coverage, not empirical validity of every named metric on the reader's application.
Item-level teaching audit: prompts and dataset curation¶
| Source IDs | Inspected teaching evidence | Finding |
|---|---|---|
| P1, P5 | Seven dataset roles, weak-set redesign | Concrete dataset composition, access boundaries and solved release-design task. |
| P2–P3 | Controlled prompt experiment, reproducibility manifest | Versioned intervention and paired quality/safety/resource comparison; not a claim that an actual prompt was optimized successfully. |
| P4 | Crossed demonstration-set/order example | Worked averages and interaction, selection-bias defense, per-item and prompt-length requirements. |
| P6 | Synthetic manifest and review, incident promotion controls | Concrete generation metadata and executed rejection/promotion mechanics; human validity is not inferred from a manifest. |
| P7–P8 | Mixed-trace workshop, sampling study | Worked clustering/repair and weighted-sampling calculations; local controls do not demonstrate processing millions of real traces. |
| P9–P10 | Twenty-prompt selection, exposed-case rejection, assigned-population impact | Concrete selection, leakage and offline-to-impact counterexamples with solutions. No universal numeric correction for arbitrary adaptive reuse is claimed. |
This scoped P1–P10 audit adds the missing P4 worked comparison and reuses existing dataset mechanics. It establishes teaching routes, not live prompt improvement or production-data qualification.
Item-level teaching audit: LLM judges¶
| Source IDs | Inspected teaching evidence | Finding |
|---|---|---|
| J1, J3 | Judging designs, rubric schema, authority exercise | Concrete criterion/schema and decision examples distinguish relative preference from absolute acceptability. |
| J2, J5 | Katas 74–75, lineage/name intervention | Retained order/rubric/verbosity controls and worked panel/name calculations. These are not live-model bias estimates. |
| J4, J6, J8 | Katas 08–10, calibration decision workshop, 120-case denominator explanation | Worked qualification, disagreements, authority restriction, abstention and requalification; independent labels remain necessary for real qualification. |
| J7 | Generator/judge joint-error calculation | Adds the generator-conditioned escape rate, covariance, phi correlation and prevalence-dependent acceptance error missing from judge-panel comparisons alone. |
All eight questions now have explicit teaching routes with concrete examples or solved decisions. This scoped audit does not certify the rest of the source corpus or claim that the synthetic reference labels were independently collected.
Item-level teaching audit: human evaluation¶
The H1–H8 source questions were checked against the following concrete teaching sections. These are educational coverage findings, not independent-human study results.
| Source IDs | Worked teaching evidence | Scope |
|---|---|---|
| H1–H2 | Decision/task contrasts, annotation specification, weak-study redesign and solution | Defines construct, evidence, population, blinding, reviewer overlap and decision authority. |
| H3 | Anchored ordinal scale and pairwise alternatives | Concrete scale and both-unacceptable counterexample; no universal superiority claim for either format. |
| H4–H5 | Eight-case pilot, Kata 72, adjudication record | Item-level disagreement, exact nominal kappa calculation, missingness, rubric repair and preserved original labels. |
| H6 | Ratings-versus-cases planning calculation | Worked approximate precision and annotation-cost calculation; explicit independence, label-error and slice limitations. Not a general power calculator. |
| H7 | Reviewer qualification, pilot H03/H06/H07 | Distinguishes authoritative state, policy expertise, language ambiguity and subjective tone rather than declaring one reviewer universally correct. |
| H8 | Assignment-confounding control, rater-effects model and solution | Worked offset/treatment decomposition, disconnected-design non-identification and model-family cautions. No fitted empirical mixed-effects model is claimed. |
This closes the H6/H8 worked-method gaps identified in this pass. It does not establish coverage of unrelated question families or replace actual reviewer qualification.
Item-level teaching audit: calibration and statistics¶
This pass checked the retained C1–C8 and S1–S8 questions against the teaching sections below. “Worked” means the cited section supplies concrete inputs and an explained result; it does not mean empirical qualification. The findings are limited to these sixteen questions.
| Source IDs | Exact teaching evidence | Finding |
|---|---|---|
| C1–C3 | Ten probability/outcome rows and score calculations; zero-ECE counterexample | Worked definitions, probability-versus-correctness distinction, Brier/NLL/ECE calculations and failure case. |
| C4 | Threshold decisions; qualification/workload workshop | Worked coverage, selective error and review-load decisions; independent validation required. |
| C5 | Fixed-bin support, mean-confidence and accuracy table | Reliability-diagram coordinates and binning method are supplied. A plotted diagram is not supplied by this section; readers can draw the listed coordinates. |
| C6–C8 | Meaning-frequency entropy exercise; sequence likelihood normalization | Worked string/meaning distinction, length-normalization reversal and consistently-wrong counterexample. No claim to reproduce the source's theoretical hallucination theorem. |
| S1–S2, S7 | First divergent boundary, Kata 81, manifest | Concrete disagreement/join exercise, solution and retained-field contract. External-framework execution is still not demonstrated. |
| S3–S4 | Paired bootstrap code; zero-discordance counterexample | Code preserves pairing; worked counterexample explains why a narrow interval can fail. The short bootstrap function is instructional, not a validated general-purpose API. |
| S5 | Benchmark decision exercise; capstone decision defense | Worked evidence-demand and hold/block reasoning, including quality, risk, costs and scope. |
| S6 | Twenty-prompt selection exercise; sequential decision study | Worked family threshold, adjusted p-value, independence counterexample, union-bound explanation and hold/validation solution. Simultaneous selection and repeated looks are kept distinct; supplied p-values are assumed valid, not derived from model trials. |
| S8 | Repeated-trial command and five-cluster hold; unequal-cluster exercise | Worked distinction between repeats, independent customers and weighting. General hierarchical resampling with stochastic generation remains a methodological extension, not implemented by the short bootstrap snippet. |
Do not infer acceptance of the other 102 questions from this subset or treat the source answer outlines as full worked solutions.
Fundamentals & evaluation metrics¶
| Concepts | Book home | Example/artifact |
|---|---|---|
| Product objective, failure cost, evaluation contract, unit, task/case/trial/trace/outcome | Foundations | CX refund contract and vocabulary table |
| Deterministic, reference, semantic, executable, human, and online evidence | Foundations; Metrics | Evidence hierarchy and metric selection examples |
| BLEU, ROUGE, semantic similarity, task and operational metrics | Metrics | Text-metric limitations and executable-state counterexample |
| Seven independent evaluation surfaces | Agent & System Evals; Build Lab | Timeout-after-commit surface table |
Benchmarks & contamination¶
| Concepts | Book home | Example/artifact |
|---|---|---|
| MMLU, BIG-bench, GPQA, GSM8K, HumanEval, SWE-bench, HELM | Benchmark Reproducibility | Benchmark construct map |
| Saturation, prompt sensitivity, answer extraction, harness effects | Benchmark Reproducibility | Four-item two-harness worked example |
| Leakage, semantic overlap, evaluation awareness, private and temporal sets | Dataset Design; Benchmark Reproducibility | Contamination controls and task-detail probes |
| Reproducible manifests and comparable budgets | Benchmark Reproducibility | Run manifest and cross-harness exercise |
Human evaluation & annotation¶
| Concepts | Book home | Example/artifact |
|---|---|---|
| Reviewer selection, training, qualification, blinding | Human Evaluation | Annotation protocol |
| Categorical, ordinal, pairwise, ranking, tie, both-unacceptable | Human Evaluation | Response-format design table |
| Inter-rater reliability, disagreement, label noise, ambiguity | Human Evaluation | Eight-conversation pilot |
| Adjudication, protocol drift, sampling bias, power | Human Evaluation; Metrics | Adjudication record and repair exercise |
LLM-as-a-judge¶
| Concepts | Book home | Example/artifact |
|---|---|---|
| Pointwise, pairwise, ordinal, reference-guided, claim-level judging | LLM as a Judge | Grounded-status rubric |
| Position, verbosity, identity, self-preference, anchoring, correlated failure | LLM as a Judge | Bias-control suite |
| Human qualification, confusion matrix, false pass/block, abstention | LLM as a Judge | 120-case synthetic calibration report |
| Versioning, shadow comparison, authority and requalification | LLM as a Judge | Criterion/slice authority decision |
Robustness & adversarial evaluation¶
| Concepts | Book home | Example/artifact |
|---|---|---|
| Behavioral testing, CheckList-style capabilities, perturbations | Robustness, Safety & Fairness | Six-case transformation suite |
| Metamorphic relations and semantic-preserving validation | Robustness, Safety & Fairness | Paraphrase-invariance code |
| Prompt injection, indirect injection, jailbreaks, adaptive threats | Robustness, Safety & Fairness | Threat manifest and crossed attack suite |
| Long context, missing/stale evidence, tool faults | Robustness; RAG; Agents | Perturbation and timeout examples |
Bias & fairness¶
| Concepts | Book home | Example/artifact |
|---|---|---|
| Language, segment, dialect, accessibility, and workflow slices | Robustness, Safety & Fairness | English/Turkish crossed design |
| BBQ, StereoSet, RealToxicityPrompts and benchmark limitations | Robustness; Benchmark Reproducibility | Safety/fairness benchmark map |
| Disparity, support, uncertainty, case mix, intersectionality | Robustness; Metrics | Crossed-slice reporting checklist |
| Privacy and governance of attributes | Robustness; Production | Governed sampling guidance |
Calibration & uncertainty¶
| Concepts | Book home | Example/artifact |
|---|---|---|
| Accuracy versus confidence, reliability curves, ECE, Brier, NLL | Metrics | Probability-calibration section |
| Semantic uncertainty and repeated generation | Metrics; Agent Evals | Repeatability and pass@k/pass^k |
| Judge-to-human calibration and threshold sensitivity | LLM as a Judge | Calibration report |
| Selective prediction, abstention, escalation, risk–coverage | LLM as a Judge | Cascade and authority example |
| Simulator and offline-to-production calibration | Agent Evals; Production | Simulator checklist and predictive-validity metrics |
Prompt evaluation & dataset curation¶
| Concepts | Book home | Example/artifact |
|---|---|---|
| Error analysis before generic grader creation | Dataset Design | Trace clustering workflow |
| Golden, optimization, regression, adversarial, calibration, sealed, shadow roles | Dataset Design | Canonical seven-role table |
| Synthetic generation, provenance, review, deduplication | Dataset Design | Generation manifest |
| Failure mining, promotion, versioning, supersession, retirement | Dataset Design; Production | Incident-to-regression worked example |
| Dataset health, coverage, freshness, disagreement, contamination | Dataset Design | Health report and coverage ledger |
Reproducibility & statistics¶
| Concepts | Book home | Example/artifact |
|---|---|---|
| Estimands, denominators, units, pairing, clustered repeats | Metrics | Experiment estimand and paired bootstrap |
| Confidence intervals, rare events, power, MDE | Metrics | Rule-of-three explanation and design checklist |
| Non-inferiority, superiority, equivalence, inconclusive | Metrics; Release Gates | Candidate/baseline decision example |
| Multiplicity and exploratory versus confirmatory slices | Metrics | Reporting rules |
| Manifests, prompts, extractors, inference settings, exclusions | Benchmark Reproducibility | Reproducibility manifest |
Tooling, MLOps & continuous evaluation¶
| Concepts | Book home | Example/artifact |
|---|---|---|
| Promptfoo, DeepEval, Inspect, LM Evaluation Harness, LightEval | System Studies; Benchmark Reproducibility | Tool-selection and cross-harness exercises |
| Phoenix, Langfuse, Opik, MLflow, Weave, OpenTelemetry | System Studies; Production | Portable trace/adapter contract |
| Trace → dataset → evaluator → experiment → gate | Production; Build Lab | Production sample and release packet |
| CI test pyramid and cost-aware cascade | Release Gates; Judge Calibration | Gate policy and course transfer tasks |
| Licensing, self-hosting, maintenance, export, vendor lifecycle | System Studies | Portability checklist |
RAG, agents & system evaluation¶
| Concepts | Book home | Example/artifact |
|---|---|---|
| Retrieval recall/precision/ranking and evidence-set completeness | RAG & Research Evals | Versioned retrieval case |
| Answer correctness, groundedness, atomic facts, citations | RAG & Research Evals | Claim-to-evidence table |
| Tool selection/arguments, trajectory, partial orders, state | Agent & System Evals | Timeout trace and trajectory policy |
| Multi-turn memory, commitments, simulator, termination | Agent & System Evals | Session/simulator specification |
| Environment, harness, scaffolding, budgets, recovery | Agent & System Evals; Benchmark Reproducibility | System manifest and mutant validation |
Interpretability¶
| Concepts | Book home | Example/artifact |
|---|---|---|
| Explanation plausibility versus causal faithfulness | Robustness, Safety & Fairness | Rationale counterexample |
| Evidence removal, tool-result intervention, counterfactual testing | Robustness; RAG | Causal intervention checklist |
| Limits of free-form rationales and chain-of-thought as evidence | Robustness | Explicit non-privileged-output rule |
| Mechanistic method evaluation: toy ground truth, necessity/sufficiency, causal completeness, predictive intervention, stability, coverage, false discoveries | Robustness, Safety & Fairness; References | Known-mechanism refund example and method-evaluation table |
Safety, RLHF & alignment¶
| Concepts | Book home | Example/artifact |
|---|---|---|
| Harmful compliance, false refusal, unsafe partial compliance | Robustness, Safety & Fairness | Crossed benign/adversarial suite |
| HarmBench, JailbreakBench, StrongREJECT, XSTest | Robustness; Benchmark Reproducibility | Benchmark map and caveats |
| Preference data, reward models, Constitutional AI concepts | Robustness, Safety & Fairness | Preference/reward evaluation section |
| Reward overoptimization, Goodhart, specification gaming | Robustness; Agent Evals | Proxy-failure examples |
| Red teaming, permission boundaries, irreversible actions | Robustness; Release Gates | Threat manifest and hard invariants |
Deep Research evaluation¶
| Concepts | Book home | Example/artifact |
|---|---|---|
| HLE, GAIA, BrowseComp and different capability constructs | RAG & Research Evals; Benchmark Reproducibility | Research benchmark map |
| HLE dataset revision changes and GAIA answer leakage | RAG & Research Evals | Version-bound comparison and retrieval-integrity controls |
| PersonQA stale references and evaluator correction | RAG & Research Evals | Dated adjudication and scorer-version replay |
| Long-output adaptation; deployed versus capability-eliciting configurations | RAG & Research Evals | Separate report protocol and configuration claims |
| Short-answer reliability versus open-report validity | RAG & Research Evals | Completeness map |
| Source authority, diversity, contradiction, and synthesis | RAG & Research Evals | Source metadata and coverage table |
| Citation entailment, completeness, correctness, and quality | RAG & Research Evals | Atomic-claim grading example |
| Live-web fragility, stale truth, timestamps, content fingerprints | RAG & Research Evals | Research-run manifest |
| Search/test-time compute, pass@k, cost, latency, abstention | RAG; Metrics | Budget and risk–coverage sections |
| Browsing prompt injection and safety | Robustness; RAG | Indirect-injection controls |
Observability and release gates¶
Operational-depth check — 7 September 2026: the source's ownership, traceability, incident, cost and resource-planning requirements have concrete teaching routes: the release chapter's requirement/owner/evidence table and expiring exception record; the production chapter's hard-rollback canary example; the telemetry and incident-promotion execution studies; and the build chapter's capacity worksheet. The evaluation-budget micro-kata closes the formula-only cost treatment with a \(960/\)1,120 worked comparison and explicit omitted-cost assumptions. This checks those operational requirements, not the source's legal dates or every external case-study claim. Hosted book conformance and deployment passed for 8c79990 in run 34092166560; that is book delivery evidence, not a deployed customer application's rollout.
| Concepts | Book home | Example/artifact |
|---|---|---|
| Evals, observability, and gates as one closed loop | Foundations; Production; Release Gates | Four planes and incident loop |
| Leading/lagging signals, SLOs, cost per success | Production; Metrics | Signal contract |
| Trace lineage across data, retrieval, model, tools, policy, outcome | Production | Portable trace |
| Shadow, canary, A/B, progressive rollout, hysteresis, rollback | Production; Release Gates | Release manifest |
| Delayed outcomes, drift, sampling, asynchronous judging | Production | Production sample and sampling exercise |
| Incident response and production-to-regression promotion | Production; Dataset Design | Canary regression story |
| Governance, owner roles, privacy, retention, exception handling | Release Gates; Production | Exception record and release packet |
| Economics and evaluator cascade | Metrics; Judge; Release Gates | Cost-aware cascade |
Course and project practice¶
Code-first exercise check — 7 September 2026: the source's progressive sequence maps to the typed CX runner, full trial artifacts, mixed-trace analysis (Katas 66–67), semantic qualification lab, paired slice diagnostics, incident promotion and CI gate controls. Its explicit strict/subset/superset trajectory exercise now has an executable comparison and solution. The book adapts account/invoice operations to the existing customer/order refund domain; it does not claim to have executed the source's independent twenty-trace human pilot or every suggested model/tool-description intervention. Those learner experiments remain distinct from the retained synthetic controls.
The research corpus recommends one evolving lab rather than unrelated notebooks. The CX Eval Lab is the implemented core; the projects below have different maturity levels. A required learning artifact is not automatically completed proof.
| Project | Book home | Required learning artifact | Current evidence status |
|---|---|---|---|
| Failure-analysis lab | Dataset Design | Actionable taxonomy and promoted cases | Executed promotion plus a replayed eight-trace workshop with versioned teaching categories, overlapping counts and proposed cause tests; independent analyst-derived taxonomy and field validation remain open |
| Golden-dataset lab | Dataset Design | Version, roles, access, health report | Versioned local promotion and known protected-inventory checks; complete organizational lineage and access control unproven |
| Human-evaluation lab | Human Evaluation | Protocol, labels, agreement, adjudication | Katas 72–73 analyze the existing synthetic labels with confusion tables, missing/unknown support, resampling sensitivity and a reviewer-assignment negative control; actual independent human pilot and validity remain unproven |
| Judge-calibration lab | LLM as a Judge; Semantic Grading Lab | Confusion, slices, abstention, authority | Executed synthetic compilation, adapter and qualification controls; Katas 74–75 add deterministic order/length/rubric diagnostics and shared-error voting controls. Actual model sensitivity and independent empirical qualification remain open |
| RAG evaluation lab | RAG & Research Evals; Knowledge-to-Action Study | Retrieval and claim attribution | Executed two-document retrieval/action controls and a same-byte-budget multi-evidence packing study; natural-language interpretation and complete long-report claim/citation evaluation remain open |
| Agent evaluation lab | Agent Evals; Build Lab | Seven-surface trace and state report | Executed mock CX, recovery and joined evidence paths; not demonstrated across all seven surfaces in real model workflows |
| Red-team lab | Robustness, Safety & Fairness | Threat-linked crossed suite | Threat models, fixtures and local integrity controls; adaptive model red-teaming and external risk transfer remain open |
| Benchmark reproduction lab | Benchmark Reproducibility; System Studies | Cross-harness manifest and differences | Worked extractor counterexample and reproduction protocol; actual pinned LM Evaluation Harness/LightEval comparison not retained |
| Production gate lab | Production; Release Gates | Shadow/canary simulation and incident receipt | Executed local exposure simulation, current decision composition and incident promotion; real traffic controller and observed cloud CI remain unverified |
The Courses & Learning Paths page adds the requested DeepLearning.AI courses and requires an artifact receipt for each.
The Interview & Design Drills page retains the full thirteen-theme taxonomy and converts the source bank into answer contracts, counterprompts, all eight whiteboard patterns, senior scenarios, and evidence receipts. The source report remains the exhaustive 118-question collection; the portal supplies the structured practice and full teaching chapters rather than duplicating every wording variant.
Machine-readable synthetic examples¶
The prose examples also have repository-local JSON companions. They are synthetic teaching data, never production evidence:
| Artifact | Path | What its test proves |
|---|---|---|
| Human annotation pilot | evals/cx-support/examples/human-annotations-v1.json |
Eight unique items, planned reviewer overlap, and explicit adjudication |
| Judge qualification summary | evals/cx-support/examples/judge-calibration-v1.json |
The six confusion cells total 120 and the false-pass calculation matches the report |
| Evolving dataset change | evals/cx-support/examples/dataset-change-v1.1.json |
Version lineage, regression role, incident source, and approval are present |
| RAG atomic claims | evals/cx-support/examples/rag-claims-v1.json |
Correctness, evidence support, and unverifiable current truth remain separate fields |
| Production canary | evals/cx-support/examples/production-canary-v1.json |
A duplicate-refund hard invariant forces rollback despite better averages |
| Probability calibration | evals/cx-support/examples/probability-calibration-v1.json |
Row-level predictions recompute Brier, NLL, binned ECE, and selective thresholds |
| Research-to-practice evidence | evals/cx-support/examples/practice-evidence-v1.json |
Every maturity claim retains primary sources, an explicit non-claim, local adoption rule, and bounded authority |
| Paired evidence receipt | evals/cx-support/examples/paired-evidence-receipt-v1.json |
Point difference, lower-bound decision, independent-cluster count, and synthetic authority ceiling remain consistent |
| Advanced agent scorer fixtures | evals/cx-support/examples/advanced-protocols-v1.json |
Factorial effects and simulator error recompute from hand-authored inputs; declared persistence, serving, skills, knowledge/action, voice, multi-agent, and integrity failures remain explicit |
tests/test_book_examples.py validates these internal relations. The artifacts make the examples inspectable; they do not turn a synthetic study into an empirical result.
Coverage status and remaining proof¶
| Layer | Status | What remains before the complete system claim |
|---|---|---|
| Concept explanation | Broad treatment across Part I and later studies | Maintain source-by-source depth review; a chapter mapping alone is not completion |
| Bite-sized synthetic examples | Examples and katas across the deep chapters | Close subject-specific worked-method gaps below; maintain visual/readability checks |
| Templates and inspectable artifacts | Original JSON companions plus retained execution packets | Distinguish schema/fixture consistency from actual execution and empirical transfer |
| Existing deterministic CX slice | Running | Preserve and reverify |
| Replay source and input consistency | Source/input verification and historical replay now compose with current qualification, campaign accounting and outer decisions; Katas 49–50 and 58–59 retain the joined path | Authenticate provenance and installed environment, obtain authoritative current registry inputs and bind a decision atomically to real exposure |
| Human annotation and judge calibration execution | Full design plus consistency-tested synthetic artifacts | Real reviewers/model runs are required before an empirical claim |
| RAG and multi-turn agent execution | Executed lexical retrieval/oracle/full-context controls, scripted clarification and process-level mock-payment recovery; atomic-claim examples remain synthetic | Extend to multi-evidence packing, full report evidence, compaction, agent-mediated restart and multi-turn model trials before claiming system-level qualification |
| Statistical release evidence | Manifest, repeated paired runner, clustered teaching interval, minimum-evidence hold, and immutable receipt are executable | Replace teaching samples with a registered independent measured population and method appropriate to the release estimand |
| Modern agent architectures | Skills, voice, multi-agent, serving, factorial and simulator scorer contracts have hand-authored fixtures; containment and cross-run cache studies now execute local mechanisms | Actual skill selection, delegation/merge and voice observations; matched-budget comparisons and field validation remain required |
| Shadow/canary production control | Executed local cohort/exposure simulation plus configured CI conformance workflow | Connect approved evidence to a real application controller; observe cloud gates and rollback through the deployed path |
Dated depth audit — 7 September 2026¶
This pass reread the four supplied editable reports (deep-research-report-9, -10, -13, -15) and the complete Code-First and Evals/Observability/Release-Gates Markdown reports. It compared their substantive learning requirements with current chapters and relevant code. It does not claim a new reread of every PDF/DOCX rendering or a fresh verification of every external claim. Earlier corpus mappings remain available above.
| Source requirement | Current gap after the recent implementations | Next completion evidence |
|---|---|---|
| Code-First: reporting fixed/regressed cases and comparisons by slice | Katas 62–63 join legacy artifacts into slice-change lists; Katas 64–65 add native structural/joint diagnostics, explicit fixture trust and retrospective feature slicing without rewriting original release receipts | Authenticate preregistration; use representative observations and qualified uncertainty/multiplicity methods before a deployment claim |
| Code-First: human-first error analysis | Katas 66–67 replay eight mixed traces into learner views and a distinct teaching key; observations, proposed causes, overlap-aware counts and matched baseline contrasts are explicit | Independent human analysis of representative traces, iterative taxonomy refinement, new targeted interventions and field validation; the authored taxonomy is not discovered human evidence |
| Reports 9/10: human annotation and judge validation | Katas 72–73 compute synthetic item-level agreement, unknown/missing denominators, exact empirical resampling sensitivity and pass-propensity/assignment controls while preserving original adjudications | Independent human sampling/annotation and validity checks; reviewer-pool uncertainty, connected incomplete-panel designs, judge order/rubric/verbosity sensitivity and actual model runs for qualification claims |
| Reports 9/10: judge bias and correlated errors | Katas 74–75 retain 300 deterministic control calls over 60 matched presentations, identity-normalized contrasts and a panel inheriting cloned errors | Independently reviewed invariance cases, repeated live-judge calls, metered pairwise transport, real cross-judge error dependence and pointwise release-grader qualification |
| Reports 9/10: representative and adaptive sampling | Katas 70–71 implement fixed-frame proportional/risk-enriched estimation, exact inclusion weights, retained sample-to-estimate joins, false-clear counterexamples and interval qualification across all 65 binary count populations at the registered sizes | Actual probability-sampled traffic and label provenance; adaptive/overlapping selection, nonresponse and noisy-label corrections, customer-level sampling, temporal transfer and real exposure decisions |
| Code-First: retrieval interventions | Katas 68–69 execute 24 ranking/packing/action controls with duplicates, stale sources, same-body-byte budgets and an oracle diagnostic; retrieved-unit recall rises while recency-first packing loses evidence | Model-backed interpretation, measured complete-request token budgets, independently qualified source authority, component ablations and representative field validation |
| Report 13: research-agent construct validity and temporal fragility | Katas 76–77 retain compact citation/reassessment evidence; Katas 78–80 add complete 878/1,020-word authored memos, nine frozen sources, format-bound extraction controls, compound/omission diagnostics and fixed-obligation synthesis review | Independent full-prose annotation and extraction qualification beyond marked findings; semantic contradiction/synthesis validity, actual research-agent comparisons, authenticated correction authority and corpus provenance |
| Report 15: cross-harness reproduction | Kata 81 executes invented per-item disagreement and identity-join checks; boundary/likelihood/replay guidance is developed, but no reviewed external-framework run is retained in the book | Audited compatible environment, pinned model/task/prompts, native per-item requests and scores across two harnesses, observed boundary differences and a tested reconciliation |
| Observability report: semantic telemetry through rollout | Katas 83–84 execute real SDK/OTLP HTTP delivery to a bounded loopback receiver, raw-payload privacy checks, independent inventory accounting, retry/deduplication and pending feedback joins | Durable authenticated collector and feedback, per-tool/distributed instrumentation, asynchronous loss/backpressure controls and a verified application rollout path |
These are depth gaps in already-covered subjects, not reasons to discard the existing chapters. Paired slice reporting, mixed-trace error analysis, evidence packing, fixed-design sampling, synthetic rater analysis, judge-sensitivity controls and report/citation reassessment now have local worked artifacts. The substantial report workshop executes finding-parser diagnostics against authored reviews; it does not validate arbitrary-prose extraction or independent semantic judgment. The telemetry study now demonstrates local export loss and feedback controls, not durable distributed operation or live rollout. External framework reproduction remains pending dependency review; actual reviewer studies and model-backed transfer retain their separate prerequisites. No single study closes the entire curriculum.
This ledger distinguishes a comprehensive book from a fully implemented production platform. The book can explain and demonstrate synthetic artifacts before every platform feature is executable, but it must state that boundary honestly.
Maintenance rule¶
When a new research document is added:
- extract its substantive topics;
- map each topic to an existing chapter or create an explicit new home;
- add an example, counterexample, artifact, or exercise where useful;
- add primary sources to the claim ledger;
- record temporal or vendor-lifecycle caveats;
- update this page before claiming complete coverage.