Evals Primer
A practical introduction to evaluation design, judge calibration, regression testing, and production feedback loops for AI systems.
Books
Public technical primers and longer-form drafts, available for free access and feedback.
I strongly believe books in the AI era should be continuously rewritten, much like the continual learning paradigm. These drafts are my attempt to write continual-learning books: living technical references that stay open, improve over time, and invite feedback as the field changes.
Technical primers
Focused, browsable guides to the systems and control surfaces behind production AI.
A practical introduction to evaluation design, judge calibration, regression testing, and production feedback loops for AI systems.
An architectural guide to composing models, orchestration, data, interfaces, and reliability into production AI applications.
A systems view of agent loops, harnesses, control boundaries, approvals, recovery, and dependable workflow execution.
A field guide to tool interfaces, the Model Context Protocol, permission boundaries, schemas, and reliable integration patterns.
A practical treatment of retrieval, ranking, evidence assembly, context construction, and grounded generation.
A control-oriented primer on identity, authorization, policy, auditability, and governance for AI applications and agents.
Long-form drafts
Living manuscripts that remain open in Google Docs for reading and feedback.
Notes and book draft on inference architecture, serving, latency, throughput, and GPU/cloud economics.
Book draft on production agent systems, durable harnesses, memory, tool boundaries, evals, and governance.
Book draft on RAG, context engineering, knowledge graphs, semantic connectors, and enterprise data-to-AI patterns.
Book draft on security patterns for AI systems, agents, data access, tool use, and AI application risk.
Book draft on modern LLMs, reasoning models, multimodal systems, and inference-time compute.
Book draft on reinforcement learning concepts, environments, evaluation loops, and agent training patterns.
Book draft on AI-assisted coding, agentic development workflows, code review, and software delivery loops.