The article is ~746 words. Here it is in HTML format:

A new textbook from Harvard University is reframing how engineers think about machine learning systems. Titled "Machine Learning Systems" and subtitled "The physics of AI engineering," the two-volume work by Vijay Janapa Reddi treats ML infrastructure not as a collection of tools and frameworks but as an engineering discipline governed by physical constraints. It is free to read online under a Creative Commons license, published by MIT Press, and has already attracted more than 367,000 readers across 210 countries.

From Algorithms to Engineering

The premise behind the book is straightforward. Students learn how to train machine learning models, but fewer are taught how to engineer the systems that make those models useful in production. As AI capabilities grow, progress will depend less on developing new algorithms and more on building engineers who can design systems that are scalable, efficient, and reliable. The textbook bridges that gap with a principles-first approach, giving readers quantitative tools to diagnose bottlenecks, predict trade-offs, and reason about system performance from first principles rather than memorizing deployment recipes.

The author frames the central challenge with what the book calls the Iron Law of ML Systems: total latency decomposes into three measurable components, a data term constrained by memory bandwidth, a compute term bounded by hardware utilization and efficiency, and a latency term capturing orchestration overhead. This equation anchors the entire curriculum, reminding readers that every ML system ultimately obeys the laws of physics.

Two Volumes, Four Stages

Volume I, subtitled "Introduction to Machine Learning Systems," covers foundations, build, optimize, and deploy on single machines. It walks readers through data engineering, neural network computation and architectures, framework internals, training infrastructure, model compression, hardware acceleration, benchmarking, serving systems, and responsible engineering. Volume II, "Machine Learning Systems at Scale," extends the same principles to distributed infrastructure, covering multi-machine training, fleet operations, fault tolerance, and governance.

Each volume progresses through four stages. The first builds conceptual foundations with mental models that underpin all systems work. The second walks through complete workflows from data pipelines through training. The third transforms theoretical understanding into systems that run efficiently under real resource constraints. The fourth navigates serving infrastructure, operations, and responsible engineering practices including fairness, privacy, security, and environmental sustainability treated as engineering problems with measurable solutions.

An Ecosystem of Learning Tools

The textbook is accompanied by a suite of interactive resources. TinyTorch lets readers build an ML framework from scratch across 20 progressive modules with no abstraction magic. MLSys·im provides first-principles performance modeling, taking a single command and producing a detailed breakdown of every bottleneck in a given hardware configuration. Interactive Marimo notebooks allow learners to change a parameter and immediately see what breaks, building intuition about trade-offs through experimentation.

For hands-on deployment, hardware kits support projects targeting Arduino, Seeed, Grove, and Raspberry Pi platforms, where learners confront real memory limits and power budgets rather than simulated constraints. StaffML offers physics-grounded interview questions, drills, and mock interviews tailored to ML systems roles, complete with vaults and progress tracking.

Built for the Classroom

The project includes an Instructor Hub with the AI Engineering Blueprint, a two-semester curriculum complete with syllabi, pedagogy guides, rubrics, and a TA handbook. Thirty-five lecture slide decks include speaker notes and 281 original SVG diagrams. The material is already derived from Harvard's CS249r course and the TinyML edX program.

The project's stated mission is to educate one million AI engineers by 2030. With 28,000 GitHub stars, 32,000 average monthly readers, and contributors from around the world already submitting pull requests, the curriculum is being shaped by a broad community rather than a single author. The hardcover edition arrives through MIT Press in November 2026, priced at $135.

What Sets It Apart

Several existing books address MLOps or machine learning systems at a practitioner level, but this text distinguishes itself by refusing to treat tools as the subject. The instruction emphasizes enduring principles over current frameworks, quantitative reasoning over hand-waving, and physical constraints over feature checklists. The author argues that MLOps guides tell you how to wire up a feature store and a pipeline with today's tools, whereas this textbook teaches you to reason about why those tools exist and where they break.

That philosophy extends to the companion tools themselves. TinyTorch exists so learners understand a framework by building one. MLSys·im exists so learners model constraints before choosing hardware. The hardware kits exist so learners deploy under real conditions. Every component of the curriculum is designed to make the invisible physics of ML engineering visible.