Second Edition · 2026
Book cover: a pixel grid rising through wireframe geometry and a neural network into a photorealistic hummingbird, with the title Building Vision AI, From Pixels to Generative Models

Building Vision AI From Pixels to Generative Models

A practitioner's guide to image processing, classical computer vision, deep learning, and generative vision models.

Alexander (Sasha) Apartsin, Ph.D. & Yehudit Aperstein, Ph.D.

Vision is the richest channel through which intelligence meets the world. This book is one connected journey through the theories, models, and engineering practices for systems that see, interpret, and create images. It starts with the signal-processing bedrock of image formation, builds through classical computer vision and deep learning, then moves into generative models that synthesize images, video, and 3D worlds, before closing with the evaluation, safety, and deployment concerns that govern real systems.

4 parts 39 chapters 219 sections 7 appendices & a capstone

The Four-Part Arc

Each part stands on the one before it; together they span sixty years of vision in one continuous build.

How This Book Teaches

Five habits, kept in every chapter from the first pixel to the last sample.

Worked Pipelines

Every chapter builds complete, runnable systems (a document scanner, a lane detector, a CIFAR-10 classifier end to end), never isolated snippets.

Library Shortcuts

After each from-scratch build, a shortcut callout shows the same task in a few lines of OpenCV, scikit-image, PyTorch, or diffusers, and names exactly what the library handles for you.

A Callout System

Pitfalls, math asides, practical industry examples, and cross-references are typeset as distinct boxes, so you can read deep or skim fast and never miss a trap.

Exercises & Labs

Each chapter closes with hands-on exercises that extend its worked pipelines, from quick checks to small projects you can put in a portfolio.

Classical Ideas Return Learned

Convolution becomes the CNN layer, denoising becomes diffusion, inpainting becomes generative editing, and multi-view geometry returns in NeRF. One story, told twice.

The Hands-On AI Science Series

Building Vision AI is one of nine connected books, each a deep, build-it-yourself guide to a major field of AI.

Hands-On AI Science is a series of in-depth guides to the major fields of artificial intelligence. Every book goes deep into the theory, models, and internals, covering the classical foundations and the most recent ideas, then shows you how to build each one in Python with the modern libraries and tools that get the job done. The writing stays plain and light (illustrations, analogies, mental models, worked examples, and a little fun) without trading away rigor or coverage. Each volume is self-contained and complete enough to anchor a full course on its subject.

Building Language AI

From Tokens to Agents.

Read online · Kindle

Building Vision AI

From Pixels to Generative Models.

You are here

Building Temporal AI

From Forecasting to Sequential Decision Making.

Read online · Kindle

Building Scalable AI

From Big Data Algorithms to Distributed Intelligence.

Read online · Kindle

Building Embodied AI

From Perception to Autonomous Action.

Read online

Building Agentic AI

From Goals to Autonomous Systems.

Read online

Building Discovery AI

From Vibe Coding to Autonomous Science.

Read online

Building Neuromorphic AI

From Spiking Neurons to Edge Intelligence.

Read online

Building Quantum AI

From Qubits to Quantum Machine Learning.

Read online

Read the full About the Hands-On AI Science Series note.