bojieli/ai-infra-book

This is a free, open-source book (available in PDF and EPUB) that teaches engineers how AI systems — like the ones powering ChatGPT or DeepSeek — are actually built and optimized under the hood, using math and real hardware constraints to explain design decisions. It covers everything from how AI models are run efficiently on specialized chips to how large-scale distributed systems coordinate thousands of processors working together.

4.8k3504 contributorsPythonsource ↗

§ 1 — what it does

This is a free, open-source book (available in PDF and EPUB) that teaches engineers how AI systems — like the ones powering ChatGPT or DeepSeek — are actually built and optimized under the hood, using math and real hardware constraints to explain design decisions. It covers everything from how AI models are run efficiently on specialized chips to how large-scale distributed systems coordinate thousands of processors working together.

§ 2 — why it matters

As AI infrastructure costs and performance become major competitive differentiators, founders and technical leaders who deeply understand these systems can make smarter build-vs-buy decisions, optimize spending on AI compute, and design products around realistic latency and cost constraints. The book's rapid adoption (4,700+ stars in a short time) signals strong demand from engineers trying to close the knowledge gap between using AI APIs and understanding what drives their cost and speed.

§ 3 — why it’s trending

As AI infrastructure becomes one of the most sought-after skill sets in tech, engineers are clearly hungry for resources that go beyond surface-level tutorials — this open-source book by Li Bojie picked up 4,620 stars in a single week, suggesting it spread rapidly through engineering communities and team Slack channels rather than any single viral moment. The 273 commits over the past 30 days signal that this isn't a static PDF dump but an actively maintained project, which likely explains why people are bookmarking it rather than just clicking and moving on. With 204 new forks this week, builders aren't just reading — they're pulling it down locally, running the companion calculation tools, and treating it as a working reference for real system design decisions.

§ 4 — related entries

4 entries

ROCm/ATOM

70/100

Breakout

ATOM is an open-source tool that makes it faster and easier to run AI language models on AMD hardware, offering similar capabilities to popular AI serving systems but optimized specifically for AMD's chip ecosystem. Think of it as a performance-tuned engine that sits between your AI application and AMD's hardware, making sure the models run as efficiently as possible.

why it matters: As businesses look to reduce dependence on Nvidia's dominant AI chips, tools like ATOM that unlock AMD hardware for AI workloads become strategically valuable — potentially offering cost savings and supply chain flexibility. For builders evaluating infrastructure choices, this signals a maturing AMD AI ecosystem that could soon offer a credible alternative for deploying AI-powered products at scale.

184149100 contributorsPython

PyTorch is the leading open-source framework used to build and train AI models, powering everything from image recognition to large language models like the ones behind ChatGPT-style products. It gives developers a flexible, Python-based environment to experiment with and deploy neural networks — the underlying technology that enables machines to learn from data.

why it matters: With over 100,000 stars and 6,600 contributors, PyTorch has become the de facto standard for AI research and production, meaning most cutting-edge AI products being built today are likely running on it. For founders and investors, understanding PyTorch adoption is a strong signal of serious AI development — and building familiarity with its ecosystem is increasingly a strategic advantage as AI becomes central to nearly every product category.

103k30.0k6.6k contributorsPython

ROCm/aiter

65/100

Hot

AITER is AMD's open-source software library that makes AI workloads run faster on AMD graphics cards, acting as a performance layer between AI frameworks and AMD hardware. Think of it as a set of highly optimized building blocks that AI software can use to squeeze maximum speed out of AMD GPUs when running or training AI models.

why it matters: As AI infrastructure costs soar, AMD GPUs represent a real alternative to Nvidia's dominance, and AITER is the critical software glue that makes that hardware viable for production AI products — giving builders a second competitive supplier to negotiate against. With 200 contributors and strong adoption signals, this project signals that the AMD AI ecosystem is maturing fast, which matters for anyone making long-term bets on AI infrastructure costs and availability.

565585200 contributorsPython

trycua/cua

64/100

Hot

Cua is an open-source platform that gives AI agents their own isolated computers to operate — including virtual Mac and Windows machines — so they can browse the web, click buttons, fill forms, and use software just like a human would. It also includes tools for testing and training these AI agents to ensure they perform reliably across different tasks.

why it matters: As AI agents move from answering questions to actually doing work on computers, builders need infrastructure to run and evaluate those agents safely at scale — Cua provides exactly that, positioning itself as a foundational layer for the next wave of AI-powered automation products. With nearly 25,000 stars and over 100 contributors, it signals strong developer momentum in a space where companies like Anthropic, OpenAI, and startups are racing to own the 'AI that uses computers' category.

25.5k1.8k102 contributorsHTML

form 27-b — subscription

THE TUESDAY BRIEFING

The repos that moved this week, why they matter, and what to watch next. One email. No noise.