AI Engineering From Scratch
分阶段 AI 工程课:测验、计划与作品产出
Rohit Ghumare · 46,253 stars · 8/8 04:44 → 8/8 14:57(10 小时)
长周期 AI 工程课:用测验、学习计划与作品集把「从基础到 RAG/Agent/多模态」串成可执行路径。内容极广,适合把它当总课表,而不是周末刷完的速成帖。认真跟完的人会留下可展示产出。
为什么收录 · 比资源列表更接近真正的工程训练营:测验、计划与作品集把长期学习变成可执行路径,适合需要外部结构管束自己的人。
这本书强在哪
- 形态接近训练营:分级测验、计划、复习队列,而不是 awesome 列表
- 章节体量极大,覆盖从基础到现代应用工程的广谱主题
- 强调作品产出,适合认真走完一条路的自学者
- 可按阶段跳读,但建议用测验决定是否跳过
建议怎么学
- 01先做定位测验,按弱项回补,不要从第一页硬啃到最后
- 02每个大阶段只保留 1 个主项目,避免并行过多
- 03RAG / Agent 阶段与站内对应专书交叉,用本书当「总课表」
- 04Capstone 前强制写架构说明与失败回顾
适合谁 / 前置
- 编程基础扎实,能长期跟练
- 愿意做测验、复习与作品集,而不是只收藏链接
- 数学与 ML 基础可边学边补
学完得到什么
- 按阶段完成从基础到 RAG / Agent / 多模态的工程训练
- 产出可展示的作品与学习记录
- 习惯用测验与复习队列管理长期学习
- 具备独立选型下一阶段专题(Infra / 训练 / Agent)的能力
目录
来自站内阅读器镜像;点击章节直接阅读
setup and tooling12 章
- 01Dev Environmentmarkdown
- 02Git & Collaborationmarkdown
- 03GPU Setup & Cloudmarkdown
- 04APIs & Keysmarkdown
- 05Jupyter Notebooksmarkdown
- 06Python Environmentsmarkdown
- 07Docker for AImarkdown
- 08Editor Setupmarkdown
另有 4 章 · 进入站内阅读查看完整目录
math foundations22 章
- 13Linear Algebra Intuitionmarkdown
- 14Vectors, Matrices & Operationsmarkdown
- 15Matrix Transformationsmarkdown
- 16Calculus for Machine Learningmarkdown
- 17Chain Rule & Automatic Differentiationmarkdown
- 18Probability and Distributionsmarkdown
- 19Bayes' Theoremmarkdown
- 20Optimizationmarkdown
另有 14 章 · 进入站内阅读查看完整目录
ml fundamentals18 章
- 35What Is Machine Learningmarkdown
- 36Linear Regressionmarkdown
- 37Logistic Regressionmarkdown
- 38Decision Trees and Random Forestsmarkdown
- 39Support Vector Machinesmarkdown
- 40K-Nearest Neighbors and Distancesmarkdown
- 41Unsupervised Learningmarkdown
- 42Feature Engineering & Selectionmarkdown
另有 10 章 · 进入站内阅读查看完整目录
deep learning core13 章
- 53The Perceptronmarkdown
- 54Multi-Layer Networks and Forward Passmarkdown
- 55Backpropagation from Scratchmarkdown
- 56Activation Functionsmarkdown
- 57Loss Functionsmarkdown
- 58Optimizersmarkdown
- 59Regularizationmarkdown
- 60Weight Initialization and Training Stabilitymarkdown
另有 5 章 · 进入站内阅读查看完整目录
computer vision28 章
- 66Image Fundamentals — Pixels, Channels, Color Spacesmarkdown
- 67Convolutions from Scratchmarkdown
- 68CNNs — LeNet to ResNetmarkdown
- 69Image Classificationmarkdown
- 70Transfer Learning & Fine-Tuningmarkdown
- 71Object Detection — YOLO from Scratchmarkdown
- 72Semantic Segmentation — U-Netmarkdown
- 73Instance Segmentation — Mask R-CNNmarkdown
另有 20 章 · 进入站内阅读查看完整目录
nlp foundations to advanced29 章
- 94Text Processing — Tokenization, Stemming, Lemmatizationmarkdown
- 95Bag of Words, TF-IDF, and Text Representationmarkdown
- 96Word Embeddings — Word2Vec from Scratchmarkdown
- 97GloVe, FastText, and Subword Embeddingsmarkdown
- 98Sentiment Analysismarkdown
- 99Named Entity Recognitionmarkdown
- 100POS Tagging and Syntactic Parsingmarkdown
- 101CNNs and RNNs for Textmarkdown
另有 21 章 · 进入站内阅读查看完整目录
speech and audio17 章
- 123Audio Fundamentals — Waveforms, Sampling, Fourier Transformmarkdown
- 124Spectrograms, Mel Scale & Audio Featuresmarkdown
- 125Audio Classification — From k-NN on MFCCs to AST and BEATsmarkdown
- 126Speech Recognition (ASR) — CTC, RNN-T, Attentionmarkdown
- 127Whisper — Architecture & Fine-Tuningmarkdown
- 128Speaker Recognition & Verificationmarkdown
- 129Text-to-Speech (TTS) — From Tacotron to F5 and Kokoromarkdown
- 130Voice Cloning & Voice Conversionmarkdown
另有 9 章 · 进入站内阅读查看完整目录
transformers deep dive16 章
- 140Why Transformers — The Problems with RNNsmarkdown
- 141Self-Attention from Scratchmarkdown
- 142Multi-Head Attentionmarkdown
- 143Positional Encoding — Sinusoidal, RoPE, ALiBimarkdown
- 144The Full Transformer — Encoder + Decodermarkdown
- 145BERT — Masked Language Modelingmarkdown
- 146GPT — Causal Language Modelingmarkdown
- 147T5, BART — Encoder-Decoder Modelsmarkdown
另有 8 章 · 进入站内阅读查看完整目录
generative ai15 章
- 156Generative Models — Taxonomy & Historymarkdown
- 157Autoencoders & Variational Autoencoders (VAE)markdown
- 158GANs — Generator vs Discriminatormarkdown
- 159Conditional GANs & Pix2Pixmarkdown
- 160StyleGANmarkdown
- 161Diffusion Models — DDPM from Scratchmarkdown
- 162Latent Diffusion & Stable Diffusionmarkdown
- 163ControlNet, LoRA & Conditioningmarkdown
另有 7 章 · 进入站内阅读查看完整目录
reinforcement learning12 章
- 171MDPs, States, Actions & Rewardsmarkdown
- 172Dynamic Programming — Policy Iteration & Value Iterationmarkdown
- 173Monte Carlo Methods — Learning from Complete Episodesmarkdown
- 174Temporal Difference — Q-Learning & SARSAmarkdown
- 175Deep Q-Networks (DQN)markdown
- 176Policy Gradient — REINFORCE from Scratchmarkdown
- 177Actor-Critic — A2C and A3Cmarkdown
- 178Proximal Policy Optimization (PPO)markdown
另有 4 章 · 进入站内阅读查看完整目录
llms from scratch24 章
- 183Tokenizers: BPE, WordPiece, SentencePiecemarkdown
- 184Building a Tokenizer from Scratchmarkdown
- 185Data Pipelines for Pre-Trainingmarkdown
- 186Pre-Training a Mini GPT (124M Parameters)markdown
- 187Scaling: Distributed Training, FSDP, DeepSpeedmarkdown
- 188Instruction Tuning (SFT)markdown
- 189RLHF: Reward Model + PPOmarkdown
- 190DPO: Direct Preference Optimizationmarkdown
另有 16 章 · 进入站内阅读查看完整目录
llm engineering17 章
- 207Prompt Engineering: Techniques & Patternsmarkdown
- 208Few-Shot, Chain-of-Thought, Tree-of-Thoughtmarkdown
- 209Structured Outputs: JSON, Schema Validation, Constrained Decodingmarkdown
- 210Embeddings & Vector Representationsmarkdown
- 211Context Engineering: Windows, Budgets, Memory, and Retrievalmarkdown
- 212RAG (Retrieval-Augmented Generation)markdown
- 213Advanced RAG (Chunking, Reranking, Hybrid Search)markdown
- 214Fine-Tuning with LoRA & QLoRAmarkdown
另有 9 章 · 进入站内阅读查看完整目录
multimodal ai25 章
- 224Vision Transformers and the Patch-Token Primitivemarkdown
- 225CLIP and Contrastive Vision-Language Pretrainingmarkdown
- 226From CLIP to BLIP-2 — Q-Former as Modality Bridgemarkdown
- 227Flamingo and Gated Cross-Attention for Few-Shot VLMsmarkdown
- 228LLaVA and Visual Instruction Tuningmarkdown
- 229Any-Resolution Vision: Patch-n'-Pack and NaFlexmarkdown
- 230Open-Weight VLM Recipes: What Actually Mattersmarkdown
- 231LLaVA-OneVision: Single-Image, Multi-Image, Video in One Modelmarkdown
另有 17 章 · 进入站内阅读查看完整目录
tools and protocols23 章
- 249The Tool Interface — Why Agents Need Structured I/Omarkdown
- 250Function Calling Deep Dive — OpenAI, Anthropic, Geminimarkdown
- 251Parallel Tool Calls and Streaming with Toolsmarkdown
- 252Structured Output — JSON Schema, Pydantic, Zod, Constrained Decodingmarkdown
- 253Tool Schema Design — Naming, Descriptions, Parameter Constraintsmarkdown
- 254MCP Fundamentals — Primitives, Lifecycle, JSON-RPC Basemarkdown
- 255Building an MCP Server — Python + TypeScript SDKsmarkdown
- 256Building an MCP Client — Discovery, Invocation, Session Managementmarkdown
另有 15 章 · 进入站内阅读查看完整目录
agent engineering42 章
- 272The Agent Loop: Observe, Think, Actmarkdown
- 273ReWOO and Plan-and-Execute: Decoupled Planningmarkdown
- 274Reflexion: Verbal Reinforcement Learningmarkdown
- 275Tree of Thoughts and LATS: Deliberate Searchmarkdown
- 276Self-Refine and CRITIC: Iterative Output Improvementmarkdown
- 277Tool Use and Function Callingmarkdown
- 278Agent Memory — Virtual Context and Memory Pagingmarkdown
- 279Memory Blocks and Sleep-Time Computemarkdown
另有 34 章 · 进入站内阅读查看完整目录
autonomous systems22 章
- 314The Shift from Chatbots to Long-Horizon Agentsmarkdown
- 315STaR, V-STaR, Quiet-STaR — Self-Taught Reasoningmarkdown
- 316AlphaEvolve — Evolutionary Coding Agentsmarkdown
- 317Darwin Godel Machine — Open-Ended Self-Modifying Agentsmarkdown
- 318AI Scientist v2 — Workshop-Level Autonomous Researchmarkdown
- 319Automated Alignment Research (Anthropic AAR)markdown
- 320Recursive Self-Improvement — Capability vs Alignmentmarkdown
- 321Bounded Self-Improvement Designsmarkdown
另有 14 章 · 进入站内阅读查看完整目录
multi agent and swarms25 章
- 336Why Multi-Agent?markdown
- 337Heritage of FIPA-ACL and Speech Actsmarkdown
- 338Communication Protocolsmarkdown
- 339The Multi-Agent Primitive Modelmarkdown
- 340Supervisor / Orchestrator-Worker Patternmarkdown
- 341Hierarchical Architecture and Its Failure Modemarkdown
- 342Society of Mind and Multi-Agent Debatemarkdown
- 343Role Specialization — Planner, Critic, Executor, Verifiermarkdown
另有 17 章 · 进入站内阅读查看完整目录
infrastructure and production28 章
- 361Managed LLM Platforms — Bedrock, Vertex AI, Azure OpenAImarkdown
- 362Inference Platform Economics — Fireworks, Together, Baseten, Modal, Replicate, Anyscalemarkdown
- 363GPU Autoscaling on Kubernetes — Karpenter, KAI Scheduler, Gang Schedulingmarkdown
- 364Serving Engine Internals — PagedAttention, Continuous Batching, Chunked Prefillmarkdown
- 365EAGLE-3 Speculative Decoding in Productionmarkdown
- 366Prefix-Cache Serving — RadixAttention and KV Reusemarkdown
- 367Hardware-Specialized Inference Compilation — FP8 and NVFP4 on Blackwellmarkdown
- 368Inference Metrics — TTFT, TPOT, ITL, Goodput, P99markdown
另有 20 章 · 进入站内阅读查看完整目录
ethics safety alignment30 章
- 389Instruction-Following as Alignment Signalmarkdown
- 390Reward Hacking and Goodhart's Lawmarkdown
- 391The Direct Preference Optimization Familymarkdown
- 392Sycophancy as RLHF Amplificationmarkdown
- 393Constitutional AI and RLAIFmarkdown
- 394Mesa-Optimization and Deceptive Alignmentmarkdown
- 395Sleeper Agents — Persistent Deceptionmarkdown
- 396In-Context Scheming in Frontier Modelsmarkdown
另有 22 章 · 进入站内阅读查看完整目录
capstone projects85 章
- 419Capstone 01 — Terminal-Native Coding Agentmarkdown
- 420Capstone 02 — RAG over Codebase (Cross-Repo Semantic Search)markdown
- 421Capstone 03 — Real-Time Voice Assistant (ASR to LLM to TTS)markdown
- 422Capstone 04 — Multimodal Document QA (Vision-First PDF, Tables, Charts)markdown
- 423Capstone 05 — Autonomous Research Agent (AI-Scientist Class)markdown
- 424Capstone 06 — DevOps Troubleshooting Agent for Kubernetesmarkdown
- 425Capstone 07 — End-to-End Fine-Tuning Pipeline (Data to SFT to DPO to Serve)markdown
- 426Capstone 08 — Production RAG Chatbot for a Regulated Verticalmarkdown
另有 77 章 · 进入站内阅读查看完整目录
策展亮点章节
仓库数据
6b10711e(main)· 许可证 MIT License (MIT)。 上游更新后可通过同步脚本刷新。