工程实战开源书高阶EN8–16 周★ 19k↑+652 章可站内阅读
Machine Learning Engineering Open Book
大规模机器学习工程开源书
Stas Bekman· 18,581 stars· 8/11 01:02 → 8/11 16:09(15 小时)
Stas Bekman 的机器学习工程开源书,聚焦大规模训练、推理、网络与编排等工程主题,强调可落地的系统实践。适合已经会训模型、想补齐训练基础设施与性能工程能力的从业者。
为什么收录 · 工程实战层架补上大规模 ML 工程开源书,填补系统性能与训练基础设施长期缺口,适合进阶读者。
ML工程分布式训练推理性能
这本书强在哪
- GitHub 高星真学习内容
- 可导入站内阅读
- 材料结构清晰
- 适合系统跟学
建议怎么学
- 01按大纲顺序推进
- 02每章留下自己的实验记录
- 03卡点时回看对应知识库文章
适合谁 / 前置
- Python 基础
- 愿意跟练代码或笔记
学完得到什么
- 掌握该教程主线知识点
- 能复现关键实验或练习
- 建立可对照的学习笔记
- 为下一阶段课程打底
目录
来自站内阅读器镜像;点击章节直接阅读
文档4 章
insights3 章
compute8 章
storage1 章
network4 章
orchestration7 章
training12 章
- 25Trainingmarkdown
- 26Model Parallelismmarkdown
- 27Software Tune Up For The Best Performancemarkdown
- 28Fault Tolerancemarkdown
- 29Reproducibilitymarkdown
- 30Avoiding, Recovering From and Understanding Instabilitiesmarkdown
- 31Understanding Training Loss Patternsmarkdown
- 32Checkpointsmarkdown
- 33Selecting Training Hyper-Parameters And Model Initializationsmarkdown
- 34Tensor Precision / Data Typesmarkdown
- 35Emulate a Multi-Node Setup Using Just a Single Nodemarkdown
- 36Re-Train HF Hub Models From Scratch Using Finetuning Examplesmarkdown
inference1 章
debug7 章
- 38Debugging and Troubleshootingmarkdown
- 39Debugging PyTorch Programsmarkdown
- 40A Backup of Scriptsmarkdown
- 45Faster Debug and Development with Tiny Models, Tokenizers and Datasetsmarkdown
- 46NCCL: Debug and Performancemarkdown
- 47Diagnosing Hangings and Deadlocks in Multi-Node Multi-GPU Python Programsmarkdown
- 48Underflow and Overflow Detectionmarkdown
testing1 章
courses1 章
model parallelism1 章
stabs2 章
策展亮点章节
ComputeTrainingInferenceNetworkOrchestrationDebug
仓库数据
Stars18,581
Forks1,198
主要语言Python
协议CC-BY-SA-4.0
创建2020/9/3
更新2026/8/11
aidebugginggpusinferencelarge-language-modelsllm
7bc7c305(master)· 许可证 需人工确认 (NOASSERTION)。 上游更新后可通过同步脚本刷新。