The DevOps Engineer's Guide to AIOps, MLOps, and LLMOps
Buch, Englisch, Format (B × H): 155 mm x 235 mm
ISBN: 979-8-8688-3254-3
Verlag: APRESS L.P.
This book is a timely and authoritative guide for platform and DevOps engineers navigating the rapid convergence of AI and operations. As traditional responsibilities expand beyond infrastructure and CI/CD pipelines, engineers must now design and operate intelligent systems from anomaly detection pipelines and ML models to LLM-powered applications. This book provides the first unified framework that brings together AIOps, MLOps, and LLMOps, translating complex AI concepts into practical, production-ready strategies. Grounded in over a decade of real-world experience and reinforced by peer-reviewed research, it equips readers with the knowledge to build scalable, intelligent platforms using open, vendor-neutral tooling.
Structured across five comprehensive parts, the book progresses from foundational concepts to practical implementations. It begins with platform engineering fundamentals, OpenTelemetry-based observability, and AI-assisted infrastructure as code. It then dives into AIOps, covering anomaly detection, ML-driven FinOps, and AI-powered chaos engineering. The MLOps section walks through the complete model lifecycle pipelines, serving, monitoring, and drift detection, tailored specifically for platform engineers. The LLMOps section explores prompt management, RAG architectures, DevOps tooling powered by LLMs, and governance practices for secure AI systems. The final section integrates these disciplines into a unified platform architecture, complete with Kubernetes-based reference implementations, migration strategies, and organizational best practices. Each chapter includes hands-on examples in Python, Kubernetes, and Terraform, along with measurable benchmarks and reproducible projects.
By the end of this book, readers will be able to design, build, and operate a fully integrated intelligent operations platform that unifies AIOps, MLOps, and LLMOps. They will gain practical skills to deploy AI-driven systems at scale, implement observability pipelines as ML-ready data sources, manage model and LLM lifecycles in production, and evolve their organizations toward AI-first operations. This book empowers engineers to move beyond reactive DevOps toward autonomous, intelligent platforms that define the future of modern infrastructure.
What will you learn:
- Build ML-based anomaly detection pipelines from OpenTelemetry telemetry with benchmarked accuracy
- Deploy, serve, and monitor ML models on Kubernetes with automated lifecycle management
- Operate LLM applications using RAG, prompts, monitoring, and governance in production
- Apply AI techniques to FinOps, chaos engineering, IaC generation, and incident response
- Design unified platforms integrating AIOps, MLOps, and LLMOps on shared infrastructure layers
Who is it for:
This book targets DevOps engineers, platform engineers, and SREs with 3–10 years of experience who are already skilled in Kubernetes, Terraform, CI/CD, and basic Python, and are now taking on AI/ML workloads without formal training. It also supports engineering managers evaluating AIOps, MLOps, and LLMOps adoption. Readers are expected to have hands-on experience with Linux, containers, and at least one major cloud platform, while all required AI/ML concepts are taught in a practical, operations-focused context.
Zielgruppe
Professional/practitioner
Autoren/Hrsg.
Fachgebiete
Weitere Infos & Material
Part 1: Foundations.- Chapter 1: The Convergence of AI and Operations.- Chapter 2: Platform Engineering as the Foundation.- Chapter 3: Observability for Intelligent Systems.- Chapter 4: Infrastructure as Code in the AI Era.- Part 2: AIOps.- Chapter 5: Introduction to AIOps.- Chapter 6: Building Anomaly Detection Pipelines.- Chapter 7: FinOps Meets AIOps.- Chapter 8: AI-Driven Chaos Engineering.- Part 3: MLOps.- Chapter 9: MLOps Fundamentals for Platform Engineers.- Chapter 10: ML Pipeline Orchestration.- Chapter 11: Model Serving and Inference at Scale.- Chapter 12: ML Monitoring and Drift Detection.- Part 4: LLMOps.- Chapter 13: LLMOps: A New Operational Discipline.- Chapter 14: Building LLM-Powered DevOps Tools.- Chapter 15: RAG Architectures for Operations.- Chapter 16: Securing and Governing AI Systems.- Part 5: Putting It All Together.- Chapter 17: The Unified Ops Platform.- Chapter 18: The Future of Intelligent Operations.- Appendix A: Tool Comparison Matrix (AIOps/MLOps/LLMOps).- Appendix B: OpenTelemetry Quick Reference.- Appendix C: Terraform Modules for AI Infrastructure.- Appendix D: Further Reading and Research Papers.




