Liebe Besucherinnen und Besucher,
aufgrund unseres Sommerfestes sind wir am 03. September 2026 bis 14 Uhr erreichbar. Am 04. September 2026 sind wir wieder wie gewohnt für Sie da. Vielen Dank für Ihr Verständnis.
Ihr Team von Sack Fachmedien
A Handbook
Buch, Englisch, Format (B × H): 155 mm x 235 mm
ISBN: 978-981-9228-49-2
Verlag: Springer
Data quality has become a decisive foundation for large foundation models, shaping their capability, reliability, alignment, and real-world applicability. provides a systematic and practice-oriented guide to data engineering for foundation models. Moving beyond a narrow focus on large language models, the book covers the data lifecycle behind language models, vision-language models, multimodal understanding systems, text-to-image and text-to-video generative models, reasoning models, agentic systems, and domain-specific AI applications.
The book presents a full-stack framework for building high-quality data pipelines for foundation-model development. It covers large-scale pre-training data engineering, including data sourcing, acquisition, cleaning, deduplication, decontamination, tokenization, serialization, efficient loading, and quality evaluation. It also addresses multimodal data engineering for image-text, document, video, and audio data, as well as post-training and alignment data construction, including SFT, preference data, RLHF, Chain-of-Thought reasoning data, tool-use data, agent memory, and multi-turn interaction data.
The book further examines data-centric AI systems, including synthetic data factories, knowledge distillation, enterprise-grade RAG and multimodal RAG pipelines, online feedback loops, knowledge updating, DataOps platforms, data governance, privacy protection, federated learning, and compliance-aware data engineering. Through end-to-end projects and reproducible system designs, readers gain hands-on experience with distributed pre-training data pipelines, domain-specific SFT datasets, multimodal instruction data factories, reasoning data flywheels, agent tool-use data factories, enterprise DataOps platforms, privacy-preserving pipelines, open-source model reproduction, and text-to-video training data pipelines. Using modern tools such as Ray, Spark, Dask, Parquet, WebDataset, vector databases, DVC, MLflow, and Airflow, this handbook equips data engineers, MLOps and DataOps professionals, AI researchers, and technical product teams to build reliable, scalable, and continuously improving foundation-model systems.
Zielgruppe
Professional/practitioner
Autoren/Hrsg.
Fachgebiete
Weitere Infos & Material
Part 1: Foundations and Infrastructure.- Chapter 1: Data-Centric Paradigm for LLMs.- Chapter 2: LLM Data Lifecycle and Quality Framework.- Chapter 3: AI-Native Data Stack and Cost Management.- Part 2: Text Pre-training Data Engineering.- Chapter 4: Data Acquisition and Licensing for LLMs.- Chapter 5: Data Cleaning and Decontamination for LLMs.- Chapter 6: Tokenization, Serialization, and Data Loading.- Chapter 7: Data Evaluation and Iterative Improvement.- Part 3: Multimodal Data Engineering.- Chapter 8: Image-Text Data Processing.- Chapter 9: Caption Refinement and Document Understanding.- Chapter 10: Video and Audio Data Processing.- Chapter 11: Cross-Modal Alignment and Fusion.- Part 4: Instruction and Preference Data Engineering.- Chapter 12: Instruction Tuning Data Design.- Chapter 13: Preference Data and Reward Signals.- Chapter 14: Annotation Platforms and Data Quality Assurance.




