Modern Data Engineering

Modern Data Engineering

Build Reliable Pipelines, Warehouses, and Analytics Systems for Real-World Data WorkflowsBy Owen Fletcher
Michael Caine
Listen with Sir Michael Caine™ and 1,000+ voices
Length7h 33m

About this audiobook

**Master the foundations of modern data engineering and learn how to build reliable, scalable data systems from the ground up.** Data is one of the most valuable assets in today's organizations—but raw data alone has little value. It must be collected, organized, transformed, tested, and delivered before it can power analytics, business intelligence, dashboards, and artificial intelligence. That is where modern data engineering comes in. **Modern Data Engineering** provides a practical, beginner-friendly guide to the principles, architectures, and workflows used to build dependable data platforms. Rather than focusing on a single vendor or tool, this book teaches the core concepts that remain valuable across today's rapidly evolving data ecosystem. Inside, you'll learn how to: * Understand the role of a modern data engineer * Design reliable data pipelines from source to analytics * Work with databases, data warehouses, data lakes, and lakehouses * Build efficient ETL and ELT workflows * Ingest data from APIs, databases, files, and SaaS platforms * Store and organize data using modern file formats and partitioning strategies * Transform raw data into analytics-ready datasets using SQL * Design fact tables, dimension tables, and star schemas * Orchestrate automated workflows and pipeline scheduling * Understand batch and real-time data processing * Improve data quality through testing and validation * Implement governance, metadata management, security, and access control * Build scalable cloud-based data engineering solutions * Complete an end-to-end data engineering project suitable for your professional portfolio Whether you're an aspiring data engineer, software developer, data analyst, cloud professional, or IT student looking to expand your technical skills, this book offers a structured learning path that combines essential theory with practical design principles. By the end of the book, you'll understand how modern data platforms are built, how data flows through an organization, and how to design scalable, trustworthy data systems that support informed decision-making and future-ready applications. If you're ready to build the skills behind today's data-driven organizations, **Modern Data Engineering** is your practical guide to getting started.

Audiobook details

GenreTechnology
Length7 hrs 33 mins
Narrated byListen with 1,000+ voices
FormateBook with Audio
LanguageEnglish

Table of contents

1Modern Data Engineering
154Modeling as an Ongoing Process
2Chapter 1: What Data Engineering Is
155Chapter 10: Workflow Orchestration
3Chapter 2: The Modern Data Engineering Ecosystem
156What Orchestration Means
4Chapter 3: Understanding Data Sources
157Scheduling Pipelines
5Chapter 4: Data Ingestion
158Task Dependencies
Show all chapters
6Chapter 5: ETL, ELT, and Pipeline Design
159Directed Acyclic Graphs
7Chapter 6: Data Storage and File Formats
160Retries and Backfills
8Chapter 7: Data Warehousing
161Retries
9Chapter 8: Data Transformation with SQL
162Backfills
10Chapter 9: Data Modeling for Analytics
163Monitoring Workflow Status
11Chapter 10: Workflow Orchestration
164Designing Dependable Automated Workflows
12Chapter 11: Real-Time Data and Streaming
165Common Orchestration Mistakes
13Chapter 12: Data Quality and Testing
166Orchestration as Production Discipline
14Chapter 13: Governance, Security, and Metadata
167Chapter 11: Real-Time Data and Streaming
15Chapter 14: Cloud Data Engineering
168Batch vs Real-Time Processing
16Introduction:
169Batch Processing
17The Role of Modern Data Engineering
170Real-Time Processing
18What Modern Data Engineering Means
171Choosing Between Batch and Real Time
19Why Companies Need Reliable Data Pipelines
172Events, Messages, and Topics
20How Data Supports Analytics, AI, Dashboards, and Decision-Making
173Events
21What Readers Will Learn from This Book
174Messages
22How to Use This Book Effectively
175Topics
23Chapter 1: What Data Engineering Is
176Producers and Consumers
24The Role of a Data Engineer
177Producers
25Data Engineering vs Data Science vs Analytics
178Consumers
26How Data Moves Through an Organization
179Stream Processing Basics
27Common Responsibilities of Data Engineers
180Filtering
28Skills Needed in Modern Data Engineering
181Transformation
29The Data Engineer’s Mindset
182Enrichment
30Chapter 2: The Modern Data Engineering Ecosystem
183Aggregation
31The Modern Data Stack
184Stateful Processing
32Databases, Warehouses, Lakes, and Lakehouses
185Output
33Databases
186Common Streaming Use Cases
34Data Warehouses
187Fraud Detection
35Data Lakes
188Real-Time Monitoring and Alerts
36Lakehouses
189Product Analytics
37Comparing the Storage Options
190Personalization and Recommendations
38Batch Systems and Streaming Systems
191IoT and Sensor Data
39Batch Processing
192Financial and Trading Systems
40Streaming Processing
193Operational Dashboards
41Choosing Between Batch and Streaming
194Data Replication
42Cloud Platforms and Managed Services
195When Real-Time Data Is Necessary
43How Tools Work Together in a Data Platform
196Challenges of Late or Out-of-Order Data
44Keeping the Ecosystem Simple
197Event Time vs Processing Time
45Chapter 3: Understanding Data Sources
198Late-Arriving Events
46Application Databases
199Out-of-Order Events
47APIs and Third-Party Services
200Duplicate Events
48Files Such as CSV, JSON, and Parquet
201Ordering, State, and Reliability
49Logs and Event Data
202Designing Streaming Systems Carefully
50SaaS Platforms
203Real-Time Data in a Modern Data Platform
51Common Problems with Source Data
204Practical Mindset for Streaming
52Evaluating a Data Source
205Chapter 12: Data Quality and Testing
53Chapter 4: Data Ingestion
206Why Data Quality Matters
54What Data Ingestion Means
207Completeness, Accuracy, Freshness, and Consistency
55Full Refresh Ingestion
208Completeness
56Incremental Ingestion
209Accuracy
57Change Data Capture
210Freshness
58API-Based Ingestion
211Consistency
59File-Based Ingestion
212Null Checks and Duplicate Checks
60Avoiding Duplicates and Missing Records
213Null Checks
61Choosing the Right Ingestion Strategy
214Duplicate Checks
62Chapter 5: ETL, ELT, and Pipeline Design
215Schema Validation
63What a Data Pipeline Is
216Business Rule Testing
64ETL vs ELT
217Data Contracts
65ETL: Extract, Transform, Load
218Building Trust in Data Pipelines
66ELT: Extract, Load, Transform
219Where to Place Data Quality Checks
67Choosing Between ETL and ELT
220Handling Data Quality Failures
68Pipeline Stages
221Designing a Practical Data Testing Strategy
69Extraction
222Data Quality as a Culture
70Loading
223Chapter 13: Governance, Security, and Metadata
71Transformation
224What Data Governance Means
72Validation
225Metadata and Data Catalogs
73Publishing
226Technical Metadata
74Monitoring
227Business Metadata
75Raw, Staging, and Transformed Data
228Operational Metadata
76Raw Data
229Data Catalogs
77Staging Data
230Data Lineage
78Transformed Data
231Why Lineage Matters
79Idempotent Pipeline Design
232Upstream and Downstream Lineage
80Handling Pipeline Failures
233Column-Level Lineage
81Designing Pipelines for Maintainability
234Ownership and Documentation
82Chapter 6: Data Storage and File Formats
235Data Ownership
83Row-Based vs Column-Based Storage
236Documentation
84Row-Based Storage
237Access Control
85Column-Based Storage
238Principle of Least Privilege
86Choosing Between Row-Based and Column-Based Storage
239Role-Based Access Control
87Data Warehouses
240Attribute-Based and Policy-Based Access
88Data Lakes
241Read, Write, and Administrative Permissions
89Lakehouses
242Service Accounts and Pipeline Access
90Object Storage
243Auditing Access
91CSV, JSON, Parquet, Avro, and ORC
244Sensitive Data Handling
92CSV
245Data Classification
93JSON
246Masking and Tokenization
94Parquet
247Hashing
95Avro
248Encryption
96ORC
249Minimization
97Choosing the Right File Format
250Retention
98Partitioning and Storage Organization
251Privacy and Compliance Basics
99Designing Storage for Real-World Use
252Privacy Principles
100Chapter 7: Data Warehousing
253Compliance Considerations
101What a Data Warehouse Is
254Governance in Pipeline Design
102Why Warehouses Are Used for Analytics
255Balancing Governance and Productivity
103OLTP vs OLAP
256Common Governance Mistakes
104OLTP Systems
257The Data Engineer’s Role in Governance
105OLAP Systems
258Governance as a Foundation for Trust
106Comparing OLTP and OLAP
259Chapter 14: Cloud Data Engineering
107Fact Tables and Dimension Tables
260Why Data Engineering Moved to the Cloud
108Fact Tables
261Cloud Storage and Compute
109Dimension Tables
262Cloud Storage
110How Facts and Dimensions Work Together
263Cloud Compute
111Star Schema Basics
264Managed Databases and Warehouses
112Benefits of a Star Schema
265Managed Databases
113Star Schema Example
266Managed Data Warehouses
114Designing Warehouse Tables for Reporting
267Serverless Data Processing
115Avoiding Common Warehouse Design Mistakes
268Scaling Pipelines
116Building Trust in the Warehouse
269Scaling Storage
117Practical Warehouse Design Mindset
270Scaling Compute
118Chapter 8: Data Transformation with SQL
271Scaling Pipeline Design
119Why SQL Is Essential for Data Engineering
272Cost Control
120Cleaning Raw Data
273Storage Costs
121Joins, Aggregations, and Window Functions
274Compute Costs
122Joins
275Warehouse Query Costs
123Aggregations
276Streaming Costs
124Window Functions
277Cost Visibility
125Building Reusable Transformation Logic
278Cloud Platform Design Principles
126Creating Analytics-Ready Tables
279Design for Business Value
127Common SQL Transformation Mistakes
280Separate Storage and Compute Where Useful
128Writing SQL for Maintainability
281Use Managed Services Wisely
129The Role of SQL in the Larger Pipeline
282Build for Security from the Beginning
130Chapter 9: Data Modeling for Analytics
283Automate Repeatable Work
131Why Data Modeling Matters
284Design for Observability
132Facts, Dimensions, and Measures
285Control Costs Intentionally
133Facts
286Design for Failure
134Dimensions
287Keep Architecture Understandable
135Measures
288Practical Cloud Data Engineering Mindset
136How Facts, Dimensions, and Measures Work Together
289Chapter 15: Building a Complete Data Engineering Project
137Star Schema vs Snowflake Schema
290Planning the Final Project
138Star Schema
291Choosing Source Data
139Snowflake Schema
292Designing the Architecture
140Choosing Between Star and Snowflake
293Ingesting Raw Data
141Slowly Changing Dimensions
294Storing and Organizing Datasets
142Type 1: Overwrite the Old Value
295Transforming and Modeling Data
143Type 2: Preserve Historical Versions
296Adding Quality Checks
144Type 3: Store Limited History in Columns
297Scheduling the Pipeline
145Choosing an SCD Strategy
298Documenting the Project for a Portfolio
146Reporting Tables
299Presenting the Project Professionally
147Detail Reporting Tables
300Extending the Project
148Summary Reporting Tables
301Practical Project Checklist
149Wide Reporting Tables
302Final Project Mindset
150Designing Good Reporting Tables
303About the Author
151Designing Models for Business Users
304About the Book
152Balancing Flexibility and Simplicity
305Acknowledgement
153Common Data Modeling Mistakes
Operation Pedestal
Operation PedestalMax Hastings12h 29m$29 · $0.00
When the Heavens Went on Sale
When the Heavens Went on SaleAshlee Vance18h 20m$40 · $0.00
Hands of Time
Hands of TimeRebecca Struthers8h 8m5 (1)$26 · $0.00
Everybody Has a Podcast (Except You)
Everybody Has a Podcast (Except You)Justin McElroy, Travis McElroy, Griffin McElroy5h 9m$24 · $0.00
The Collected Works
The Collected WorksNikola Tesla50h 9m$1 · $0.00
Never Lost Again
Never Lost AgainBill Kilday10h 1m$29 · $0.00
How the Internet Happened
How the Internet HappenedBrian McCullough13h 28m$23 · $0.00
Collected Writings of Nikola Tesla
Collected Writings of Nikola TeslaNikola Tesla, Thomas Commerford Martin21h 50m$1 · $0.00
Just Aspire
Just AspireAjai Chowdhry9h 13m$29 · $0.00
The Little Book of Aliens
The Little Book of AliensAdam Frank8h 20m$26 · $0.00
The inventions, researches and writings of Nikola Tesla (Annotated)
The inventions, researches and writings of Nikola Tesla (Annotated)Thomas Commerford Martin16h 59m$2 · $0.00
Kargil
KargilV.P. Malik15h 18m$24 · $0.00
Fire on the Horizon
Fire on the HorizonTom Shroder, John Konrad8h 23m$26 · $0.00
The Smell of Kerosene (Annotated)
The Smell of Kerosene (Annotated)National Aeronautics and Space Administration, Donald L. Mallick, Peter W. Merlin11h 32m$2 · $0.00
The Boy Who Harnessed the Wind
The Boy Who Harnessed the WindWilliam Kamkwamba, Bryan Mealer10h 5m$29 · $0.00
Mars Rover Curiosity
Mars Rover CuriosityRob Manning, William L. Simon7h 43m$20
Across the Airless Wilds
Across the Airless WildsEarl Swift10h 8m$29 · $0.00
Console Wars
Console WarsBlake J. Harris20h 41m$46 · $0.00
Improvised Munitions Handbook
Improvised Munitions HandbookU.S. Department of Defense4h 38m$2 · $0.00
Escaping Gravity
Escaping GravityLori Garver10h 54m$20 · $0.00

You may also like