4 days EN / DE Max 16

Data Engineering on Google Cloud

Get hands-on experience with designing and building data processing systems on Google Cloud. This course uses lectures, demos, and hands-on labs to show you how to design data processing systems, build end-to-end data pipelines, analyze data, and implement machine learning. This course covers structured, unstructured, and streaming data.

€2.900,00 excl. VAT

Individual scheduling

The courses are held as dedicated group sessions. Once you've booked, we'll coordinate a date that works for your team and send invitations to all participants.

Prerequisites

  • Understanding of data engineering principles, including ETL/ELT processes, data modeling, and common data formats (Avro, Parquet, JSON).
  • Familiarity with data architecture concepts, specifically Data Warehouses and Data Lakes.
  • Proficiency in SQL for data querying.
  • Proficiency in a common programming language (Python recommended).
  • Familiarity with using Command Line Interfaces (CLI).
  • Familiarity with core Google Cloud concepts and services (Compute, Storage, and Identity management).

What you'll learn

  • Design scalable data processing systems in Google Cloud.
  • Differentiate data architectures and implement data lakehouse and pipeline concepts.
  • Build and manage robust streaming and batch data pipelines.
  • Utilize AI/ML tools to optimize performance and gain process and data insights.

Course modules
The role of a data engineer Data sources versus data sinks Data formats Storage solution options on Google Cloud Metadata management options on Google Cloud Sharing datasets using Analytics Hub
Replication and migration architecture The gcloud command-line tool Moving datasets Datastream
Extract and load architecture The bq command-line tool BigQuery Data Transfer Service BigLake
Extract, load, and transform (ELT) architecture SQL scripting and scheduling with BigQuery Dataform
Extract, transform, and load (ETL) architecture Google Cloud GUI tools for ETL data pipelines Batch data processing using Dataproc Streaming data processing options Bigtable and data pipelines
Automation patterns and options for pipelines Cloud Scheduler and Workflows Cloud Composer Cloud Run Functions Eventarc
The classics: Data lakes and data warehouses The modern approach: Data lakehouse Choosing the right architecture
Building a data lake foundation Introduction to Apache Iceberg open table format BigQuery as the central processing engine Combining operational data in AlloyDB Combining operational and analytical data with federated queries Real world use case
BigQuery fundamentals Partitioning and clustering in BigQuery Introducing BigLake and external tables
Data governance and security in a unified platform Demo: Data Loss Prevention Analytics and machine learning on the lakehouse Real-world lakehouse architectures and migration strategies
Review Best practices
Batch data pipelines and their use cases Processing and common challenges
Design batch pipelines Large scale data transformations Dataflow and Serverless for Apache Spark Data connections and orchestration Execute an Apache Spark pipeline Optimize batch pipeline performance
Batch data validation and cleansing Log and analyze errors Schema evolution for batch pipelines Data integrity and duplication Deduplication with Serverless for Apache Spark Deduplication with Dataflow
Orchestration for batch processing Cloud Composer Unified observability Alerts and troubleshooting Visual pipeline management
Course learning objectives Course prerequisites The use case, company, challenge, and mission
Introduction to streaming data pipelines on Google Cloud Streaming ETL Streaming AI/ML Streaming applications Reverse ETL
Architectural considerations for Pub/Sub and Managed Service for Apache Kafka Dataflow: The processing powerhouse BigQuery: The analytical engine Bigtable: The solution for operational data
What you've accomplished Next steps
Data Engineering on Google Cloud