This hands-on training introduces developers and data engineers to Apache Beam, a unified programming model for building portable and scalable data processing pipelines, with deployment on Google Cloud Dataflow. Participants will begin with the foundations of the Beam model—its architecture, execution flow, and key abstractions—and gradually move into working with core transforms like ParDo, Map, Filter, CoGroupByKey, and Composite Transforms.
The course emphasizes both batch and streaming use cases, showcasing how to connect Beam with Google Cloud Pub/Sub, set up streaming projects on GCP, and run real-time data pipelines on Google Dataflow. Learners will also understand how Beam handles type safety, data encoding, and advanced pipeline features such as side inputs and multiple outputs. A real-world case study on identifying defaulter customers helps reinforce learning through application.
Duration: 3 Days
Course Code: BDT 508
Learning Objectives:
By the end of this course, participants will be able to:
This course is ideal for:
Module 1: Introduction to Apache Beam
Module 2: Beam Setup and Basic Concepts
Module 3: Working with Beam Transforms
Module 4: Pipeline Logic and Advanced Transforms
Module 5: Side Inputs and Outputs
Module 6: Case Study – Identifying Bank Defaulters
Module 7: Type Hints and Coders in Beam
Module 8: Introduction to Streaming in Beam
Module 9: Apache Beam with Google Cloud Dataflow
Training Material Provided: