This hands-on developer-focused training provides comprehensive coverage of the Hadoop and Apache Spark ecosystems, enabling participants to build and optimize Big Data applications using Scala or Python. The training begins with Hadoop essentials—HDFS, YARN, MapReduce, Hive, Sqoop, Pig, HBase, Flume—and progresses into Apache Spark, covering RDDs, SparkSQL, optimization techniques, and iterative algorithms including machine learning use cases.
Participants will work through practical examples and real-world scenarios, gaining fluency in working with large-scale data processing tools, ingest mechanisms, and querying frameworks. The course is language-flexible and supports parallel development in both Scala and Python, making it ideal for developers, data engineers, and analysts transitioning into Big Data platforms.
Duration: 5 Days
Course Code: BDT 505
Learning Objectives:
By the end of this course, participants will be able to:
This course is ideal for:
Module 1: Big Data and the Hadoop Ecosystem
Module 2: HDFS and Hadoop Architecture
Module 3: MapReduce and Data Ingestion with Sqoop
Module 4: Hive and Impala Basics
Module 5: Working with Hive and Impala
Module 6: Data Formats and Avro Schema Handling
Module 7: Advanced Hive Concepts and Partitioning
Module 8: Apache Flume and HBase
Module 9: Introduction to Apache Pig
Module 10: Apache Spark Fundamentals
Module 11: Deeper Dive into Spark RDDs
Module 12: Developing Spark Applications
Module 13: Spark Parallelism and Execution
Module 14: Spark RDD Optimization Techniques
Module 15: Spark Algorithms and ML Use Cases
Module 16: Spark SQL and DataFrames
Training Material Provided: