This course will teach participants how to use Apache Hadoop and Apache Spark to solve sophisticated data science problems, producing valuable insights in a wide range of scenarios.
Day one focuses on data science basics, including data acquisition, scrubbing and manipulation, as well as a general overview of data science applications as well as the analytics and machine learning processes typically employed. A number of practical use cases are examined during class and lab sessions.
Day two focuses on Apache Hadoop and its ecosystem along with the types of data science applications typically handled by the Hadoop platform. The course outlines the statistical methods used to produce actionable business insights with MapReduce, Python, Hive and other tools.
Day three begins with an overview of the Apache Spark platform and its machine learning library, MLlib.
Participants will learn how to perform entity ranking, implement recommendation engines and perform other common data science tasks using Spark batch, streaming, graph and machine learning capabilities.
BDT62 / 3 Days
In this course, participants will:
Yes (Digital format)