Course Overview
This course will introduce Apache Spark. The students will learn how Spark fits into the Big Data ecosystem, and how to use Spark for data analysis.
This class is taught with Python language and using Jupyter environment.
What you’ll learn
During the Data Analytics With Hadoop And Spark course, students will learn:
- Spark ecosystem
- Spark Shell
- Spark Data structures (RDD / Dataframe / Dataset)
- Spark SQL
- Modern data formats and Spark
- Spark & Hadoop & Hive
Requirements
- Analyst background (familiarity with SQL, Scripting ..etc)
Target audiences
- Data Analysts , Business Analysts
Curriculum
- 1 Section
- 8 Lessons
- 3 Days
Expand all sectionsCollapse all sections
- Topics8
- 1.1Spark Introduction Big Data, Hadoop, Spark Spark concepts and architecture Spark components overview Labs : Installing and running Spark
- 1.2First Look at Spark Spark shell Spark web UIs Analyzing dataset – part 1 Labs: Spark shell exploration
- 1.3Spark Data structures Partitions Distributed execution Operations: transformations and actions Labs: Unstructured data analytics using RDDs
- 1.4Caching Caching overview Various caching mechanisms available in Spark In memory file systems Caching use cases and best practices Labs: Benchmark of caching performance
- 1.5Dataframes / Datasets Dataframes Intro Loading structured data (json, CSV) using Dataframes Using schema Specifying schema for Dataframes Labs : Dataframes, Datasets, Schema
- 1.6Spark SQL Spark SQL concepts and overview Defining tables and importing datasets Querying data using SQL Handling various storage formats : JSON / Parquet / ORC Labs: querying structured data using SQL; evaluating data formats
- 1.7Spark and Hadoop Hadoop Primer: HDFS / YARN Hadoop + Spark architecture Running Spark on Hadoop YARN Processing HDFS files using Spark Spark & Hive
- 1.8Workshops These are group workshops Attendees will work on solving real-world data analysis problems using Spark

