Big Data Analysis with PySpark

Lecture Series

Big Data Analysis with PySpark

Two lectures on distributed computing: a code-free introduction to MapReduce, Hadoop & Spark, followed by a fully hands-on implementation lecture that builds four real PySpark programs on the same Olist e-commerce data from Tutorial 3.

Concepts — No Install Required

MapReduce & Distributed Computing Concepts

A gentle, complete introduction: what MapReduce solves and why, the map/shuffle/reduce model, how Hadoop implements it, why Spark replaced raw Hadoop, and the core PySpark vocabulary — all in plain language, nothing to install.

45-60 min Beginner 4 steps
Start Lecture 1
Hands-On Implementation

PySpark Implementation: Step by Step

Install Java & PySpark, then build four real programs on the Olist dataset: Word Count, a key–value RDD job, a DataFrame join, and a trained MLlib regression pipeline.

100-130 min Advanced 6 steps, 5 exercises
Start Lecture 2

Prerequisites

Dataset Used