Big Data Analysis — Tutorial Series

Big Data Analysis: Choose Your Track

Two hands-on, self-paced tutorial series covering how modern systems store and process data at scale — one on document databases, one on distributed computation. Pick one to get started, or come back later for the other.

Document Databases

Big Data Analysis with MongoDB

Five tutorials from zero to a real sharded cluster: environment setup, CRUD & aggregation pipelines, nested documents & $lookup joins on 100K+ real orders, and horizontal scaling with sharding.

5 tutorials 5-7 hours total Beginner → Ultra Advanced
Start This Track
Distributed Computation

Big Data Analysis with PySpark

Two lectures: a gentle, code-free intro to MapReduce, Hadoop & Spark concepts, then a hands-on implementation lecture — Word Count, key-value RDDs, DataFrame joins, and a real machine learning pipeline.

2 lectures, 10 steps 145-190 min Advanced
Start This Track

Prerequisites (Both Tracks)