Big Data Analysis: Choose Your Track
Two hands-on, self-paced tutorial series covering how modern systems store and process data at scale — one on document databases, one on distributed computation. Pick one to get started, or come back later for the other.
Big Data Analysis with MongoDB
Five tutorials from zero to a real sharded cluster: environment setup, CRUD & aggregation
pipelines, nested documents & $lookup
joins on 100K+ real orders, and horizontal scaling with sharding.
Big Data Analysis with PySpark
Two lectures: a gentle, code-free intro to MapReduce, Hadoop & Spark concepts, then a hands-on implementation lecture — Word Count, key-value RDDs, DataFrame joins, and a real machine learning pipeline.
Prerequisites (Both Tracks)
- A computer running Ubuntu/Debian or Fedora/RHEL Linux
- Internet connection for downloading packages
- Basic familiarity with the terminal / command line
- The two tracks are independent — start with either one, no need to finish both