Big Data Analysis with MongoDB & PySpark — Tutorial Series
Hands-on big data analysis — MongoDB & NoSQL from fundamentals to sharding, then distributed computing and PySpark on the same dataset.
Instructor: Istiaq Ahmed Fahad
Term: August 2026
Location: Institute of Information Technology (IIT) & IASDS, University of Dhaka
Course Overview
This lecture series bridges relational (SQL) and document (MongoDB/NoSQL) thinking, showing how evolving, nested application data maps from tables and relationships to flexible documents. It covers document models, terminology mapping, and when MongoDB is the right fit for real-world use cases, then scales the same analysis up to distributed computing with MapReduce and PySpark.
Venues
The series was delivered to students of two institutes, University of Dhaka:
Tutorial Tracks
The hands-on material is published as two tracks:
- Big Data Analysis with MongoDB — Tutorials 1–4: environment setup, MongoDB fundamentals, advanced real-world analysis, and sharding.
- Big Data Analysis with PySpark — Tutorials 5–6: distributed computing concepts and a hands-on PySpark implementation.
Prerequisites
- Basic SQL and data modeling; no prior MongoDB required.
- For the PySpark track: basic Python familiarity (loops, functions) and roughly 1 GB of free disk for Java and PySpark. No MongoDB or Docker needed — the PySpark track is independent of the MongoDB series.
Schedule
| Topic | Materials |
|---|---|
| Lecture 1 — MongoDB Fundamentals & NoSQL Overview SQL foundations, where relational modeling becomes difficult, same data as tables vs documents, what is NoSQL, why MongoDB. | |
| Lecture 2 — Data Modeling & Queries SQL vs MongoDB terminology map, embedding vs referencing, when MongoDB fits (e-commerce, IoT, CMS, user profiles). | |
| Lecture 3 — Advanced MongoDB & Big Data Analysis Document model (Database → Collections → Documents), indexing, sharding, horizontal scaling and scalable analytics. | |
| Lecture 4 — Sharding & Distributed MongoDB Building a real sharded cluster in Docker — config servers, shards and a mongos router, shard key selection and rebalancing on the Olist orders collection. | |
| Lecture 5 — MapReduce & Distributed Computing Concepts What MapReduce solves and why, the map/shuffle/reduce model, how Hadoop implements it, why Spark replaced raw Hadoop, and the core PySpark vocabulary. | |
| Lecture 6 — PySpark Implementation Installing Java and PySpark, then building four programs on the Olist dataset — Word Count, a key–value RDD job, a DataFrame join and a trained MLlib regression pipeline. |