Big Data Analysis with MongoDB & PySpark — Tutorial Series

Hands-on big data analysis — MongoDB & NoSQL from fundamentals to sharding, then distributed computing and PySpark on the same dataset.

Instructor: Istiaq Ahmed Fahad

Term: August 2026

Location: Institute of Information Technology (IIT) & IASDS, University of Dhaka

Course Overview

This lecture series bridges relational (SQL) and document (MongoDB/NoSQL) thinking, showing how evolving, nested application data maps from tables and relationships to flexible documents. It covers document models, terminology mapping, and when MongoDB is the right fit for real-world use cases, then scales the same analysis up to distributed computing with MapReduce and PySpark.

Venues

The series was delivered to students of two institutes, University of Dhaka:

Tutorial Tracks

The hands-on material is published as two tracks:

Prerequisites

  • Basic SQL and data modeling; no prior MongoDB required.
  • For the PySpark track: basic Python familiarity (loops, functions) and roughly 1 GB of free disk for Java and PySpark. No MongoDB or Docker needed — the PySpark track is independent of the MongoDB series.

Schedule

Topic Materials
Lecture 1 — MongoDB Fundamentals & NoSQL Overview

SQL foundations, where relational modeling becomes difficult, same data as tables vs documents, what is NoSQL, why MongoDB.

Lecture 2 — Data Modeling & Queries

SQL vs MongoDB terminology map, embedding vs referencing, when MongoDB fits (e-commerce, IoT, CMS, user profiles).

Lecture 3 — Advanced MongoDB & Big Data Analysis

Document model (Database → Collections → Documents), indexing, sharding, horizontal scaling and scalable analytics.

Lecture 4 — Sharding & Distributed MongoDB

Building a real sharded cluster in Docker — config servers, shards and a mongos router, shard key selection and rebalancing on the Olist orders collection.

Lecture 5 — MapReduce & Distributed Computing Concepts

What MapReduce solves and why, the map/shuffle/reduce model, how Hadoop implements it, why Spark replaced raw Hadoop, and the core PySpark vocabulary.

Lecture 6 — PySpark Implementation

Installing Java and PySpark, then building four programs on the Olist dataset — Word Count, a key–value RDD job, a DataFrame join and a trained MLlib regression pipeline.