Course Information
Big Data platforms is a 5 ECTS Master's level advanced course package. It is split into a 3 ECTS open to all students free of charge MOOC course (DATA140031 Big Data Platforms MOOC, 3 ECTS) and an in person controlled exam (140032 Big Data Platforms MOOC EXAM, 2 ECTS) organized through the Open University. University of Helsinki students can (and should) take both courses together to be able to include the course package in their CS Master curriculum and for UH students both course parts are free of charge. See Lecture 1 slides for additional details of the course arrangements.
This course focuses on big data platforms and on key algorithmic ideas and methods used to implement them. After completing this course you are able to list many of the key technologies used in big data processing and to select suitable methods for solving challenging big data processing tasks using cloud computing technologies. You will also be able to compare the scalability and fault tolerance implications of using the selected methodologies.
Main topics are:
- distributed computing,
- Warehouse-Scale Computers,
- fault tolerance in distributed systems,
- distributed file systems,
- distributed batch processing with the MapReduce and the Apache Spark (PySpark) computing frameworks, and
- distributed cloud based databases.
The course material will consist of lecture materials and exercises provided by the lecturer.
Course Target Audience
The course is suitable to those who are interested in big data platforms employed in cloud computing and have previous knowledge in programming, database systems and command line tools. Optional course in Data Science Master's Program. Also suitable for Computer Science Master's Program students. The course is suitable to University of Helsinki exchange students.
Course Prerequisites
To attend this course, you must have:
- basic programming skills (Python),
- skills to work with command line tools in Linux, and
- basic knowledge in database systems (SQL).
Lecture Schedule
The Lectures of the course will be will Zoom based lectures. Slides and video recording of each of the lectures will typically be made available a within 48 hours of the live lecture session. The link to the Zoom lectures is:
https://helsinki.zoom.us/j/65695344424?pwd=TGxVUEFkY2RxQm9FYXBVN2NGWWxmQT09
| Lecture date | Lecture time (EEST) | |
|---|---|---|
| Lecture 1 | Tue 1.9.2026 | 10:15-11:45 |
| Lecture 2 | Thu 3.9.2026 | 12:15-13:45 |
| Lecture 3 | Tue 8.9.2026 | 10:15-11:45 |
| Lecture 4 | Thu 10.9.2026 | 12:15-13:45 |
| Lecture 5 | Tue 15.9.2026 | 10:15-11:45 |
| Lecture 6 | Thu 17.9.2026 | 12:15-13:45 |
| Lecture 7 | Tue 22.9.2026 | 10:15-11:45 |
| Lecture 8 | Thu 24.9.2026 | 12:15-13:45 |
| Lecture 9 | Thu 1.10.2026 | 12:15-13:45 |
| Lecture 10 | Tue 6.10.2026 | 10:15-11:45 |
| Lecture 11 | Tue 13.10.2026 | 10:15-11:45 |
| Backup Lecture Slot | Thu 15.10.2026 | 12:15-12:45 |
Lecture Slides and Videos
The Lecture slides contain all the material needed to pass the course, the videos go through this material and contain no additional information needed for the quizzes.
| Lecture Slides | Lecture Videos | |
|---|---|---|
| Lecture 1 | Lecture 1 slides | Lecture 1 video (YouTube) |
| Lecture 2 | Lecture 2 slides | Lecture 2 video (YouTube) |
| Lecture 3 | Lecture 3 slides | Lecture 3 video (YouTube) |
| Lecture 4 | Lecture 4 slides | Lecture 4 video (YouTube) |
| Lecture 5 | Lecture 5 slides | Lecture 5 video (YouTube) |
| Lecture 6 | Lecture 6 slides | Lecture 6 video (YouTube) |
| Lecture 7 | Lecture 7 slides | Lecture 7 video (YouTube) |
| Lecture 8 | Lecture 8 slides | Lecture 8 video (YouTube) |
| Lecture 9 | Lecture 9 slides | Lecture 9 video (YouTube) |
| Lecture 10 | Lecture 10 slides | |
| Lecture 11 |
Home Exercise Schedule
The course will contain programming exercises where you will be using the Spark framework to solve Big Data processing tasks. We will be using the Python programming language based PySpark interface and will be doing several database query type analytics queries. Therefore basic programming skills using Python and knowledge about database programming, especially using the SQL query language will be very helpful for completing the home exercises.
The schedule for the home exercises was announced in the first Lecture:
| Release Date | Due Date (16:00 EEST) | |
|---|---|---|
| Introduction to Spark + RDD Programming | 8.9.2026 | 22.9.2026 |
| Dataframe Programming | 15.9.2026 | 29.9.2026 |
| Machine Learning (MLlib) | 22.9.2026 | 6.10.2026 |
| Graphframe Programming | 29.9.2026 | 13.10.2026 |
| Structured Streaming | 6.10.2026 | 20.10.2026 |
| Extras (Optional for extra points) | 13.10.2026 | 31.10.2026 |
The Home Exercise System
To complete the home exercises, you will need to utilize the container-based Jupyter Notebook system. Detailed instructions can be found via the following link:
The home exercises will be released according to the schedule outlined above, accessible through the following link:
To submit your assignments, please use the submission box located at the bottom of this page. Please note that the submission box will only become visible once you are logged in.
Course Discord Channel
The course has a Discord channel for helping fellow students. Lecturer and Course Assistant will periodically also join in the conversation. You can join to the groups through the link:
https://study.cs.helsinki.fi/discord/join/bdp
Passing the 3 ECTS MOOC part of the course
You need to pass home exercises by their respective deadlines listed above. Minimum 60% from home exercise totals are needed to pass MOOC part of the course.
Passing the 2 ECTS MOOC EXAM part of the course
You need to register for the DATA140032 course either through Sisu (UH students) or Open University (other students). The Exam is an in-person Exam done in Examinarium during the time period 19.10.2026-9.12.2026.
Use of Large Language Models in the Course
Large language models (LLMs) have recently developed rapidly as versatile tools. While they have useful applications, they can also conflict with learning objectives. Permitted use cases are always dependent on the course.
General language models can produce incorrect, misleading, or irrelevant information. Therefore, it is the student's responsibility to ensure the accuracy and relevance of the information. It is also worth noting that specialized tools generally yield better results than language models.
Presenting generated text or code as one's own can be interpreted as plagiarism. More information can be found at the following link: https://studies.helsinki.fi/instructions/article/what-cheating-and-plagiarism
In this course, the use of language models is completely prohibited.
Submission Box
If you encounter any issue, please report on Discord