Syllabus

Title
0708 Data Processing 2: Scalable Data Processing, Legal & Ethical Foundations of Data Science
Instructors
Assoz.Prof PD Dr. Sabrina Kirrane
Contact details
Type
PI
Weekly hours
2
Language of instruction
Englisch
Registration
09/03/26 to 11/25/26
Registration via LPIS
Notes to the course
Dates
Day Date Time Room
Tuesday 12/01/26 12:30 PM - 04:30 PM D5.0.001
Wednesday 12/09/26 09:00 AM - 01:00 PM LC.-1.022 (P&S)
Tuesday 12/15/26 09:00 AM - 01:00 PM LC.-1.022 (P&S)
Tuesday 12/22/26 09:00 AM - 01:00 PM LC.-1.022 (P&S)
Tuesday 01/12/27 12:30 PM - 04:30 PM TC.1.02
Tuesday 01/19/27 12:00 PM - 04:00 PM TC.3.03
Wednesday 01/20/27 10:00 AM - 01:00 PM LC.-1.022 (P&S)
Wednesday 01/27/27 10:00 AM - 01:00 PM LC.-1.022 (P&S)
Contents

This fast-paced class is intended for students interested in scalable handling of big data, understanding legal fundamentals and ethical frameworks in dealing with data in an international context. The course focuses on gaining fundamental knowledge in dealing with large amounts of data and learning about efficient and scalable processing methods. Throughout the course there will be an emphasis on important aspects regarding legal and ethical principals related to data processing and data science.

Learning outcomes

Students in the course will learn about the scalable handling of big data, understanding legal fundamentals and ethical frameworks in dealing with data in an international context.

This includes:

  • Basic knowledge about different scalable data processing frameworks and paradigms, including:
    • Batch processing vs. stream processing
    • Data pipeline architecture
    • Batch processing with Apache Spark
    • Stream processing with Apache Kafka
  • Legal and Ethical frameworks
    • Codes of Conduct
    • Intellectual Property Rights / handling of different licensing schemes
    • Algorithmic bias
    • Relevant European Regulations
Attendance requirements
  • Attendance at the course is compulsory and will be monitored (attendance list). The minimum attendance for an assessment is 80% of the units. Any absence must be reported in a timely manner and a reason for absence must be provided.
  • Attendance at the first and main examination dates is mandatory for students and is exempt from the above-mentioned exception to compulsory attendance.
  • An unjustified and excused absence in the first unit may result in loss of place.
  • The possibility of taking a substitute exam at a later date only exists if the main exam was missed for a valid reason (e.g. illness with a medical certificate).
Teaching/learning method(s)
 Unit 1: Introduction
  • Introduction to Data Processing 2
  • Horizontal & vertical scalability
  • Batch processing vs. stream processing
  • Typical data pipeline architecture
  • Core ethical principles
  • Copyright
  • Algorithmic bias

Unit 2: Apache Spark Basics

  • Resilient Distributed Datasets (RDDs)
  • Datasets and DataFrames
  • Transformations and actions
  • Directed Acyclic Graphs (DAGs)
  • Spark Structured Query Language (SQL)

Unit 3: Apache Spark MLlib

  • Spark’s machine learning Library (MLlib)
  • MLlib classification and regression algorithms 
  • Featurization: extracting, transforming, and selecting features
  • ML pipelines 

Unit 4: Apache Kafka and Spark Streaming

  • Apache Kafka topics and partitions
  • Apache Kafka producers and consumers
  • Apache Spark structured streaming 
  • Combining Kafka's reliable messaging with Apache Sparks structured streaming processing

Unit 5: Relevant European Regulations

  • The European Strategy for Data
  • The General Data Protection Regulation
  • Copyright in the Digital Single Market
  • The Directive on Open Data
  • The Data Act & the Data Governance Acts
  • The Artificial Intelligence Act

Unit 6: Future Outlook

  • Multimodal analytics
  • Decentralised architectures
  • Big data and generative AI
  • Future trends

Unit 7: Examination Prep

Unit 8: Final Exam

Assessment

In-class Exercises & Quizzes: 40%

Take-home Project: 30% 

In-class Examination: 30% 

 

Grading Scheme:

90−100 Sehr gut (Really good) is the best possible grade and indicates outstanding performance with no or only minor errors.

80−89 Gut (Good) is the next-highest grade and is given for performance that is above-average standard but with some errors.

64−79 Befriedigend (Satisfactory) indicates generally sound work with a number of notable errors.

51−63 Genügend (Sufficient) is the lowest passing grade and is given if the standard has been met but with a significant number of shortcomings.

0−50 Nicht genügend (Insufficient) is the lowest possible grade and the only failing grade.

Prerequisites for participation and waiting lists

Students need to register for course 1 of SBWL Data Science before registering for this course.

Please be aware that for all courses in this SBWL registration is only possibly for students who successfully have completed the entry course (Einstieg in die SBWL: Data Science).

Note that for courses within the SBWL "Data Science" we can only accept students enrolled in one of WU's bachelor programmes who qualify for starting an SBWL; particularly, we cannot accept students from other courses and programmes enrolled at WU as 'Mitbeleger' only.

Readings

Please log in with your WU account to use all functionalities of read!t. For off-campus access to our licensed electronic resources, remember to activate your VPN connection connection. In case you encounter any technical problems or have questions regarding read!t, please feel free to contact the library at readinglists@wu.ac.at.

Availability of lecturer(s)

During the lecture and based on individual appointments. To request an appointment send an email to the lecturers with the subject “[Data Processing 2]”.

Last edited: 2026-08-13



Back