- Introduction to Data Processing 2
- Horizontal & vertical scalability
- Batch processing vs. stream processing
- Typical data pipeline architecture
- Core ethical principles
- Copyright
- Algorithmic bias
Unit 2: Apache Spark Basics
- Resilient Distributed Datasets (RDDs)
- Datasets and DataFrames
- Transformations and actions
- Directed Acyclic Graphs (DAGs)
- Spark Structured Query Language (SQL)
Unit 3: Apache Spark MLlib
- Spark’s machine learning Library (MLlib)
- MLlib classification and regression algorithms
- Featurization: extracting, transforming, and selecting features
- ML pipelines
Unit 4: Apache Kafka and Spark Streaming
- Apache Kafka topics and partitions
- Apache Kafka producers and consumers
- Apache Spark structured streaming
- Combining Kafka's reliable messaging with Apache Sparks structured streaming processing
Unit 5: Relevant European Regulations
- The European Strategy for Data
- The General Data Protection Regulation
- Copyright in the Digital Single Market
- The Directive on Open Data
- The Data Act & the Data Governance Acts
- The Artificial Intelligence Act
Unit 6: Future Outlook
- Multimodal analytics
- Decentralised architectures
- Big data and generative AI
- Future trends
Unit 7: Examination Prep
Unit 8: Final Exam