Learning series

PFP Labs Learning Series

Independent PFP Labs reading paths. Track names describe the topic, not an affiliation.

Featured artwork from PySpark Mental Model: Lazy DataFrames for Media Analytics
Databricks / Spark

Spark & Databricks

An eight-part PySpark reading path on synthetic streaming-media data, from DataFrames to a medallion pipeline.

Open reading path
Series cover: Spark for Streaming Media Analytics

8 published chapters

Spark for Streaming Media Analytics

Eight PySpark chapters on synthetic playback data, with a runnable companion lab.

Open reading path
  1. Chapter 1: PySpark Mental Model: Lazy DataFrames for Media Analytics — cover artwork01PySpark Mental Model: Lazy DataFrames for Media Analytics9 min read
  2. Chapter 2: PySpark Ingestion: Explicit Schemas and a Delta Bronze Table — cover artwork02PySpark Ingestion: Explicit Schemas and a Delta Bronze Table9 min read
  3. Chapter 3: PySpark Window Functions: Viewer Sessions and Completion Rates — cover artwork03PySpark Window Functions: Viewer Sessions and Completion Rates10 min read
  4. Chapter 4: Spark Performance: Shuffles, Partitions, and Broadcast Joins — cover artwork04Spark Performance: Shuffles, Partitions, and Broadcast Joins9 min read
  5. Chapter 5: Spark Structured Streaming: Watermarks and Concurrent Viewers — cover artwork05Spark Structured Streaming: Watermarks and Concurrent Viewers10 min read
  6. Chapter 6: Testing PySpark Pipelines: Data Quality Checks That Catch Bugs — cover artwork06Testing PySpark Pipelines: Data Quality Checks That Catch Bugs9 min read