AI assistance is most useful once you can tell good Spark code from plausible Spark code. That is why this chapter comes after plans, windows, streaming, and tests: those skills are exactly what you use to review what an assistant proposes.
Genie Code is Databricks’ AI coding and data assistant in the workspace. According to the documentation, it can generate, explain, optimize, and fix code in notebooks, the SQL editor, pipelines editor, and dashboards, and it is governed by your Unity Catalog permissions. Availability, pricing, and features depend on your workspace and account settings, so check the current docs.
Where it fits in the workflow
- Drafting a first version of a transformation you can already describe precisely.
- Explaining an unfamiliar plan or error message.
- Fixing errors with Diagnose Error and Quick Fix, which you still accept or reject.
- Translating between PySpark and SQL so you can compare plans.
Databricks itself recommends reviewing generated code before running it. Treat every suggestion as a pull request from a fast, confident colleague who has not read your metric definitions.
Example prompts to try
These are example prompts for the synthetic tables in this series, not output from a specific session. Reference tables with @ so the assistant uses real column names.
1Using @bronze_playback, write PySpark that assigns a session_n per2viewer_id, starting a new session after more than 30 minutes of3inactivity ordered by event_ts. Use window functions, not a UDF.4Return the code only, and explain how ties in event_ts are handled.1Explain this physical plan line by line. Identify each Exchange,2why it exists, and whether a broadcast join could remove it.3Do not change the code yet.A review checklist
| Check | How |
|---|---|
| Metric definition | Does the code match the written definition, including boundaries like > versus >=? |
| Nulls and duplicates | Run the transformation on the tiny test fixtures from part 6 |
| Plan shape | explain(): unexpected shuffles, cartesian joins, Python UDFs |
| Scope of change | Did it modify only the cell you asked about? |
| Data access | Does it read only the tables the task needs? |
Exercises
- Ask for a sessionization function, then run it against the part 6 tests. Record what failed, if anything.
- Ask the assistant to optimize the genre metrics query, and compare plans before and after.
- Write your own team prompt template that states metric definitions up front.
- Note one suggestion you rejected and why.
Official documentation
