Conversation bubbles flow through explicit memory to evaluated replies

Practice · Six-chapter chatbot lab

Building an OpenAI Chatbot

Start with a failing support chatbot. Repair its information path, connect the Responses API, inspect evals, and make conversation state explicit.

6 published chapters · TypeScript & Python · Local exercises

From theory to working code

Your chatbot reading path

40 min total
  1. Chapter 1: Start with a Failing Chatbot
    Part 016 min read

    Start with a Failing Chatbot

    Build a failing OpenAI chatbot lab with an executable return-policy contract, offline eval cases, and clear simulator limits.

    Read chapter
  2. Chapter 2: Fix Retrieval and Conversation History
    Part 025 min read

    Fix Retrieval and Conversation History

    Fix chatbot retrieval and conversation history for short numeric replies, then test 30/61-day boundaries without overstating offline evals.

    Read chapter
  3. Chapter 3: Connect the Responses API
    Part 037 min read

    Connect the Responses API

    Connect a local TypeScript or Python chatbot to the OpenAI Responses API with structured output, error handling, and explicit live/offline modes.

    Read chapter
  4. Chapter 4: Debug with Evals and Local Traces
    Part 046 min read

    Debug with Evals and Local Traces

    Debug chatbot evals with local trace provenance, report scope, case-level comparisons, and separate diagnoses for ambiguous numeric replies.

    Read chapter
  5. Chapter 5: Make Conversation State Explicit
    Part 059 min read

    Make Conversation State Explicit

    Make chatbot conversation state explicit: preserve order IDs as strings, enforce application guards, and inspect the historical before/after trace.

    Read chapter
  6. Chapter 6: Read Results Honestly and Extend the Lab
    Part 067 min read

    Read Results Honestly and Extend the Lab

    Interpret chatbot eval results honestly, separate model calls from application guards, and extend the lab with held-out tests and quality checks.

    Read chapter

One lab, three teaching checkpoints

Use the README for Node 22.18+ or Python 3.11+ setup. The codex branches reconstruct the starting, history, and state checkpoints; reported scores and traces are historical evidence, not newly rerun results.

pfplabs/chatbot-eval-lab

An independent PFP Labs teaching series. OpenAI identifies the technology; no vendor affiliation or endorsement is implied.