A customer wants to return a product. The chatbot asks when they bought it. The customer replies:
29What does that number mean?
If the previous question was “How many days ago did you buy it?”, it probably means 29 days. If the previous question was “What is your order number?”, it means something completely different.
That distinction became the most useful lesson in this project. We started with a small support chatbot, added evaluations, connected a real OpenAI model, and watched an apparently simple conversation break in several different ways.
The final saved live run passed all 16 cases. Getting there required more than adding instructions to a prompt: we had to inspect requests, preserve history, distinguish identifiers from customer facts, and decide which behaviors belonged in ordinary application code.
I am separating that walkthrough into six practical chapters. Start by making a wrong answer observable, then repair the information path, connect a live backend, inspect evidence, and finally assign explicit responsibilities to application state.
What we built
The chatbot answers return-policy questions using this deliberately small policy:
Returns are accepted within 30 days. Premium members may return items within 60 days.
The expected behavior is straightforward:
| Known facts | Expected action |
|---|---|
| Purchase age is unknown | Ask how many days ago it was purchased |
| Purchase is at most 30 days old | Eligible, regardless of membership |
| Purchase is 31–60 days old; membership unknown | Ask about membership |
| Purchase is 31–60 days old; premium member | Eligible |
| Purchase is 31–60 days old; standard member | Ineligible |
| Purchase is older than 60 days | Ineligible |
| No relevant policy was retrieved | Return no_context |
We implemented matching TypeScript and Python versions. TypeScript was the main learning path; Python provided a useful comparison of the same request flow.
Start with a deliberately weak design
The first version accepted a question and conversation history, retrieved a policy, and built model messages.
Historical design, simplified:
1function respond(question: string, history: Message[] = []) {2 const context = retrieve(question);3 4 const messages: Message[] = [5 {6 role: "system",7 content:8 "Answer customer questions using the provided policy. " +9 "Be concise and give a direct answer.",10 },11 {12 role: "user",13 content: `Policy:\n${context}\n\nQuestion:\n${question}`,14 },15 ];16 17 return generate(messages);18}Three details matter:
historyis accepted but never used.- Retrieval receives only the current message.
- The instruction asks for a direct answer but does not explain when clarification is necessary.
The role labels themselves are not the bug. A system message for instructions and a user message for a question are reasonable. The problem is the information and behavior surrounding those messages.
For a 45-day-old purchase with unknown membership, the chatbot denied the return. It had the premium-member exception in its context but did not ask the missing question.
Learning: An incorrect response is an observation. It does not, by itself, tell us whether retrieval, prompting, state, or grading failed.
Make the expected behavior executable
We put examples in cases.json:
1{2 "id": "policy-unknown-45",3 "category": "missing_information",4 "question": "I bought this 45 days ago. Can I return it?",5 "expected": "ask_membership"6}The answer contract separated a machine-readable decision from customer-facing text:
1type Decision =2 | "eligible"3 | "ineligible"4 | "ask_days"5 | "ask_membership"6 | "no_context";7 8type Answer = {9 decision: Decision;10 text: string;11};The central grading rule was intentionally simple:
const passed = actual.decision === testCase.expected;This checks the selected action without requiring exact sentence matching. It also creates a blind spot: a correct decision with a misleading explanation still passes. We will return to that limitation.
The initial simulator scored 5/8. Adding clarification guidance brought the TypeScript simulator to 6/8, leaving conversation-history failures.
Use a simulator to learn the plumbing—but know its limits
Before making live API calls, we used a deterministic model simulator in model.ts. It extracted recognizable day and membership phrases and returned predefined responses.
It even recognized clarification guidance through a programmed keyword check:
1const shouldClarify = ["ask", "missing", "member"].every(2 (word) => systemInstructions.toLowerCase().includes(word),3);That is a teaching mechanism, not an approximation of how a real language model reads instructions.
We refactored the simulator into small functions for reading messages, extracting facts, and deciding eligibility. This made it easier to trace data through the application without mixing that exercise with network calls or variable model behavior.
Learning: A simulator can verify wiring and exercise a workflow. Its passing score does not predict real-model performance.
Exercise: explain one failed decision
Use the existing starting branch in a local checkout. Run the offline baseline, then inspect policy-unknown-45. Write down the one missing fact that changes the answer before changing the clarification instruction. Keep the expected decision unchanged. Save your work locally before switching teaching branches.
1node typescript/evaluate.ts --offline2node typescript/evaluate.ts --offline --case=policy-unknown-45Read the trace for the targeted case. The first design accepts history as a parameter but does not put it into messages. That is a concrete application defect you can demonstrate without calling a model. A useful experiment changes one instruction, retains the same cases, and records the failing case IDs rather than chasing a greener total.
Do not edit the grader merely to match the bot. The policy gives standard members 30 days and premium members 60; at 45 days, membership changes the result. At 20 days, it does not. A clarification should request a fact that can affect the decision, not collect every possible fact.
The simulator’s keyword convention means an instruction containing ask, missing, and member changes programmed behavior. That makes this checkpoint excellent for learning the harness, but weak evidence about language understanding. The next chapter follows the natural reply that those first eight cases missed.
Lab source and official references