Back to blog

Building an OpenAI Chatbot · Part 1 of 6

Start with a Failing Chatbot

Building an OpenAI Chatbot · Part 1 of 6

Chris Eberl

Chris Eberl

Founder • Engineering Leader, GenAI, Data, ML

AI EngineeringOct 8, 20266 min read
Start with a Failing Chatbot

A customer wants to return a product. The chatbot asks when they bought it. The customer replies:

Text
29

What does that number mean?

If the previous question was “How many days ago did you buy it?”, it probably means 29 days. If the previous question was “What is your order number?”, it means something completely different.

That distinction became the most useful lesson in this project. We started with a small support chatbot, added evaluations, connected a real OpenAI model, and watched an apparently simple conversation break in several different ways.

The final saved live run passed all 16 cases. Getting there required more than adding instructions to a prompt: we had to inspect requests, preserve history, distinguish identifiers from customer facts, and decide which behaviors belonged in ordinary application code.

I am separating that walkthrough into six practical chapters. Start by making a wrong answer observable, then repair the information path, connect a live backend, inspect evidence, and finally assign explicit responsibilities to application state.

What we built

Return-policy decision contract: no context, unknown age, 30-day eligibility, membership at 31–60 days, and ineligibility after 60 days.
Clarification depends on the missing fact that can change eligibility.

The chatbot answers return-policy questions using this deliberately small policy:

Returns are accepted within 30 days. Premium members may return items within 60 days.

The expected behavior is straightforward:

Known factsExpected action
Purchase age is unknownAsk how many days ago it was purchased
Purchase is at most 30 days oldEligible, regardless of membership
Purchase is 31–60 days old; membership unknownAsk about membership
Purchase is 31–60 days old; premium memberEligible
Purchase is 31–60 days old; standard memberIneligible
Purchase is older than 60 daysIneligible
No relevant policy was retrievedReturn no_context

We implemented matching TypeScript and Python versions. TypeScript was the main learning path; Python provided a useful comparison of the same request flow.

Start with a deliberately weak design

The first version accepted a question and conversation history, retrieved a policy, and built model messages.

Historical design, simplified:

Code
1function respond(question: string, history: Message[] = []) {2  const context = retrieve(question);3 4  const messages: Message[] = [5    {6      role: "system",7      content:8        "Answer customer questions using the provided policy. " +9        "Be concise and give a direct answer.",10    },11    {12      role: "user",13      content: `Policy:\n${context}\n\nQuestion:\n${question}`,14    },15  ];16 17  return generate(messages);18}

Three details matter:

  • history is accepted but never used.
  • Retrieval receives only the current message.
  • The instruction asks for a direct answer but does not explain when clarification is necessary.

The role labels themselves are not the bug. A system message for instructions and a user message for a question are reasonable. The problem is the information and behavior surrounding those messages.

For a 45-day-old purchase with unknown membership, the chatbot denied the return. It had the premium-member exception in its context but did not ask the missing question.

Learning: An incorrect response is an observation. It does not, by itself, tell us whether retrieval, prompting, state, or grading failed.

Make the expected behavior executable

We put examples in cases.json:

JSON
1{2  "id": "policy-unknown-45",3  "category": "missing_information",4  "question": "I bought this 45 days ago. Can I return it?",5  "expected": "ask_membership"6}

The answer contract separated a machine-readable decision from customer-facing text:

Code
1type Decision =2  | "eligible"3  | "ineligible"4  | "ask_days"5  | "ask_membership"6  | "no_context";7 8type Answer = {9  decision: Decision;10  text: string;11};

The central grading rule was intentionally simple:

Code
const passed = actual.decision === testCase.expected;

This checks the selected action without requiring exact sentence matching. It also creates a blind spot: a correct decision with a misleading explanation still passes. We will return to that limitation.

The initial simulator scored 5/8. Adding clarification guidance brought the TypeScript simulator to 6/8, leaving conversation-history failures.

Use a simulator to learn the plumbing—but know its limits

Before making live API calls, we used a deterministic model simulator in model.ts. It extracted recognizable day and membership phrases and returned predefined responses.

It even recognized clarification guidance through a programmed keyword check:

Code
1const shouldClarify = ["ask", "missing", "member"].every(2  (word) => systemInstructions.toLowerCase().includes(word),3);

That is a teaching mechanism, not an approximation of how a real language model reads instructions.

We refactored the simulator into small functions for reading messages, extracting facts, and deciding eligibility. This made it easier to trace data through the application without mixing that exercise with network calls or variable model behavior.

Learning: A simulator can verify wiring and exercise a workflow. Its passing score does not predict real-model performance.

Exercise: explain one failed decision

Use the existing starting branch in a local checkout. Run the offline baseline, then inspect policy-unknown-45. Write down the one missing fact that changes the answer before changing the clarification instruction. Keep the expected decision unchanged. Save your work locally before switching teaching branches.

Shell
1node typescript/evaluate.ts --offline2node typescript/evaluate.ts --offline --case=policy-unknown-45

Read the trace for the targeted case. The first design accepts history as a parameter but does not put it into messages. That is a concrete application defect you can demonstrate without calling a model. A useful experiment changes one instruction, retains the same cases, and records the failing case IDs rather than chasing a greener total.

Do not edit the grader merely to match the bot. The policy gives standard members 30 days and premium members 60; at 45 days, membership changes the result. At 20 days, it does not. A clarification should request a fact that can affect the decision, not collect every possible fact.

The simulator’s keyword convention means an instruction containing ask, missing, and member changes programmed behavior. That makes this checkpoint excellent for learning the harness, but weak evidence about language understanding. The next chapter follows the natural reply that those first eight cases missed.

Read next

Newsletter

New posts, straight from Chris

A short note from me whenever a new article goes live — product engineering, AI workflows, IoT, indie apps, and engineering leadership. No spam, unsubscribe anytime.

By subscribing, you agree to our Privacy Policy. We do not share your email.