The traces separated two problems: empty retrieval and an identifier mistaken for a customer fact. More prompt text was not a guarantee. I wanted the application to preserve what it knew, and to enforce the small invariants it could already decide.
Move the empty-context decision into code
The application knows whether retrieval returned anything. It does not need a model call to decide that.
The final orchestration includes:
1if (!context) {2 return finish(3 {4 decision: "no_context",5 text:6 "I couldn't find a relevant policy. " +7 "Could you describe what you need help with?",8 },9 false,10 );11}finish() packages the answer, next state, and trace. Its second argument records whether the backend was called.
This fixed fresh-number-no-context deterministically, without an API request. It did not fix the order-number case because that conversation did have a retrieved policy.
Learning: Implement an invariant directly when the application already has the information needed to enforce it.
Make conversation state explicit
Conversation history and structured state serve different purposes:
1// History preserves what was said.2{ role: "assistant", content: "What is your order number?" }3 4// State records what answer the application is waiting for.5{ pendingQuestion: "order_number" }We introduced:
1type PendingQuestion =2 | "purchase_age"3 | "membership"4 | "order_number"5 | null;6 7type ConversationState = {8 pendingQuestion: PendingQuestion;9 daysSincePurchase: number | null;10 isPremiumMember: boolean | null;11 orderNumber: string | null;12};Now a numeric reply has a destination:
1if (state.pendingQuestion === "order_number" && /^\d+$/.test(reply)) {2 state.orderNumber = reply;3 state.pendingQuestion = null;4 return state;5}We deliberately keep order numbers as strings: "007123" should not become 7123.
When the pending question is purchase age, a numeric reply instead updates daysSincePurchase. A yes/no reply to a membership question updates isPremiumMember.
The state helper returns a copy. If the API call fails later in the turn, the chat loop does not commit a partially updated conversation.
State alone is not the entire fix
Merely passing a state object to a model would still leave it free to misinterpret the response. For the specific ambiguity we understood, we added an application branch:
1if (2 previousState.pendingQuestion === "order_number" &&3 /^\d+$/.test(question.trim()) &&4 state.daysSincePurchase === null5) {6 return finish(7 {8 decision: "ask_days",9 text: "How many days ago did you buy it?",10 },11 false,12 );13}After receiving order number "29", the returned state is:
1{2 "pendingQuestion": "purchase_age",3 "daysSincePurchase": null,4 "isPremiumMember": null,5 "orderNumber": "29"6}The app knows the order number and still does not know the purchase age. It asks for the missing fact without spending another model call on this ambiguity.
This is intentionally narrow behavior. It does not make the regex parser a general natural-language understanding system.
Carry state through the complete chat loop
The original walkthrough’s key requirement was to carry answer.state into the next call and clear both history and state on reset. The current public chat.ts below shows the complete loop, including failed-turn handling; it is not a hosted service or a new implementation.
1import { createInterface } from "node:readline/promises";2import { stdin, stdout } from "node:process";3import { respond } from "./chatbot.ts";4import type { Message } from "./types.ts";5import { backend, model } from "./backend.ts";6import { createState } from "./state.ts";7 8const history: Message[] = [];9let state = createState();10const terminal = createInterface({ input: stdin, output: stdout });11console.log(`Practice chatbot (${backend}: ${backend === "offline" ? "simulator" : model}). Type /quit to exit or /reset to clear history.`);12try {13 while (true) {14 let question: string;15 try {16 question = (await terminal.question("You: ")).trim();17 } catch (error) {18 if ((error as NodeJS.ErrnoException).code === "ERR_USE_AFTER_CLOSE") break;19 throw error;20 }21 if (question === "/quit") break;22 if (question === "/reset") {23 history.length = 0;24 state = createState();25 continue;26 }27 if (!question) continue;28 let answer;29 try {30 answer = await respond(question, history, state);31 } catch (error) {32 console.error(`Error: ${(error as Error).message}`);33 continue;34 }35 console.log(`Bot: ${answer.text} [${answer.decision}]`);36 state = answer.state;37 history.push(38 { role: "user", content: question },39 { role: "assistant", content: answer.text },40 );41 }42} finally {43 terminal.close();44}The chatbot maps ask_days to a pending purchase-age question and ask_membership to a pending membership question. Other decisions clear the pending question.
The Python version follows the same pattern:
1state = create_state()2answer = respond(question, history, state)3state = answer["state"]Start mid-conversation tests with explicit state
The order-number fixture now includes:
1{2 "id": "number-without-days-question",3 "category": "missing_information",4 "history": [5 { "role": "user", "content": "I want to return a product" },6 { "role": "assistant", "content": "What is your order number?" }7 ],8 "initialState": {9 "pendingQuestion": "order_number",10 "daysSincePurchase": null,11 "isPremiumMember": null,12 "orderNumber": null13 },14 "question": "29",15 "expected": "ask_days"16}That starting state is supplied by the test; the harness does not infer it from assistant prose.
This distinction matters. The chatbot does not currently initiate an order-number question itself. This fixture tests an explicitly seeded application state, not an end-to-end order lookup workflow.
Other mid-conversation cases similarly supply initial facts. We also checked sequential follow-ups outside the fixture harness so state was actually carried from one response into the next.
The aha moment: the same “29,” a different responsibility
This was the conversation:
1Customer: I want to return a product2Assistant: What is your order number?3Customer: 29Before the fix, an actual console result from our debugging session was:
1FAIL number-without-days-question: expected=ask_days actual=eligible (714ms)2 Since your purchase was made 29 days ago, you are eligible to return the product.The customer never said “29 days.” The chatbot converted an identifier into a fact about purchase age. A fluent answer hid an unsupported assumption.
Here is a shortened excerpt from the final saved eval report. Fields have been omitted for readability; the values below are unchanged:
1{2 "id": "number-without-days-question",3 "history": [4 {5 "role": "user",6 "content": "I want to return a product"7 },8 {9 "role": "assistant",10 "content": "What is your order number?"11 }12 ],13 "question": "29",14 "expected": "ask_days",15 "actual": {16 "decision": "ask_days",17 "text": "How many days ago did you buy it?",18 "trace": {19 "modelCalled": false,20 "stateBefore": {21 "pendingQuestion": "order_number",22 "daysSincePurchase": null,23 "isPremiumMember": null,24 "orderNumber": null25 },26 "stateAfterReply": {27 "pendingQuestion": null,28 "daysSincePurchase": null,29 "isPremiumMember": null,30 "orderNumber": "29"31 }32 },33 "state": {34 "pendingQuestion": "purchase_age",35 "daysSincePurchase": null,36 "isPremiumMember": null,37 "orderNumber": "29"38 }39 },40 "passed": true41}Read the trace in three steps:
- Before the reply:
pendingQuestionisorder_number. The application is waiting for an identifier. - After interpreting “29”:
orderNumberbecomes"29", whiledaysSincePurchasestaysnull. We have learned an order number—not a purchase age. - After choosing the response: the app asks “How many days ago did you buy it?” and sets the next pending question to
purchase_age.
The revealing field is:
"modelCalled": falseThis particular PASS verifies an application branch. It is not evidence that the model learned to interpret order numbers reliably. The original 714 ms failure and this skipped call also do not establish a general latency benchmark.
Two details keep the example honest: the test supplies the starting order_number state explicitly, and the trace is recorded by our application, not automatically generated by the Responses API. The current chatbot does not initiate an order-number collection workflow itself.
That is the practical distinction I want to remember: inspect whether an answer is grounded in facts the conversation supplied, then decide whether the missing guarantee belongs in model guidance or ordinary application logic.
Exercise: preserve an ID, then learn an age
On main, run the offline regression suite and inspect number-without-days-question. In a local test, seed pendingQuestion: order_number explicitly, reply with 007123, then pass the returned state into a purchase-age reply and a membership reply. Check the state at every boundary. Finally reset the interactive session and verify both history and facts are empty.
1node typescript/evaluate.ts --offline --case=number-without-days-question2node typescript/chat.ts --offlineThe important boundary is transactional, not just typed: applyReply returns a copy, respond packages a next state, and the chat loop commits that state and history only after a successful response. If a call throws, the loop continues without adding half a turn. Clearing history without state would leave facts from a supposedly fresh session; clearing state without history would leave old prose available to retrieval.
The final chapter asks what these successful branches prove, what still uses the model, and which tests would make the next iteration more informative.