The first checkpoint made missing membership visible, but the remaining failures were conversational. I wanted to stop guessing at prompts and follow exactly what reached retrieval and the simulator.
Discover the natural follow-up that the tests missed
An interactive session exposed a new bug:
1You: I want to return a product2Bot: How many days ago did you buy it? [ask_days]3You: 294Bot: I couldn't find a relevant policy. [no_context]This was not one single model failure. It had three causes:
- Retrieval searched only
"29", which did not contain the wordreturn. - Previous messages did not reach the simulator.
- The simulator recognized
"29 days", but not the standalone reply"29".
We first updated retrieval to consider the user conversation:
1const retrievalQuery = [2 ...conversation3 .filter((message) => message.role === "user")4 .map((message) => message.content),5 question,6].join("\n");7 8const context = retrieve(retrievalQuery);We also inserted history before the current question:
1const messages: Message[] = [2 { role: "system", content: instructions },3 ...conversation,4 { role: "user", content: currentRequest },5];The simulator received a narrow follow-up parser for numeric ages and yes/no membership answers. We added cases for those exchanges, including boundaries at 30 and 61 days.
The expanded offline suite passed 16/16 in both languages.
The retrieval implementation was still just a keyword check. Searching all user history fixed the immediate follow-up issue but could also retain an old return topic after the conversation moved on. Fixing one failure does not remove every limitation of the approach.
A narrow parser, not general understanding
The public lab’s model.ts makes the follow-up rule concrete. This excerpt is copied from the current repository; statement is the user-only text after removal of the Policy/Question wrapper. It is an offline teaching rule, not a Responses API feature.
1const previous = messages[index - 1];2 // Interpret short answers only after the corresponding simulator question.3 if (previous?.role === "assistant") {4 if (previous.content === "How many days ago did you buy it?" && /^\d+$/.test(statement)) {5 return [`${statement} days`];6 }7 if (previous.content === "Are you a premium member?") {8 if (/^yes[.!]?$/i.test(statement)) return ["a premium member"];9 if (/^no[.!]?$/i.test(statement)) return ["not a premium member"];10 }11 }Notice the exact assistant sentence comparisons. A bare number becomes a day phrase only after the purchase-age question. Yes and no become membership facts only after the membership question. An identifier after another question should remain an identifier. Rephrasing the assistant question can defeat this exact-match parser; the exercise should expose that boundary rather than hide it.
Expand cases where the policy changes
1[2 {3 "id": "followup-days-30",4 "question": "30",5 "history": [6 {7 "role": "user",8 "content": "I want to return a product"9 },10 {11 "role": "assistant",12 "content": "How many days ago did you buy it?"13 }14 ],15 "expected": "eligible"16 },17 {18 "id": "followup-days-61",19 "question": "61",20 "history": [21 {22 "role": "user",23 "content": "I want to return a product"24 },25 {26 "role": "assistant",27 "content": "How many days ago did you buy it?"28 }29 ],30 "expected": "ineligible"31 }32]These two teaching fixtures show the boundary shape, not a replacement case file. Use the repository’s actual shared cases for scored runs. Thirty days is eligible without membership. Sixty-one is beyond both windows. For 31–60 days, ask membership only if it is unknown; a premium reply and a standard reply must diverge.
In this simulator, only user statements supply customer facts. The policy’s numbers must not be read as an age the customer supplied. The latest matching explicit statement wins; absent facts remain null. This still does not support arbitrary corrections, multiple purchases, or phrases such as four weeks ago.
The two fixes solve different information losses. Joining user history into retrieval makes the return keyword available. Including assistant turns in model messages supplies the question that gives a short reply meaning. Fixing just one can still leave the other broken. Trace both paths before modifying the parser.
1node typescript/evaluate.ts --offline2python3 evaluate.py --offline3node typescript/chat.ts --offlineExercise: follow 29 through three layers
Compare codex/01-start with codex/02-history. In the interactive offline session, ask to return a product and answer the purchase-age question with 29. Inspect the retrieval query, policy, and constructed messages. Then test 30 and 61 in fresh sessions. Change the simulator’s preceding question wording in your own scratch copy and observe what the parser no longer recognizes.
Do not treat the 16-case score as a continuation of an identical eight-case benchmark. The suite grew and the simulator changed. Keep a note of branch, case count, and backend with every result. The original project also checked both language implementations offline; that is a useful wiring comparison, not evidence that Python and TypeScript calls to a real model are equivalent.
At this point the lab can carry a short answer through the expected offline path. The next chapter keeps that plumbing and changes the backend. That is where a passing simulator stops being reassuring.
Lab source and official references