The first live baseline reached 13/16, not the simulator’s 16/16. That result was useful because it gave us failures to investigate. It was not useful enough to tell us which layer had failed.
Understand where the traces come from
We wrote our own traces. They are not an automatic trace document downloaded from the Responses API.
| Information | Source |
|---|---|
| Retrieved policy and constructed messages | Our chatbot.ts |
| Model answer, usage, response ID | The API response, extracted by backend.ts |
| Measured latency | Our local timers |
| Expected decision and pass/fail result | Our eval harness |
| JSON report file | Our evaluate.ts |
Originally, the trace was simply:
1return {2 ...answer,3 trace: { context, messages },4};To confirm that history was sent, we inspected actual.trace.messages and checked that the backend used input: messages in its HTTP body.
That establishes the application path to the request. It does not expose model reasoning or prove the model used each message correctly.
Know which report you are reading
The latest live TypeScript report is:
results.typescript.openai.jsonEach TypeScript live eval overwrites it, including single-case runs. Check:
1created_at2selected3executed4passed5errors6resultsA report with selected: 1 is not a full regression run. The equivalent Python and offline reports are separate files.
We initially accumulated old baselines and smoke-test reports with similar names. Cleaning them up made the debugging session easier to follow. Timestamped archives would be a useful future improvement, but are not implemented here.
Improve the prompt—and check what actually changed
One live failure asked about membership for a purchase made 30 days ago, even though the answer did not depend on membership.
We made the decision order more explicit:
1"If purchase age is missing, ask how many days ago the item was purchased. " +2"If the purchase was at most 30 days ago, it is eligible regardless of membership. " +3"Only ask about missing membership when the purchase was more than 30 " +4"but no more than 60 days ago. Purchases older than 60 days are ineligible. "The 30-day case passed. But a subsequent full run still scored 13/16, with a different set of failures.
We also added instructions to interpret short replies using the preceding question and avoid treating identifiers as ages. The order-number failure persisted even when the trace confirmed those instructions were present.
Learning: Inspect the recorded prompt before changing it again. “The new instruction was not sent” and “the instruction was sent but did not reliably work” are different diagnoses.
We experimented with a separate developer message for reference policy. The current implementation retains the policy in the final user-message wrapper as well, partly preserving the offline simulator's input contract. It is duplicated; we did not establish that message-role separation alone fixed the behavior.
Separate two superficially similar failures
These two cases both involved "29", but required different fixes.
Case A: an order number mistaken for purchase age
1Assistant: What is your order number?2User: 293 4Expected: ask_days5Actual: eligibleThe model answered as though the purchase was 29 days old. Policy was available. The missing fact was purchase age.
Case B: a number with no conversation or retrieved policy
1User: 292 3Expected: no_context4Actual: ask_membershipThe model invented a purchase-age interpretation and continued a return conversation despite retrieval finding no policy.
We had also repeated policy rules in the system prompt, so empty retrieval did not mean the model had no policy-like information. That made relying on the model to detect this condition especially weak.
Compare failure identities, not just totals
A decision-order instruction made the 30-day case pass, yet a subsequent full run still scored 13/16. The number was unchanged; the failures were not. That is why I would keep an archive before each new live run, then compare case IDs and recorded inputs. The lab itself still overwrites reports; timestamped archival is a proposed local practice, not an implemented feature.
1node typescript/evaluate.ts --case=followup-days-302# This selected run overwrites results.typescript.openai.json.3# Preserve any report you want before running the full suite.4node typescript/evaluate.tsA report’s selected count is how many cases the command asked to run. Its executed count tells you how far the runner got. Neither is a model-call count: application guards can return without calling the backend. The next two chapters explain those branches and partial-run reporting. For now, resist interpreting an incomplete report as a new full-suite score.
The trace’s messages tell you what your app constructed. Confirming input: messages in backend.ts closes the application path to the HTTP body for a model-called answer. It still does not reveal reasoning or prove every message influenced the model. For guard-generated answers, constructed messages may never be sent at all.
A local report-inspection tool
This diagnostic snippet is a new teaching utility, not an original saved result. Run it from the lab root against a report you created. It prints a small summary and case-level evidence without request headers or credentials.
1import { readFileSync } from "node:fs";2 3const report = JSON.parse(4 readFileSync("results.typescript.openai.json", "utf8"),5);6console.log({7 created_at: report.created_at,8 selected: report.selected,9 executed: report.executed,10 passed: report.passed,11 errors: report.errors,12});13for (const result of report.results) {14 console.log(result.id, result.expected, result.actual?.decision);15 console.log(result.actual?.trace?.context);16 console.log(result.actual?.trace?.messages);17}Reports contain conversation data, even when this lab uses synthetic cases. They are ignored by Git for a reason. For a real application, redact customer data and scope who can read traces. Never add authorization headers merely to make debugging more convenient.
Exercise: diagnose the two 29s
Read number-without-days-question and fresh-number-no-context in shared cases.json. Before changing anything, write a separate hypothesis for each: whether relevant policy exists, whether purchase age is supplied, and what the expected action is. Compare a selected live report with a full report, preserving each locally first. A failed or skipped live call is still evidence about the execution path, not a new historical result.
The next chapter handles these diagnoses with two different responsibilities: an empty-context guard and explicit pending-question state. It keeps the full before/after trace in one place so the architectural lesson is not buried in repeated summaries.
Lab source and official references