Back to blog

Building an OpenAI Chatbot · Part 6 of 6

Read Results Honestly and Extend the Lab

Building an OpenAI Chatbot · Part 6 of 6

Chris Eberl

Chris Eberl

Founder • Engineering Leader, GenAI, Data, ML

AI EngineeringOct 8, 20267 min read
Read Results Honestly and Extend the Lab

The state chapter made the number’s meaning explicit and showed why one successful case needed no model call. The final task is less glamorous: report the result without turning a useful lab checkpoint into a reliability claim.

Make the output useful to a human

Eval coverage boundaries: sixteen visible cases, fourteen model calls, two guards, and unproven quality and security.
A passing visible suite is a checkpoint, not general reliability.

We colored console labels while keeping reports plain JSON:

Code
1function colorStatus(status: "PASS" | "FAIL" | "ERROR"): string {2  const colors = { PASS: 32, FAIL: 31, ERROR: 33 };3  const useColor =4    process.stdout.isTTY &&5    process.env.NO_COLOR === undefined &&6    process.env.TERM !== "dumb";7 8  return useColor9    ? `\x1b[${colors[status]}m${status}\x1b[0m`10    : status;11}
  • PASS / green: expected decision matched.
  • FAIL / red: the application answered, but the decision differed.
  • ERROR / yellow: the request, response parsing, or runtime failed.

The runner stops at the first API/runtime error, saves partial results, and reports how many cases remain unrun. A connection problem should not be mistaken for 16 wrong policy decisions.

The final traces include:

Text
1actual.trace.context2actual.trace.messages3actual.trace.modelCalled4actual.trace.stateBefore5actual.trace.stateAfterReply6actual.state

For an application-generated answer, modelCalled is false. The trace's messages were constructed but not sent. Live API answers also have usage and response metadata; direct application answers do not.

Read the final result honestly

The final saved live TypeScript run used gpt-4.1-mini and produced:

CategoryPassed
Basic2/2
Policy2/2
Missing information4/4
Multi-turn8/8
Total16/16

That run made 14 model calls. The remaining two cases were handled directly by the application: missing policy and the order-number clarification.

Both offline implementations also passed 16/16. Additional checks exercised order numbers 29, 007123, and 60, verified that order numbers stayed separate from purchase age, and passed the resulting state into subsequent age and membership replies.

Here is the progression we observed—not a controlled benchmark comparing identical systems:

CheckpointObserved resultInterpretation
Original offline simulator5/8Missing clarification and history behavior
Clarification instruction in simulator6/8Programmed clarification behavior improved
Expanded offline follow-up suite16/16Wiring worked for the simulator's rules
First full live model baseline13/16Real behavior differed from the simulator
Prompt iterationsSome runs still 13/16Failure identities changed; score alone hid that
State and application guards16/16 in the final live runThese cases passed with the revised architecture

We changed prompts, state inputs, and application behavior over time. We did not prove that any one prompt change alone caused the final score.

What remains unfinished

Passing this suite is a useful checkpoint. Several important gaps remain:

  • The grader checks labels, not explanations. An eligible decision paired with “you cannot return this” could still pass. Add separate response-quality checks.
  • The parser recognizes a small set of phrases. “Four weeks ago,” corrections, negative quantities, and multiple products need additional handling and tests.
  • The retrieval layer is a keyword placeholder. Old return mentions can keep a return policy in scope after the user changes topics.
  • Policy exists in multiple places. The system instructions repeat the limits, and the policy appears in both developer and user messages. A policy update could leave contradictory copies.
  • The tests were visible while we made changes. Add held-out cases and repeated runs to assess generalization and variability.
  • The state machine follows decision labels. It assumes the model's label agrees with its customer-facing question. We do not yet validate that agreement.
  • This is not a security evaluation. Role separation, JSON schemas, and regexes do not establish resistance to prompt injection.
  • The state is session-local. It does not persist across process restarts, and there is no customer identity or production storage layer.

For this particular policy, a future version could compute all eligibility decisions in application code once validated facts are available, leaving the model to handle language. We have not implemented that complete separation here.

A debugging workflow I would reuse

For each failure:

  • Read the input, expected behavior, and actual answer.
  • Inspect the trace to verify what context and history were supplied.
  • Classify the likely failure: retrieval, missing facts, interpretation, application logic, API behavior, or grading.
  • State one hypothesis and the smallest change that tests it.
  • Run the specific case, then the full suite.
  • Inspect which cases changed—not just the aggregate score.
  • Explain what the result supports and what remains uncertain.

Codex helped create files, inspect code, and run checks. The valuable learning was in reviewing those changes and asking why they should affect the observed failure. A useful instruction to a coding agent is specific:

“Inspect this failed case's recorded messages. Tell me whether purchase age is actually present. Do not change the prompt yet.”

Or:

“Add an early return for empty policy context, preserve the trace, and verify that this branch makes no API call.”

The shift is from asking the model to behave better in general to identifying a concrete responsibility and deciding whether it belongs in a prompt, a state transition, a retrieval step, or a deterministic branch.

Run checks, then extend the lab

The companion repository already provides tests and an exercise guide. On main, use the documented commands below. These are reader-run commands, not a claim that the original live experiment was rerun for this series. No installation is required; both languages use their existing lab implementations.

Shell
1npm test2python3 -m unittest discover -s tests -p 'test_*.py'3node typescript/evaluate.ts --offline4python3 evaluate.py --offline5# Optional live run: uses your API account.6node typescript/evaluate.ts

A nonzero eval exit after a decision mismatch is intentional. It makes regressions visible to automation. An ERROR means execution itself failed; the runner stops and saves a partial report with remaining cases unrun. Never silently convert those unrun cases into ordinary failures or quietly omit them from a success percentage.

Exercise: add evidence outside the visible suite

  • Create held-out cases before editing: four weeks ago, an age correction, two products, and a changed topic. Keep them separate from the cases driving the change.
  • Add a response-quality check for an eligible label paired with a refusal sentence. Explain why the current decision-only grader would pass it.
  • Check that the pending state agrees with the customer-facing question. A correct label is not enough if the text asks something different.
  • Repeat a small live set with the same model/configuration if you choose to use your account; archive reports and compare failure identities.
  • Design session persistence with scoped identity, expiry, reset semantics, and redacted traces before using real customer data. Treat it as a proposed extension, not an existing feature.

For a small deterministic return policy, I would consider computing eligibility from validated facts in application code and using the model for language. The original project did not implement that complete separation. It still asks a model to make most eligibility decisions; the final saved run’s 14 calls make that boundary visible.

Security requires its own evaluation. Structured Outputs does not authorize a refund, message roles do not make untrusted customer text safe, and regexes do not establish injection resistance. The lab has no authentication, customer identity, order lookup, production storage, or deployment. Review SECURITY.md before adapting it and avoid feeding private data into casual debugging reports.

My reusable workflow is deliberately modest: read a failure, inspect the evidence path, name a hypothesis, change the smallest responsible layer, run the focused case, then regress the full suite. The useful result is not just a higher fraction. It is an explanation of why a guarantee belongs in retrieval, model guidance, state, or ordinary code—and what the latest run still cannot tell us.

Read next

Newsletter

New posts, straight from Chris

A short note from me whenever a new article goes live — product engineering, AI workflows, IoT, indie apps, and engineering leadership. No spam, unsubscribe anytime.

By subscribing, you agree to our Privacy Policy. We do not share your email.