Back to blog

Building an OpenAI Chatbot · Part 3 of 6

Connect the Responses API

Building an OpenAI Chatbot · Part 3 of 6

Chris Eberl

Chris Eberl

Founder • Engineering Leader, GenAI, Data, ML

AI EngineeringOct 8, 20267 min read
Connect the Responses API

The repaired history path passed the visible offline cases. The next question was not whether the simulator looked convincing; it was whether the same inputs produced the expected decisions from a real model.

Use the existing lab, not a new scaffold

Responses API flow from messages and schema through HTTP status, completion and refusal checks to validated answer.
The expected decision stays in the eval harness, not the API request.

The public companion repository already contains matching TypeScript and Python implementations, shared cases.json and answer.schema.json, backend selection, interactive chat, state helpers, and the exercise guide. Clone or download it locally. No package installation is required. Follow its README for Node 22.18+ or Python 3.11+; Node’s native TypeScript execution strips types but does not type-check.

Shell
1node --version2python3 --version3cp .env.example .env4# Add your own key to the local ignored .env only.5node typescript/evaluate.ts --offline6python3 evaluate.py --offline

Connect the real Responses API

Next, we replaced the simulator as the default backend with a real OpenAI model. The simulator remained available through --offline.

We used gpt-4.1-mini for this experiment. It is configurable through OPENAI_MODEL; this article does not claim it is the newest or universally best model.

Local configuration

Create a local .env file with your own key. The example intentionally leaves the key empty:

Code
1OPENAI_API_KEY=2OPENAI_MODEL=gpt-4.1-mini

Ignore secrets and generated artifacts:

Code
1.env2node_modules/3__pycache__/4results*.json5.DS_Store

The TypeScript backend loads the file relative to its own location:

Code
1const envPath = fileURLToPath(new URL("../.env", import.meta.url));2if (existsSync(envPath)) loadEnvFile(envPath);

The key is read at runtime and used only in the authorization header. We do not put it in messages, traces, screenshots, reports, or published code.

The actual HTTP request

This is the request shape used by the project:

Code
1const response = await fetch("https://api.openai.com/v1/responses", {2  method: "POST",3  headers: {4    Authorization: `Bearer ${key}`,5    "Content-Type": "application/json",6  },7  body: JSON.stringify({8    model,9    input: messages,10    store: false,11    max_output_tokens: 512,12    text: {13      format: {14        type: "json_schema",15        name: "support_answer",16        strict: true,17        schema,18      },19    },20  }),21  signal: AbortSignal.timeout(60_000),22});

key is a runtime variable, never a published credential. The complete parsing and error path is already available in typescript/backend.ts and backend.py in the repository.

We used the Responses API Structured Outputs format to request the decision and text object. The schema constrains the shape of a completed answer; it does not guarantee the answer is factually correct. Refusals and incomplete responses also require handling.

We used native fetch to keep the TypeScript project dependency-free. Because it runs .ts files through Node's native type stripping, this lab requires Node 22.18+. This execution path does not type-check the code. The Python implementation uses the standard library and requires Python 3.11+.

Parse the response rather than assuming the first output item is text

The raw HTTP response can contain different output item types. Our backend collects message content and extracts output_text entries:

Code
1const content = payload.output.flatMap(2  (item: { content?: { type: string; text?: string }[] }) =>3    item.content ?? [],4);5 6const text = content7  .filter((part: { type: string }) => part.type === "output_text")8  .map((part: { text?: string }) => part.text ?? "")9  .join("");

It checks completion status, refusals, and the parsed answer. Errors are reported separately from wrong answers. The app never silently substitutes the simulator after an API failure.

Handle completed, refused, and invalid answers

The repository backend checks HTTP status before parsing the payload, rejects incomplete responses, rejects refusal content, and validates the parsed decision and text. This is the current lab excerpt; the earlier output_text extraction above retains the original article’s request example.

Code
1if (!response.ok) {2    throw new Error(`OpenAI HTTP ${response.status}. Check API key, billing, rate limits, and model access.`);3  }4  const payload = await response.json();5  if (payload.status !== "completed") throw new Error(`OpenAI response did not complete (${payload.status}).`);6  const content = payload.output.flatMap((item: { content?: { type: string; text?: string }[] }) => item.content ?? []);7  if (content.some((part: { type: string }) => part.type === "refusal")) throw new Error("Model refused the request.");8  const text = content.filter((part: { type: string }) => part.type === "output_text")9    .map((part: { text?: string }) => part.text ?? "").join("");10  let answer: Answer;11  try {12    answer = JSON.parse(text);13    if (!answer || !schema.properties.decision.enum.includes(answer.decision) || typeof answer.text !== "string") {14      throw new Error();15    }16  } catch {17    throw new Error("Model returned an invalid answer object.");18  }

A schema requests the shape of a completed answer. It does not guarantee eligibility is correct or the explanation agrees with the decision. Refusal content is not a policy answer. A timeout, missing key, bad JSON, or incomplete response is an error, not a FAIL decision. The lab does not silently switch to offline mode when live mode fails.

Choose live or offline deliberately

From chatbot_lab:

Code
1npm run chat2npm run eval

For a single case:

Code
npm run eval -- --case=followup-days-30

For the simulator:

Code
npm run eval -- --offline

The direct Node commands also work and bypass unrelated npm configuration warnings:

Code
1node typescript/chat.ts2node typescript/evaluate.ts

The first full live TypeScript baseline scored 13/16, even though the offline suite scored 16/16. A Python live smoke test also produced a different decision from an earlier TypeScript request on the same example.

This did not automatically mean the two implementations had different logic. Separate model calls can produce different responses. It was a reason to inspect inputs and repeat controlled tests, not assume a language-specific bug.

Shell
1python3 chat.py2python3 evaluate.py3python3 evaluate.py --offline

Live commands use your API account and send constructed conversation messages and retrieved policy to OpenAI. The request uses store: false, a 60-second timeout, and a 512-output-token cap; the lab implements no automatic retries. Those choices do not by themselves establish a retention or privacy guarantee for every API account. Use synthetic examples and review current provider documentation before adapting the lab.

The original first full live TypeScript baseline was 13/16. The repository’s completed main branch already includes later state guards, so running main now is not reconstructing that initial baseline. The history branch is a teaching checkpoint, not an exact snapshot of the original API chronology. Record your own branch, model, selected cases, and results instead of expecting the historical number.

Exercise: inspect one real request path

First run one selected case offline. If you choose to incur API usage, configure your own local key and run the same case live. Inspect context, messages, actual decision and explanation. Read backend.ts to verify input: messages and confirm that the expected answer never enters the API request. Do not paste your key or request headers into a trace.

Once a live answer differs, inspect the recorded information before adding another sentence to the prompt. The next chapter makes those records useful and explains why two identical aggregate scores can conceal different bugs.

Read next

Newsletter

New posts, straight from Chris

A short note from me whenever a new article goes live — product engineering, AI workflows, IoT, indie apps, and engineering leadership. No spam, unsubscribe anytime.

By subscribing, you agree to our Privacy Policy. We do not share your email.