Building an Agent You Can Evaluate: From GitHub to Vercel
I wanted to learn agent evaluations end to end, not just read about them. With Codex as a coding and teaching partner, I built a fictional customer-support agent, deployed it, and worked through a course project one step at a time: define behavior, test deterministic rules, measure a baseline, make a targeted change, and check what the evidence supports.
Part 1 of a two-part series: this post covers setup and evals. Part 2 is an engineer-focused observability deep dive.
The result is a learning project, not a production support platform. The useful outcome was learning how to make agent behavior testable and improve it without treating a small demo score as a reliability guarantee.
Try it: Live demo · Source repository · Course syllabus
Start with policy, not prompting
The example is an online shop with fictional customers, orders, delivery promises, and refund requests. Before tuning a prompt, I wrote down what the system should allow:
- Order access comes from the server’s authenticated session, never a customer ID asserted in chat.
- Address changes must check ownership, shipment status, serviceability, and delivery timing.
- A mismatch between the requested state and postal PIN needs separate, explicit confirmation.
- Refund requests outside policy are declined automatically; eligible requests go to a human for review. The agent cannot approve or issue money.
That policy shaped the trust boundary. The model can select a tool and formulate arguments. Server-side tools apply ownership and policy checks against trusted session and fixture data. Tool descriptions help the model choose an action; they are not authorization controls. The human review capability is intentionally not exposed to the agent.
How the project fits together
The browser sends a chat request to /api/demo. A Vercel Node.js function runs the OpenAI Agents SDK. The SDK sends model requests to OpenAI, executes registered JavaScript tools on the server when requested, and feeds tool results back into the model. It then returns a response and correlation IDs to the browser.
Browser → Vercel function → Agents SDK → OpenAI model
│ │
└─ local policy tools
→ tool result → model response
Vercel hosts and runs the application; it does not decide which tool the agent needs. The model proposes a tool call; the server executes that function and enforces policy. One browser submission can contain multiple model turns while remaining a single inbound HTTP request.
The project is in GitHub so each lesson leaves a reviewable change. GitHub Actions runs unit/API tests and deterministic policy scenarios on pushes and pull requests. It does not call the live model: those evals require an API key, incur usage, and can vary. Live chat and model evals use OPENAI_API_KEY on the server. I set it in Vercel’s project environment settings, alongside DEMO_SESSION_SECRET; neither secret belongs in the browser bundle. Saving an environment variable requires a new deployment before the running function sees it. ChatGPT subscriptions and API Platform billing are separate (
OpenAI billing details,
Vercel environment variables).
Build an evaluation ladder
I used different checks for different failure surfaces:
- Unit and API tests check deterministic rules such as ownership denials, refund eligibility, and side effects. They run without a model.
- Deterministic policy scenarios run proposals through server tools with fixed data and a controlled clock.
- Live model evals send customer language to the agent and inspect tool results, state changes, and answer requirements.
- Conversation evals preserve a transcript and shared fixture state. One scenario asks for a refund before the promised delivery date, advances the fixture clock, and asks again. The first request is declined; a later eligible one can enter the human review queue.
- Grader tests and adversarial cases check that the eval rejects unsafe outcomes while accepting valid alternative wording and tool paths.
The evaluator needed testing too. A correct explanation that an order was “within its promised delivery window” initially failed because the phrase check was too narrow. A cross-account case required one particular tool to report denial even though another server-checked path had the same safe outcome. We changed the rubric to consider policy outcome, side effects, and response meaning instead of one phrase or one valid route.
This is a useful quality-engineering distinction: a failed eval can mean the agent is wrong, the application is wrong, or the grader is wrong. A pass rate without failure classification hides those differences.
Measure a baseline, then make a narrow change
For the optimization exercise, I chose three ordinary tasks: check an order status, list the signed-in account’s orders, and change an eligible address while reporting its revised delivery promise. I ran the same live model eval before changing the prompt or tool description.
| Run | Passes | Average tool calls | Average elapsed time |
|---|---|---|---|
| Baseline | 2/3 (67%) | 1.00 | 3,766 ms |
| Prompt instruction added | 3/3 (100%) | 1.33 | 4,802 ms |
| Prompt + tool-description guidance | 3/3 (100%) | 1.00 | 3,322 ms |
The baseline failure was specific: the address-change tool succeeded and returned a revised promised date, but the final answer omitted it. I updated the rubric to require the new delivery promise, added an instruction to include the date returned by the tool, and added a grader regression test.
That fixed the answer requirement but raised average tool calls to 1.33. Inspecting the runs showed that for a clear address-change request, the agent first looked up the order and then called the address-change tool, which already checks ownership and shipment status. I clarified the tool description so the agent could call it directly when intent was clear, while retaining lookup for questions or ambiguity. The next run stayed at 3/3 and averaged one call per scenario. We reduced redundant work without weakening server checks.
Read the numbers honestly
These measurements compare the exact runs; they do not establish production quality or a statistically significant improvement. Three scenarios are a tiny sample. “100%” means 3 of 3 cases passed, not that the agent is 100% accurate. Elapsed time includes variable network and service conditions, so timing differences are suggestive only.
Tool-call count is an efficiency proxy. This eval runner does not expose token usage, so it cannot support a dollar-cost calculation. One fewer tool call may reduce work, but it does not tell us how many tokens or dollars were saved. Once I used the optimization cases to guide changes, they were no longer a pristine holdout. I created separate final-validation cases—PIN/state mismatch, shipped order, and an untrusted lost-package claim—and ran them after the change.
For a technical leader, the lesson is the measurement loop: define the customer outcome, set a baseline, inspect a failure, make a narrow change, and evaluate again. Keep outcome quality, safety/side effects, efficiency, and operational health separate; one score cannot stand in for all four.
Continue with the observability deep dive
The next post follows one submitted message across Vercel and OpenAI: how the request ID differs from the agent trace ID, what the log and span views contain, and where an engineer should investigate each failure. Read Tracing an AI Agent Across Vercel and OpenAI.
The GitHub repository contains the code and lesson notes. The live demo uses fictional data. The project taught me to treat agent behavior as something to specify and measure, then make claims no stronger than the evidence.