Tracing an AI Agent Across Vercel and OpenAI: An Observability Deep Dive

One chat message can produce a serverless invocation, several model API calls, and local tool executions. When the answer is wrong—or missing—the challenge is finding which system owns the failure.

This post follows one request through a small support-agent project and shows what to inspect in Vercel Runtime Logs, OpenAI Agents SDK Traces, and the API Platform’s request and usage views. It is the engineering companion to Building an Agent You Can Evaluate, which covers setup, eval design, baselines, and improvement decisions.

Part 2 of a two-part series: this deep dive is for engineers following a request through the runtime and agent execution layers.

One request, multiple observability views

The browser posts to /api/demo. A Vercel Node.js function calls the Agents SDK Runner.run, which sends generation requests to OpenAI. If the model requests a tool, the SDK executes the registered JavaScript function in the server process and sends its result back to the model. The model can then make another generation before the answer returns.

Customer message
      │
      ▼
Browser ── POST /api/demo ──► Vercel Node.js function
                                   │
                                   ├── platform request metadata
                                   ├── app console events ─────► Vercel Runtime Logs
                                   │
                                   └── Agents SDK Runner.run
                                            │ traceId = this run
                                            │ groupId = conversation
                                            ▼
                                      OpenAI model ────────────► API request activity
                                            │ generation span     Usage Dashboard
                                            │
                                            └─ function span
                                                   ▼
                                          Local server tool/policy
                                                   │ result code
                                                   └──► model continues or answers

GitHub Actions / local eval runner ──► tests, outcomes, tool counts, elapsed time

Vercel executes the HTTP request; the Agents SDK orchestrates the run; the model generates text and proposes tool calls; server functions execute tools and enforce policy. One inbound request can contain multiple model calls. A later message creates a new request and trace, linked to the earlier one by the conversation group ID.

Keep the identifiers straight

In this app, api/demo.js generates a custom traceId for each submitted chat turn and uses the browser’s conversationId as the Agents SDK groupId:

const traceId = `trace_${randomUUID().replaceAll("-", "")}`;
const { conversationId } = body;

const runner = new Runner({
  workflowName: "Northstar Shop support chat",
  traceId,
  groupId: conversationId,
  traceIncludeSensitiveData: false,
});

The IDs answer different questions:

IdentifierCreated byUse it to find
Vercel request IDVercelThe platform invocation: route, status, timing, deployment, and runtime details.
App traceIdapi/demo.jsOne submitted message’s agent run, including its model and tool spans in OpenAI Traces.
conversationId / SDK groupIdBrowser/appThe set of separate agent runs from one chat conversation.
OpenAI API request IDOpenAI APIOne individual model API request, useful when investigating that request’s status or service behavior.

Do not assume similarly named trace fields are interchangeable. Vercel can show its own request, invocation, or tracing metadata. Search OpenAI Traces with the app trace ID printed in the Vercel application log; use the Vercel request ID to locate the platform request. For a whole multi-message chat, use the conversation ID/group ID.

What is logged in each system

SystemWhat it ownsWhat to inspect
BrowserSends messages and renders responses.Network result and returned IDs; not authoritative for policy or ownership.
Vercel Runtime LogsRuns the deployed function.Method/path, status, duration, deployment, platform metadata, app console lines. The app omits the full chat body.
OpenAI Agents SDK TracesCaptures the agent run.Span hierarchy, generations, functions, timings, and run outcome. Default spans include task, agent, turn, generation, and function.
OpenAI API PlatformServes model inference.HTTP Requests view for individual failed/slow calls; Usage Dashboard for aggregate usage/cost. Separate from SDK traces and Vercel logs.
Server tools and policy codeChecks session, ownership, eligibility, and allowed changes.Deterministic result codes and state. Fix policy/data bugs here and add tests.
Eval runner / GitHub ActionsMeasures scenarios and gates changes.Outcomes, pass rate, tool count, elapsed time. CI is deterministic; live evals are metered.

A log is an event, a span is a timed operation, a trace connects spans for one run, and a metric aggregates runs. A trace explains one execution; evals show behavior across cases.

The app sets traceIncludeSensitiveData: false, excluding LLM and tool inputs/outputs from traces while retaining spans. It separately logs safe result codes with the app trace ID. Thus the trace will not contain the customer’s full prompt, tool arguments, or tool-result body. See the Agents SDK tracing guide and Vercel Runtime Logs.

Correlate a Vercel log with an OpenAI trace

The following IDs are invented and the lines are shortened, but the event names and app fields match the project. Vercel’s platform request metadata is shown separately from the app’s own log messages:

Vercel platform: POST /api/demo  status=200  duration=4.1s  requestId=req_91…
app stdout:      {"event":"agent_run_started","conversationId":"conv_7f…","traceId":"trace_4a…"}
app stdout:      {"event":"agent_tool_result","traceId":"trace_4a…","tool":"request_shipping_fee_refund_review","code":"REFUND_NOT_ELIGIBLE"}
app stdout:      {"event":"agent_run_completed","conversationId":"conv_7f…","traceId":"trace_4a…","turns":5}

The actual turns value is result.newItems?.length ?? 0; it is a count of SDK result items, not necessarily model generations. The tool-result line includes a trace ID but not a conversation ID. That lets you connect the tool outcome to a run, while the start/completion lines connect that run to the conversation.

In OpenAI Traces, the same run is represented as a span tree. The exact display labels can vary by SDK version; conceptually it looks like this:

Northstar Shop support chat   traceId=trace_4a…   groupId=conv_7f…
└─ task / agent / turn
   ├─ generation span         model decides whether a tool is needed
   ├─ function span           request_shipping_fee_refund_review
   │                          payload omitted by trace redaction
   └─ generation span         model forms its final answer

The function span shows that the tool ran and where it fits in the agent loop. The Vercel agent_tool_result event gives the safe outcome code, REFUND_NOT_ELIGIBLE. Payload capture is disabled, so the trace omits the message and full tool input/output. Use the API Requests view for individual model-call errors or latency, the SDK trace for agent flow, and the Usage Dashboard for aggregate usage and cost.

The API handler calls forceFlush() before responding so the trace exporter has a chance to send the spans. If flushing fails, it writes agent_trace_flush_failed with the same conversation and trace IDs. That event is useful when the Vercel request completed but the expected OpenAI trace is absent.

Triage by symptom

Follow the failure to the system that owns it:

  1. Browser shows 5xx or times out. Start in Vercel Runtime Logs. Filter by deployment, route, time, and status; inspect duration and agent_run_failed. Check for agent_run_started and server config such as OPENAI_API_KEY. Fix the runtime or API boundary if the request failed before completion.
  2. Request succeeds, but the model picks a poor tool or stops early. Use traceId from agent_run_started to inspect generation → function → generation. Improve instructions, tool descriptions, or eval scenarios. Keep authorization and policy checks in server code.
  3. Right tool ran, but returned the wrong result or changed state incorrectly. Match agent_tool_result by traceId; inspect its code, session, fixture, and tool implementation. Fix domain logic and add a deterministic test.
  4. Tool result is right, but the answer is incomplete. Compare the final generation with the rubric. Improve instructions or response construction and add a grader regression case. For critical dates or policy disclosures, deterministic text can give stronger guarantees.
  5. Traces look fine, but pass rate or latency declines. Segment eval failures by scenario; compare tool counts and elapsed time, and check grader false positives. Improve coverage or measurement before optimizing. Three cases cannot establish a production trend.
  6. Vercel shows a request, but OpenAI Traces does not. Search the app traceId; check agent_trace_flush_failed, deployment/time range, and whether this was hosted chat. Local deterministic controls produce no OpenAI trace.

Correlate IDs, classify the failure, and fix the layer that owns it. A real system also needs retention, access control, alerting, and a policy for sensitive trace data.

Continue the evals story

For the project setup, eval design, baseline numbers, and the changes those measurements motivated, read Building an Agent You Can Evaluate: From GitHub to Vercel. The source repository and live demo are available for trying the example with fictional data.

Jacob Aloysious
Jacob Aloysious
Engineering Leader

Engineering manager at Meta with 20 years of experience across enterprise platforms, developer tools, and semiconductor software. I write about leadership, architecture, and applied AI.

Related