<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Observability | Jacob Aloysious</title><link>https://jacobaloysious.in/tag/observability/</link><atom:link href="https://jacobaloysious.in/tag/observability/index.xml" rel="self" type="application/rss+xml"/><description>Observability</description><generator>Source Themes Academic (https://sourcethemes.com/academic/)</generator><language>en-us</language><lastBuildDate>Sun, 11 Oct 2026 00:00:00 +0000</lastBuildDate><image><url>https://jacobaloysious.in/img/jacob-leadership-social.png</url><title>Observability</title><link>https://jacobaloysious.in/tag/observability/</link></image><item><title>Tracing an AI Agent Across Vercel and OpenAI: An Observability Deep Dive</title><link>https://jacobaloysious.in/post/tracing-agent-across-vercel-openai/</link><pubDate>Sun, 11 Oct 2026 00:00:00 +0000</pubDate><guid>https://jacobaloysious.in/post/tracing-agent-across-vercel-openai/</guid><description>&lt;p>One chat message can produce a serverless invocation, several model API calls, and local tool executions. When the answer is wrong—or missing—the challenge is finding which system owns the failure.&lt;/p>
&lt;p>This post follows one request through a small support-agent project and shows what to inspect in Vercel Runtime Logs, OpenAI Agents SDK Traces, and the API Platform&amp;rsquo;s request and usage views. It is the engineering companion to
&lt;a href="https://jacobaloysious.in/post/learning-agent-evals-with-codex/" target="_blank" rel="noopener">Building and Evaluating an AI Agent&lt;/a>, which covers setup, eval design, baselines, and improvement decisions.&lt;/p>
&lt;p>&lt;strong>Part 2 of a two-part series:&lt;/strong> this deep dive is for engineers following a request through the runtime and agent execution layers.&lt;/p>
&lt;h2 id="one-request-multiple-observability-views">One request, multiple observability views&lt;/h2>
&lt;p>The browser posts to &lt;code>/api/demo&lt;/code>. A Vercel Node.js function calls the Agents SDK &lt;code>Runner.run&lt;/code>, which sends generation requests to OpenAI. If the model requests a tool, the SDK executes the registered JavaScript function in the server process and sends its result back to the model. The model can then make another generation before the answer returns.&lt;/p>
&lt;pre>&lt;code class="language-text">Customer message
│
▼
Browser ── POST /api/demo ──► Vercel Node.js function
│
├── platform request metadata
├── app console events ─────► Vercel Runtime Logs
│
└── Agents SDK Runner.run
│ traceId = this run
│ groupId = conversation
▼
OpenAI model ────────────► API request activity
│ generation span Usage Dashboard
│
└─ function span
▼
Local server tool/policy
│ result code
└──► model continues or answers
GitHub Actions / local eval runner ──► tests, outcomes, tool counts, elapsed time
&lt;/code>&lt;/pre>
&lt;p>Vercel executes the HTTP request; the Agents SDK orchestrates the run; the model generates text and proposes tool calls; server functions execute tools and enforce policy. One inbound request can contain multiple model calls. A later message creates a new request and trace, linked to the earlier one by the conversation group ID.&lt;/p>
&lt;h2 id="keep-the-identifiers-straight">Keep the identifiers straight&lt;/h2>
&lt;p>In this app, &lt;code>api/demo.js&lt;/code> generates a custom &lt;code>traceId&lt;/code> for each submitted chat turn and uses the browser&amp;rsquo;s &lt;code>conversationId&lt;/code> as the Agents SDK &lt;code>groupId&lt;/code>:&lt;/p>
&lt;pre>&lt;code class="language-js">const traceId = `trace_${randomUUID().replaceAll(&amp;quot;-&amp;quot;, &amp;quot;&amp;quot;)}`;
const { conversationId } = body;
const runner = new Runner({
workflowName: &amp;quot;Northstar Shop support chat&amp;quot;,
traceId,
groupId: conversationId,
traceIncludeSensitiveData: false,
});
&lt;/code>&lt;/pre>
&lt;p>The IDs answer different questions:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Identifier&lt;/th>
&lt;th>Created by&lt;/th>
&lt;th>Use it to find&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>Vercel request ID&lt;/td>
&lt;td>Vercel&lt;/td>
&lt;td>The platform invocation: route, status, timing, deployment, and runtime details.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>App &lt;code>traceId&lt;/code>&lt;/td>
&lt;td>&lt;code>api/demo.js&lt;/code>&lt;/td>
&lt;td>One submitted message&amp;rsquo;s agent run, including its model and tool spans in OpenAI Traces.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>conversationId&lt;/code> / SDK &lt;code>groupId&lt;/code>&lt;/td>
&lt;td>Browser/app&lt;/td>
&lt;td>The set of separate agent runs from one chat conversation.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>OpenAI API request ID&lt;/td>
&lt;td>OpenAI API&lt;/td>
&lt;td>One individual model API request, useful when investigating that request&amp;rsquo;s status or service behavior.&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>Do not assume similarly named trace fields are interchangeable. Vercel can show its own request, invocation, or tracing metadata. Search OpenAI Traces with the &lt;strong>app trace ID printed in the Vercel application log&lt;/strong>; use the Vercel request ID to locate the platform request. For a whole multi-message chat, use the conversation ID/group ID.&lt;/p>
&lt;h2 id="what-is-logged-in-each-system">What is logged in each system&lt;/h2>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>System&lt;/th>
&lt;th>What it owns&lt;/th>
&lt;th>What to inspect&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>Browser&lt;/td>
&lt;td>Sends messages and renders responses.&lt;/td>
&lt;td>Network result and returned IDs; not authoritative for policy or ownership.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Vercel Runtime Logs&lt;/td>
&lt;td>Runs the deployed function.&lt;/td>
&lt;td>Method/path, status, duration, deployment, platform metadata, app &lt;code>console&lt;/code> lines. The app omits the full chat body.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>OpenAI Agents SDK Traces&lt;/td>
&lt;td>Captures the agent run.&lt;/td>
&lt;td>Span hierarchy, generations, functions, timings, and run outcome. Default spans include task, agent, turn, generation, and function.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>OpenAI API Platform&lt;/td>
&lt;td>Serves model inference.&lt;/td>
&lt;td>HTTP Requests view for individual failed/slow calls; Usage Dashboard for aggregate usage/cost. Separate from SDK traces and Vercel logs.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Server tools and policy code&lt;/td>
&lt;td>Checks session, ownership, eligibility, and allowed changes.&lt;/td>
&lt;td>Deterministic result codes and state. Fix policy/data bugs here and add tests.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Eval runner / GitHub Actions&lt;/td>
&lt;td>Measures scenarios and gates changes.&lt;/td>
&lt;td>Outcomes, pass rate, tool count, elapsed time. CI is deterministic; live evals are metered.&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>A &lt;strong>log&lt;/strong> is an event, a &lt;strong>span&lt;/strong> is a timed operation, a &lt;strong>trace&lt;/strong> connects spans for one run, and a &lt;strong>metric&lt;/strong> aggregates runs. A trace explains one execution; evals show behavior across cases.&lt;/p>
&lt;p>The app sets &lt;code>traceIncludeSensitiveData: false&lt;/code>, excluding LLM and tool inputs/outputs from traces while retaining spans. It separately logs safe result codes with the app trace ID. Thus the trace will not contain the customer&amp;rsquo;s full prompt, tool arguments, or tool-result body. See the
&lt;a href="https://openai.github.io/openai-agents-js/guides/tracing/" target="_blank" rel="noopener">Agents SDK tracing guide&lt;/a> and
&lt;a href="https://vercel.com/docs/logs/runtime" target="_blank" rel="noopener">Vercel Runtime Logs&lt;/a>.&lt;/p>
&lt;h2 id="correlate-a-vercel-log-with-an-openai-trace">Correlate a Vercel log with an OpenAI trace&lt;/h2>
&lt;p>The following IDs are invented and the lines are shortened, but the event names and app fields match the project. Vercel&amp;rsquo;s platform request metadata is shown separately from the app&amp;rsquo;s own log messages:&lt;/p>
&lt;pre>&lt;code class="language-text">Vercel platform: POST /api/demo status=200 duration=4.1s requestId=req_91…
app stdout: {&amp;quot;event&amp;quot;:&amp;quot;agent_run_started&amp;quot;,&amp;quot;conversationId&amp;quot;:&amp;quot;conv_7f…&amp;quot;,&amp;quot;traceId&amp;quot;:&amp;quot;trace_4a…&amp;quot;}
app stdout: {&amp;quot;event&amp;quot;:&amp;quot;agent_tool_result&amp;quot;,&amp;quot;traceId&amp;quot;:&amp;quot;trace_4a…&amp;quot;,&amp;quot;tool&amp;quot;:&amp;quot;request_shipping_fee_refund_review&amp;quot;,&amp;quot;code&amp;quot;:&amp;quot;REFUND_NOT_ELIGIBLE&amp;quot;}
app stdout: {&amp;quot;event&amp;quot;:&amp;quot;agent_run_completed&amp;quot;,&amp;quot;conversationId&amp;quot;:&amp;quot;conv_7f…&amp;quot;,&amp;quot;traceId&amp;quot;:&amp;quot;trace_4a…&amp;quot;,&amp;quot;turns&amp;quot;:5}
&lt;/code>&lt;/pre>
&lt;p>The actual &lt;code>turns&lt;/code> value is &lt;code>result.newItems?.length ?? 0&lt;/code>; it is a count of SDK result items, not necessarily model generations. The tool-result line includes a trace ID but not a conversation ID. That lets you connect the tool outcome to a run, while the start/completion lines connect that run to the conversation.&lt;/p>
&lt;p>In OpenAI Traces, the same run is represented as a span tree. The exact display labels can vary by SDK version; conceptually it looks like this:&lt;/p>
&lt;pre>&lt;code class="language-text">Northstar Shop support chat traceId=trace_4a… groupId=conv_7f…
└─ task / agent / turn
├─ generation span model decides whether a tool is needed
├─ function span request_shipping_fee_refund_review
│ payload omitted by trace redaction
└─ generation span model forms its final answer
&lt;/code>&lt;/pre>
&lt;p>The function span shows that the tool ran and where it fits in the agent loop. The Vercel &lt;code>agent_tool_result&lt;/code> event gives the safe outcome code, &lt;code>REFUND_NOT_ELIGIBLE&lt;/code>. Payload capture is disabled, so the trace omits the message and full tool input/output. Use the API Requests view for individual model-call errors or latency, the SDK trace for agent flow, and the Usage Dashboard for aggregate usage and cost.&lt;/p>
&lt;p>The API handler calls &lt;code>forceFlush()&lt;/code> before responding so the trace exporter has a chance to send the spans. If flushing fails, it writes &lt;code>agent_trace_flush_failed&lt;/code> with the same conversation and trace IDs. That event is useful when the Vercel request completed but the expected OpenAI trace is absent.&lt;/p>
&lt;h2 id="triage-by-symptom">Triage by symptom&lt;/h2>
&lt;p>Follow the failure to the system that owns it:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Browser shows 5xx or times out.&lt;/strong> Start in Vercel Runtime Logs. Filter by deployment, route, time, and status; inspect duration and &lt;code>agent_run_failed&lt;/code>. Check for &lt;code>agent_run_started&lt;/code> and server config such as &lt;code>OPENAI_API_KEY&lt;/code>. Fix the runtime or API boundary if the request failed before completion.&lt;/li>
&lt;li>&lt;strong>Request succeeds, but the model picks a poor tool or stops early.&lt;/strong> Use &lt;code>traceId&lt;/code> from &lt;code>agent_run_started&lt;/code> to inspect generation → function → generation. Improve instructions, tool descriptions, or eval scenarios. Keep authorization and policy checks in server code.&lt;/li>
&lt;li>&lt;strong>Right tool ran, but returned the wrong result or changed state incorrectly.&lt;/strong> Match &lt;code>agent_tool_result&lt;/code> by &lt;code>traceId&lt;/code>; inspect its code, session, fixture, and tool implementation. Fix domain logic and add a deterministic test.&lt;/li>
&lt;li>&lt;strong>Tool result is right, but the answer is incomplete.&lt;/strong> Compare the final generation with the rubric. Improve instructions or response construction and add a grader regression case. For critical dates or policy disclosures, deterministic text can give stronger guarantees.&lt;/li>
&lt;li>&lt;strong>Traces look fine, but pass rate or latency declines.&lt;/strong> Segment eval failures by scenario; compare tool counts and elapsed time, and check grader false positives. Improve coverage or measurement before optimizing. Three cases cannot establish a production trend.&lt;/li>
&lt;li>&lt;strong>Vercel shows a request, but OpenAI Traces does not.&lt;/strong> Search the app &lt;code>traceId&lt;/code>; check &lt;code>agent_trace_flush_failed&lt;/code>, deployment/time range, and whether this was hosted chat. Local deterministic controls produce no OpenAI trace.&lt;/li>
&lt;/ol>
&lt;p>Correlate IDs, classify the failure, and fix the layer that owns it. A real system also needs retention, access control, alerting, and a policy for sensitive trace data.&lt;/p>
&lt;h2 id="continue-the-evals-story">Continue the evals story&lt;/h2>
&lt;p>For the project setup, eval design, baseline numbers, and the changes those measurements motivated, read
&lt;a href="https://jacobaloysious.in/post/learning-agent-evals-with-codex/" target="_blank" rel="noopener">Building and Evaluating an AI Agent: A Hands-On Case Study&lt;/a>. The
&lt;a href="https://github.com/jacobaloysious/ai-evals-support-agent" target="_blank" rel="noopener">source repository&lt;/a> and
&lt;a href="https://ai-evals-support-agent.vercel.app/" target="_blank" rel="noopener">live demo&lt;/a> are available for trying the example with fictional data.&lt;/p></description></item></channel></rss>