Get an AI Agent to Debug My API

Necco Ceresani
Necco Ceresani
·15 min read

Fifteen years building engineering platforms, currently focused on advanced AI infrastructure at yeet. I love turning the deeply complex topics into something everyone can understand. I relate deeply with the core yeet philosophy that you can just build things.

Last updated: October 2026

Quick answer. An agent debugging an API usually has three inputs: your code, your logs, and whatever you paste into the chat. None of them contain the request that failed, so the agent reasons about what the code should have done and proposes fixes you then have to test. The change that matters is giving it the traffic. With yeet installed, the agent can read the HTTP the machine actually sent and received, including bodies, read from the kernel with eBPF rather than from code anybody instrumented. Then "why is checkout returning 502s" stops being a question about your source and becomes a question it can answer from evidence: which endpoint failed, what the response body said, what the service asked Postgres and Redis while it was failing, and whether the connection was reset before an answer ever came.

An agent is good at the part of debugging that is reading and correlating, and bad at the part that is guessing. Given a repository and a stack trace it will do the reading well, form a plausible story, and suggest a change. The trouble is that the story is a hypothesis built from what the code says should happen, and the bug exists precisely because something happened that the code did not describe.

So you end up in the loop everyone recognises: the agent proposes, you deploy, the failure continues, you paste more logs, it proposes again. Each turn costs a deploy and adds no new evidence, because the input never changed. The useful question is not how to prompt it better. It is what to give it to read.

Why can't my AI agent debug an API from the code and logs?

Because both describe intentions rather than events. Source code says what the program would do given inputs it expects. A log line exists because a developer anticipated that moment being interesting. The failures that survive to production are the ones nobody anticipated, which is exactly the set that leaves no log line and contradicts what the code implies.

Three gaps show up again and again, and an agent cannot close any of them by reading harder:

  • The request that failed is gone. Your service logged that it returned 500. It did not log the body that caused it, the header that was missing, or the token that had expired. The agent is told the outcome and asked to infer the input.
  • The dependency's answer is unrecorded. When a handler fails because Postgres was slow or a third-party API returned 429, the only trace is usually a generic timeout in the caller's log. What the dependency actually said is not written down anywhere.
  • Failures below HTTP leave nothing at all. A connection reset by a load balancer, a packet dropped, a DNS answer that went stale: none of these produce an application event, so from the agent's point of view they did not happen.

An agent reasoning without those is not being unintelligent. It is doing inference with the evidence withheld.

What does an AI agent get access to when yeet is installed?

The traffic itself, read from the kernel, as a thing it can query rather than a file it has to parse. Every HTTP request a machine sends or serves passes through the socket layer on its way, and every connection underneath has state the kernel already tracks. yeet reads both with eBPF and gives the agent a typed surface over them, so the agent asks questions instead of grepping.

The API debugging use case splits what it can reach into three layers, which is a useful way to think about what you are handing over:

LayerWhat the agent can askExamples it answers
Symptoms, HTTP and TLSWhich endpoint is slow, how often it fails, what the 500s returnedResponse bodies of the failures; a client stuck in a retry storm; API drift after a deploy
DependenciesWhat the service said to its databases and peersThe slow query, and the one running 200 times; the conversation between two services; rate-limit headroom with a third-party API
Infrastructure, TCP and DNSWhat happened below HTTPWho is resetting the connections; where the kernel dropped the packets; whether the connection pool is exhausted; stale DNS answers

The property that makes this work for an agent rather than just for a dashboard is that nothing in it is pre-judged. Every metric is a plain number on the row, and the agent composes the criteria, so "which endpoint is failing badly enough to matter" is a query it writes rather than a threshold somebody else picked. An agent handed a dashboard inherits opinions; an agent handed a query surface forms its own.

How does an AI agent debug a 502 API error with yeet?

It follows the failure down the stack rather than sideways through the repository. The prompt is ordinary, something like checkout-api.prod.internal is returning 502s, figure out why, and what changes is the evidence available at each step.

It reads the failing responses and the requests underneath them, so it starts from the actual exchange rather than a status code. It checks what checkout said to Postgres, Redis and payments during the window the failures happened, which is where the cause usually is. It looks below HTTP at resets, retransmits, drops and DNS answers, because a 502 is frequently a connection that died rather than a handler that threw. And it arrives at whether the cause is the code or the infrastructure with the evidence attached, which is the distinction that decides who fixes it.

That last step is worth dwelling on, because it is the one engineers argue about. A dependency call that came back with an error status reached its destination and was refused on its merits, which is code. A dependency call that got a reset instead of an answer, from something that is not the service, is infrastructure sitting in the path. Both produce a 502 at the gateway and identical lines in its log. They are only distinguishable if you can see the exchange and the connection together, which is the thing an agent with kernel access has and an agent with logs does not.

What can I run today to let an agent debug my API?

Start with apiwatch, which is published, runs in one command, and answers the first question most people have about a box: what APIs does this thing even serve and call.

curl -fsSL https://yeet.cx | sh    # install yeet, once
yeet run gh:yeet-src/apiwatch      # clone, build, and list this box's APIs from 30s of traffic

Point it at a server you did not build and a minute later you have the list nobody wrote down: which ports answer HTTP and which process owns each one, the endpoints they serve and the status codes they return, who calls them, and which outside APIs the machine calls in turn. It comes from traffic, so an API that exists only in a config file nobody reads still shows up, and a documented route nobody calls does not.

The reason this is a good first move with an agent specifically is that it converts an unknown machine into a described one in sixty seconds, which is the context the agent was missing. --watch then alerts on the responses your real clients got, with the failing request and response in the message, rather than on an uptime checker's own synthetic requests.

Is it safe to give an AI agent production API traffic to debug?

Treat it as database access rather than as a dashboard, because that is closer to what it is. Request and response bodies carry authorization headers, session tokens, personal data and payment details, and a capture holds them regardless of whether an agent or a person reads it. The right questions are who can query it, how long it is kept, and what is removed before anything leaves the host.

The tools differ in what they do about this, so check rather than assume. apiwatch redacts by key name and recognisable shape, which catches passwords, tokens, API keys and card numbers, and it is explicit that this is not everything: a secret under an innocent key, a free-text body, or personal data like names and email addresses goes through as captured, and --bodies off is the switch when that matters. If alerts route through Slack, those bodies pass through yeet's servers on the way, which is worth knowing before you point it at a channel.

The other half of safety is what the tooling can do rather than what it can see. apiwatch observes: it reports what crossed the socket, and it does not block, retry or modify a request, and does not probe your APIs itself. An agent given that surface can learn anything the traffic contains and change nothing, which is a different risk profile from an agent with a shell on the same box.

What can't an agent see in API traffic: Go TLS, HTTP/3, gRPC?

Enough that you should know the list before you rely on it, because the gaps are specific rather than general.

  • HTTPS from Go, and from TLS stacks that are not OpenSSL or rustls. Go's crypto/tls has no shared library to hook, so a Go program's HTTPS calls appear by hostname from the TLS handshake with no status codes, which httpscope addresses with uprobes on Go's own TLS, and a Go server terminating its own TLS is invisible as HTTP. Statically linked and stripped OpenSSL is the same.
  • HTTP/3. QUIC runs over UDP, and these hooks never see it.
  • gRPC failures that arrive as HTTP 200. gRPC reports errors in a grpc-status trailer on a 200 response, so a tool reading status codes alone does not count them as failures; grpcsnoop decodes gRPC calls and their messages.
  • Failures with no HTTP response at all. A called API that refuses the connection or never answers produces no status code; you hear about it only if your app turns it into a 5xx.
  • Bodies past about 32 KiB, which arrive cut, and anything that happened before the capture started, since state lives in memory.

None of these are reasons to skip the approach. They are reasons to know which question you are asking: a Go service's outbound HTTPS and a gRPC error path need different tools than a plaintext HTTP endpoint, and the honest move is to say so rather than let an agent conclude from silence that nothing failed.

The bottom line: to debug an API, change the agent's input

The reason agent-assisted debugging stalls on API failures is almost never the model. It is that the agent is asked to explain an event using artifacts that do not contain it, so it does the only thing available and produces a well-reasoned guess.

Giving it the traffic changes the shape of the work. The agent stops proposing hypotheses for you to test by deploying and starts reading what happened: the failing request, the response body, what the dependency said, whether the connection survived. That is the same investigation a good engineer does, and the reason to hand it to an agent is that it is mostly correlation, which agents are fast at and people find tedious.

Frequently asked questions

Can Claude Code debug a production API?

It can reason about your code and your logs, which is enough for a bug that reproduces and not enough for one that only happens in production. What changes the outcome is giving it the actual traffic to read: the failing request and response, what the service asked its dependencies, and the state of the connection underneath. With yeet installed on the host, that becomes something the agent can query rather than something you have to paste in.

What does an AI agent need to debug an API properly?

The exchange that failed, not a description of it. Concretely: the request body and headers that produced the error, the response the service actually returned, what its dependencies answered during the same window, and whether the connection was reset or timed out before any answer arrived. Code and logs supply none of those reliably, which is why an agent limited to them tends to propose changes rather than identify causes.

How do I give an AI agent access to network traffic?

Install a capture on the host that reads HTTP at the kernel level and exposes it to the agent as a queryable surface. With yeet, apiwatch is the published starting point: one command lists every API the machine serves and calls from live traffic, and the agent reads that rather than guessing at the machine's shape. There is a ready-made prompt that has an agent set it up end to end.

Is eBPF API capture safe to run in production?

It observes rather than intervenes: no proxy in the path, no TLS re-termination, no sidecar to deploy, and no modification of requests. The costs that are real are CPU on the capture path, which is why narrowing to specific ports matters on a busy database or file server, and the sensitivity of captured bodies, which should be governed like database access rather than like metrics.

Does an AI agent need root to read API traffic?

The capture does, because attaching eBPF programs is privileged, but the usual arrangement keeps that privilege in a daemon rather than handing it to the agent. The agent talks to a read-only interface, so it can ask what the traffic contained without holding the capability to attach probes, change the system, or run arbitrary commands.

Can an agent see HTTPS request bodies?

For TLS stacks it can hook, yes, by reading at the library's plaintext boundary, SSL_read and SSL_write, rather than decrypting the wire, which needs no private key. OpenSSL and rustls are covered this way. Go is the significant exception, because crypto/tls has no shared library to attach to, so a Go service's HTTPS shows up as a hostname from the handshake without status codes or bodies.

What is the difference between this and an APM agent?

An APM agent runs inside your application, which is how it sees spans and internal timings, and it only covers services you instrumented in languages it supports. A kernel capture sees every process on the host, including ones nobody instrumented and ones you did not write, and it sees the connection layer underneath HTTP. They overlap less than people expect: APM is better inside your code, kernel capture is better at everything between your services.

Will my agent just flood me with alerts?

That depends on the thresholds, and it is worth checking what a tool's defaults actually are. apiwatch alerts on an API's first 5xx, learns each API's normal 4xx share over five minutes before alerting on a jump, and treats a served port going quiet as an event. It sends one message when a problem starts, one when it recovers, and a reminder every thirty minutes, with simultaneous problems grouped into one message.

Can an agent debug an API it did not write?

That is where this is strongest, because the agent's disadvantage on unfamiliar code disappears when the evidence is the traffic. It does not need to understand a service's internals to observe that an endpoint returns 500 when a particular field is absent, or that the failing requests all arrived on connections that had been idle for several minutes. Reading behaviour requires no prior knowledge of the implementation.

Do I need distributed tracing for an agent to debug microservices?

It helps if you already have it and is not a prerequisite. Tracing stitches one request across services and needs every service instrumented and propagating context. Kernel capture gives per-hop exchanges with no instrumentation, which is less complete as a narrative and immediately available on a system that was never set up for tracing, which is the common case during an unplanned incident.

Sources

  1. yeet API debugging use case, for what an agent reads at the HTTP, dependency and infrastructure layers. https://yeet.cx/use-cases/api-debugging
  2. apiwatch, the published eBPF API watchdog that lists a machine's APIs from live traffic and alerts on failures, including its redaction behaviour and its stated limits. https://github.com/yeet-src/apiwatch
  3. RFC 9110, HTTP Semantics, for the status code classes an agent reads a failure against. https://www.rfc-editor.org/rfc/rfc9110.html
  4. Linux kernel BPF program types, for the attach points a capture uses to read traffic without entering the path. https://docs.kernel.org/bpf/libbpf/program_types.html
  5. yeet documentation, the runtime the capture and the agent interface run on. https://yeet.cx/docs/

Related resources

  1. How do I let Claude investigate production on Linux, on the access model that decides what an agent can reach. https://yeet.cx/topical-takes/let-claude-investigate-production
  2. Who is resetting my TCP connections on Linux, for the connection-layer evidence behind a 502 with no handler error. https://yeet.cx/topical-takes/who-is-resetting-my-tcp-connections
  3. How to monitor HTTP traffic on Linux, on capturing requests without a sidecar or an app change. https://yeet.cx/topical-takes/monitor-http-traffic-linux