
Fifteen years building engineering platforms, currently focused on advanced AI infrastructure at yeet. I love turning the deeply complex topics into something everyone can understand. I relate deeply with the core yeet philosophy that you can just build things.
Last updated: October 2026
Quick answer. An agent debugging an API usually has three inputs: your code, your logs, and whatever you paste into the chat. None of them contain the request that failed, so the agent reasons about what the code should have done and proposes fixes you then have to test. The change that matters is giving it the traffic. With yeet installed, the agent can read the HTTP the machine actually sent and received, including bodies, read from the kernel with eBPF rather than from code anybody instrumented. Then "why is checkout returning 502s" stops being a question about your source and becomes a question it can answer from evidence: which endpoint failed, what the response body said, what the service asked Postgres and Redis while it was failing, and whether the connection was reset before an answer ever came.
An agent is good at the part of debugging that is reading and correlating, and bad at the part that is guessing. Given a repository and a stack trace it will do the reading well, form a plausible story, and suggest a change. The trouble is that the story is a hypothesis built from what the code says should happen, and the bug exists precisely because something happened that the code did not describe.
So you end up in the loop everyone recognises: the agent proposes, you deploy, the failure continues, you paste more logs, it proposes again. Each turn costs a deploy and adds no new evidence, because the input never changed. The useful question is not how to prompt it better. It is what to give it to read.
Because both describe intentions rather than events. Source code says what the program would do given inputs it expects. A log line exists because a developer anticipated that moment being interesting. The failures that survive to production are the ones nobody anticipated, which is exactly the set that leaves no log line and contradicts what the code implies.
Three gaps show up again and again, and an agent cannot close any of them by reading harder:
An agent reasoning without those is not being unintelligent. It is doing inference with the evidence withheld.
The traffic itself, read from the kernel, as a thing it can query rather than a file it has to parse. Every HTTP request a machine sends or serves passes through the socket layer on its way, and every connection underneath has state the kernel already tracks. yeet reads both with eBPF and gives the agent a typed surface over them, so the agent asks questions instead of grepping.
The API debugging use case splits what it can reach into three layers, which is a useful way to think about what you are handing over:
| Layer | What the agent can ask | Examples it answers |
|---|---|---|
| Symptoms, HTTP and TLS | Which endpoint is slow, how often it fails, what the 500s returned | Response bodies of the failures; a client stuck in a retry storm; API drift after a deploy |
| Dependencies | What the service said to its databases and peers | The slow query, and the one running 200 times; the conversation between two services; rate-limit headroom with a third-party API |
| Infrastructure, TCP and DNS | What happened below HTTP | Who is resetting the connections; where the kernel dropped the packets; whether the connection pool is exhausted; stale DNS answers |
The property that makes this work for an agent rather than just for a dashboard is that nothing in it is pre-judged. Every metric is a plain number on the row, and the agent composes the criteria, so "which endpoint is failing badly enough to matter" is a query it writes rather than a threshold somebody else picked. An agent handed a dashboard inherits opinions; an agent handed a query surface forms its own.
It follows the failure down the stack rather than sideways through the repository. The prompt is ordinary, something like checkout-api.prod.internal is returning 502s, figure out why, and what changes is the evidence available at each step.
It reads the failing responses and the requests underneath them, so it starts from the actual exchange rather than a status code. It checks what checkout said to Postgres, Redis and payments during the window the failures happened, which is where the cause usually is. It looks below HTTP at resets, retransmits, drops and DNS answers, because a 502 is frequently a connection that died rather than a handler that threw. And it arrives at whether the cause is the code or the infrastructure with the evidence attached, which is the distinction that decides who fixes it.
That last step is worth dwelling on, because it is the one engineers argue about. A dependency call that came back with an error status reached its destination and was refused on its merits, which is code. A dependency call that got a reset instead of an answer, from something that is not the service, is infrastructure sitting in the path. Both produce a 502 at the gateway and identical lines in its log. They are only distinguishable if you can see the exchange and the connection together, which is the thing an agent with kernel access has and an agent with logs does not.
Start with apiwatch, which is published, runs in one command, and answers the first question most people have about a box: what APIs does this thing even serve and call.
curl -fsSL https://yeet.cx | sh # install yeet, once
yeet run gh:yeet-src/apiwatch # clone, build, and list this box's APIs from 30s of traffic
Point it at a server you did not build and a minute later you have the list nobody wrote down: which ports answer HTTP and which process owns each one, the endpoints they serve and the status codes they return, who calls them, and which outside APIs the machine calls in turn. It comes from traffic, so an API that exists only in a config file nobody reads still shows up, and a documented route nobody calls does not.
The reason this is a good first move with an agent specifically is that it converts an unknown machine into a described one in sixty seconds, which is the context the agent was missing. --watch then alerts on the responses your real clients got, with the failing request and response in the message, rather than on an uptime checker's own synthetic requests.
Treat it as database access rather than as a dashboard, because that is closer to what it is. Request and response bodies carry authorization headers, session tokens, personal data and payment details, and a capture holds them regardless of whether an agent or a person reads it. The right questions are who can query it, how long it is kept, and what is removed before anything leaves the host.
The tools differ in what they do about this, so check rather than assume. apiwatch redacts by key name and recognisable shape, which catches passwords, tokens, API keys and card numbers, and it is explicit that this is not everything: a secret under an innocent key, a free-text body, or personal data like names and email addresses goes through as captured, and --bodies off is the switch when that matters. If alerts route through Slack, those bodies pass through yeet's servers on the way, which is worth knowing before you point it at a channel.
The other half of safety is what the tooling can do rather than what it can see. apiwatch observes: it reports what crossed the socket, and it does not block, retry or modify a request, and does not probe your APIs itself. An agent given that surface can learn anything the traffic contains and change nothing, which is a different risk profile from an agent with a shell on the same box.
Enough that you should know the list before you rely on it, because the gaps are specific rather than general.
crypto/tls has no shared library to hook, so a Go program's HTTPS calls appear by hostname from the TLS handshake with no status codes, which httpscope addresses with uprobes on Go's own TLS, and a Go server terminating its own TLS is invisible as HTTP. Statically linked and stripped OpenSSL is the same.grpc-status trailer on a 200 response, so a tool reading status codes alone does not count them as failures; grpcsnoop decodes gRPC calls and their messages.None of these are reasons to skip the approach. They are reasons to know which question you are asking: a Go service's outbound HTTPS and a gRPC error path need different tools than a plaintext HTTP endpoint, and the honest move is to say so rather than let an agent conclude from silence that nothing failed.
The reason agent-assisted debugging stalls on API failures is almost never the model. It is that the agent is asked to explain an event using artifacts that do not contain it, so it does the only thing available and produces a well-reasoned guess.
Giving it the traffic changes the shape of the work. The agent stops proposing hypotheses for you to test by deploying and starts reading what happened: the failing request, the response body, what the dependency said, whether the connection survived. That is the same investigation a good engineer does, and the reason to hand it to an agent is that it is mostly correlation, which agents are fast at and people find tedious.
It can reason about your code and your logs, which is enough for a bug that reproduces and not enough for one that only happens in production. What changes the outcome is giving it the actual traffic to read: the failing request and response, what the service asked its dependencies, and the state of the connection underneath. With yeet installed on the host, that becomes something the agent can query rather than something you have to paste in.
The exchange that failed, not a description of it. Concretely: the request body and headers that produced the error, the response the service actually returned, what its dependencies answered during the same window, and whether the connection was reset or timed out before any answer arrived. Code and logs supply none of those reliably, which is why an agent limited to them tends to propose changes rather than identify causes.
Install a capture on the host that reads HTTP at the kernel level and exposes it to the agent as a queryable surface. With yeet, apiwatch is the published starting point: one command lists every API the machine serves and calls from live traffic, and the agent reads that rather than guessing at the machine's shape. There is a ready-made prompt that has an agent set it up end to end.
It observes rather than intervenes: no proxy in the path, no TLS re-termination, no sidecar to deploy, and no modification of requests. The costs that are real are CPU on the capture path, which is why narrowing to specific ports matters on a busy database or file server, and the sensitivity of captured bodies, which should be governed like database access rather than like metrics.
The capture does, because attaching eBPF programs is privileged, but the usual arrangement keeps that privilege in a daemon rather than handing it to the agent. The agent talks to a read-only interface, so it can ask what the traffic contained without holding the capability to attach probes, change the system, or run arbitrary commands.
For TLS stacks it can hook, yes, by reading at the library's plaintext boundary, SSL_read and SSL_write, rather than decrypting the wire, which needs no private key. OpenSSL and rustls are covered this way. Go is the significant exception, because crypto/tls has no shared library to attach to, so a Go service's HTTPS shows up as a hostname from the handshake without status codes or bodies.
An APM agent runs inside your application, which is how it sees spans and internal timings, and it only covers services you instrumented in languages it supports. A kernel capture sees every process on the host, including ones nobody instrumented and ones you did not write, and it sees the connection layer underneath HTTP. They overlap less than people expect: APM is better inside your code, kernel capture is better at everything between your services.
That depends on the thresholds, and it is worth checking what a tool's defaults actually are. apiwatch alerts on an API's first 5xx, learns each API's normal 4xx share over five minutes before alerting on a jump, and treats a served port going quiet as an event. It sends one message when a problem starts, one when it recovers, and a reminder every thirty minutes, with simultaneous problems grouped into one message.
That is where this is strongest, because the agent's disadvantage on unfamiliar code disappears when the evidence is the traffic. It does not need to understand a service's internals to observe that an endpoint returns 500 when a particular field is absent, or that the failing requests all arrived on connections that had been idle for several minutes. Reading behaviour requires no prior knowledge of the implementation.
It helps if you already have it and is not a prerequisite. Tracing stitches one request across services and needs every service instrumented and propagating context. Kernel capture gives per-hop exchanges with no instrumentation, which is less complete as a narrative and immediately available on a system that was never set up for tracing, which is the common case during an unplanned incident.