Prove a Network Problem Isn't on Our Side

Necco Ceresani
Necco Ceresani
·18 min read

Fifteen years building engineering platforms, currently focused on advanced AI infrastructure at yeet. I love turning the deeply complex topics into something everyone can understand. I relate deeply with the core yeet philosophy that you can just build things.

Last updated: October 2026

Quick answer. The deadlock in a cross-team network dispute is that both sides are reading their own application metrics, which describe outcomes rather than the exchange, so neither can disprove the other. What settles it is the packet record: which side sent the first FIN or RST, the timestamp it happened, the direction of each segment, and the bytes the remote endpoint actually returned. Those are facts about the wire rather than interpretations of a dashboard, and on Linux you can collect them with a TCX eBPF tap like pktscope on yeet, without a sidecar, a proxy, or a change to the service. Lead with the direction of the first abort, because that single field decides which side acted and is the hardest to argue with.

There is a specific shape to these incidents. Your service reports errors against a dependency. The team that owns the dependency reports healthy. Both statements are true, because each team is measuring a different thing: you are measuring your requests failing, they are measuring their handlers succeeding, and the failure lives in between where neither has instrumentation.

What breaks the deadlock is evidence that does not come from either side's application. The network exchange is the one record both parties can read the same way, and it answers the questions that dashboards cannot: not whether the service was healthy, but what it sent, when, and who stopped talking first.

What counts as proof in a network dispute: a TCP RST, not a dashboard?

Facts about the exchange rather than conclusions from your monitoring. The distinction matters because the other team's monitoring is as credible as yours, and two dashboards in disagreement produce a longer meeting rather than a resolution. Evidence that travels across an organizational boundary has to be checkable by someone who does not trust your tooling.

Four things qualify. Strongest is which side sent the first FIN or RST, because that is a single field in a single packet rather than an inference, and it answers who ended the connection. Next is the timestamp of that packet, which anchors the event to a moment the other team can find in their own logs and converts "around 14:30" into a specific second. Then the direction of each segment, since showing that your host sent a complete request and received nothing, versus received a response and rejected it, separates two entirely different faults. And finally the bytes the remote returned, because a response body or status line is the endpoint speaking for itself rather than your client's interpretation of it.

What does not qualify, however true: your error rate, your p99, your retry count, or a screenshot of a dashboard. Those are outcomes. They establish that you were harmed, which is already agreed, and not what caused it, which is the dispute.

Sorted by how well each one survives contact with a team that does not trust your tooling:

EvidenceWhat it establishesCan the other team dispute itWhere it comes from
Direction of the first RST or FINWhich side ended the connectionNo, it is a control bit with a directionA packet view on your host
Packet timestamp, UTCThe exact moment, matchable to their logsOnly if clocks disagree, so sync NTPThe capture, plus a synced host
The request bytes you sentThat your side sent a complete, well-formed requestRarely, the bytes speak for themselvesThe payload pane
The response bytes returnedWhat their endpoint actually answeredNo, it is their output quoted backThe payload pane
Your error rate or p99That you were affectedYes, trivially, it says nothing about causeYour dashboard
"Our service is healthy"Nothing the other team can act onYes, they will say the sameEither side's monitoring

The top four rows share a property the bottom two lack: they are facts about the exchange rather than measurements of one endpoint. That is the whole reason packet evidence ends these arguments and dashboards extend them.

How do I show which side sent the first TCP FIN or RST on Linux?

Capture the connection's final segments and read the direction of the first FIN or RST. TCP makes this unambiguous by design: an orderly close is a FIN from whichever side finished writing, and an abort is an RST from whichever side gave up. Both are control bits defined in RFC 9293, so their meaning is not a matter of interpretation between teams. The packet carries a direction, so "who stopped first" is a fact you read rather than a theory you argue.

A tap on the traffic control hook sees this without being in the path. pktscope filters on TCP flags and labels direction as rx or tx, which is exactly the pair of fields this question needs:

yeet run gh:yeet-src/pktscope -- --iface eth0   # a TCX tap on the interface
# then, in the filter bar:
fin host 203.0.113.10                 # orderly closes with one peer
rst host 203.0.113.10                 # aborts with one peer
tx rst                                # resets that WE sent

That last filter is the one to run first, and it is worth running before you escalate. If tx rst returns your own host aborting connections, the dispute is over and the answer is uncomfortable: your side acted. Finding that out privately, in thirty seconds, is considerably better than finding it out in a shared incident channel after an hour.

When the filter comes back empty and the resets are all inbound, you have the finding that actually settles the escalation: your host did not end these connections, and here is the packet that did, with its source address and timestamp.

What did the third-party endpoint send back over the network?

Read the response off the wire rather than through your client's error handling. This matters because client libraries compress a wide range of remote behavior into a small number of exceptions, and the compression discards exactly the detail a vendor needs to act. A timeout in your logs could be a connection that was never answered, a connection answered with a 503, or a response that arrived with a body your parser rejected. Those are three different tickets.

The wire view keeps the distinction. A packet analyzer that decodes HTTP shows you the status line and headers the endpoint returned, and the payload pane shows the body bytes as they arrived. When the remote is returning a 429 with a Retry-After header, or a 200 with a truncated body, or nothing at all, each looks different on the wire and identical in a log line that says "upstream request failed".

For a vendor conversation, this is the difference between "your API is failing" and "at 14:32:07 UTC your endpoint returned HTTP 503 with this body, to this request, and our host had already sent the complete request 40 milliseconds earlier". The second version can be acted on by someone who does not work with you and does not have your dashboard.

What makes this practical is that the decoding happens for you rather than in your head. A packet analyser that understands HTTP/1 shows the request line, the status line and the headers as fields, with the raw bytes alongside them, so quoting the endpoint accurately is reading a pane rather than reconstructing a conversation from hex. The headers are frequently where the answer sits: a Retry-After that your client ignored, a Content-Length that disagrees with the body that arrived, or an upstream identifier that tells the vendor which of their nodes served you.

The limit to hold onto is TLS. Once a connection is encrypted, the request and response bodies are not readable from the wire without key material, so for HTTPS endpoints this technique gives you the connection, the handshake and the timing rather than the payload. That is often enough, because the dispute is usually about whether an answer came at all and when, but it is worth knowing before you promise a vendor a quotation of their response body.

Was it our client timeout or a slow HTTP response from their endpoint?

Measure the gap between your last request byte and their first response byte, because that interval is the one fact that separates the two stories. A client timeout fires on your clock, so "we timed out" only establishes that you stopped waiting. Whether the remote was slow, silent, or answering promptly into a connection that was already gone is a different claim, and the packets carry it.

Three shapes show up, and they are distinguishable at a glance once you are looking at both directions:

  • Request sent, nothing back at all. No response segments before your side gave up. This is the strongest case for escalation, because your host did everything and received no answer.
  • Request sent, response arrived after your deadline. Response segments exist but land after your client stopped waiting. Here the dispute is about their latency against your timeout setting, and both numbers are now visible rather than asserted.
  • Request sent, response arrived in time, client rejected it. The bytes came back promptly and your side still failed. That is almost always your bug, and finding it privately is the point of looking.

Filtering to one peer and reading both directions is what makes the distinction, so the filter to reach for pairs a host with direction:

host 203.0.113.10 tx       # what we sent them
host 203.0.113.10 rx       # what came back, and when

The third case is the uncomfortable one and it is worth checking first, for the same reason tx rst is worth checking first: a response that arrived and was discarded by your own parsing or deadline logic will otherwise be escalated as a vendor problem, and the vendor will correctly reject it.

How do I capture network traffic with eBPF, without a sidecar or proxy?

Observe at the kernel's traffic control hook, which sees frames on the interface without being inserted into the path. This is the property that makes the evidence credible and the collection safe: nothing is routed differently, no TLS is terminated and re-originated, and the service under investigation is not restarted or reconfigured.

That last point matters more than it first appears. The standard alternatives all change the thing being measured. A proxy inserted for debugging, mitmproxy or similar, substitutes its own TCP and TLS stacks for your client's, so a failure caused by your client's behavior can disappear the moment you start looking. A sidecar requires a deploy, which restarts the process and resets the state that produced the failure. A restart with more logging does the same.

A TCX tap avoids that because it is an observer rather than a participant. pktscope streams full frames up to 1536 bytes into a ring buffer and decodes them into a protocol tree where every field carries the byte range it came from, so the evidence you present can be traced back to actual bytes rather than to a tool's summary of them.

One practical limitation to plan around: this is a live analyzer rather than a capture-file workflow, and it does not currently write a pcap you can attach to a ticket. In practice that means capturing the finding while it is on screen, by recording the decoded fields, the timestamps and the hex, rather than handing over a file for the vendor to open in Wireshark. If your dispute process requires a pcap artifact, pair the live view with tcpdump -w for the archive, readable afterwards in Wireshark, and use the live view for the diagnosis.

Can I prove our host sent a complete request before the RST?

Partly, and it is worth being precise about which part, because overclaiming here is how evidence gets dismissed. You can prove your host sent what it was supposed to send, when it sent it, and that it did not abort the connection. That is a strong, narrow claim grounded in packets.

What a host-local view does not prove is that the path between you and the remote was healthy. A packet that left your interface can be dropped by any device between the two endpoints, and from your host that is indistinguishable from a remote that received the request and chose not to answer. Both look like a request sent with no response returned.

What strengthens the claim is staying inside what the packets actually show: the request left your interface complete and well formed, at a timestamp you can name, and the connection was not terminated from your end. Each of those is readable in the same capture, and none of them requires the other team to trust your monitoring.

The honest framing to use in the escalation is "our host sent a complete request at this timestamp and received no response; we did not close the connection" rather than "the network is fine and it is your problem". The first is defensible with what you have. The second invites a counterargument you cannot answer from one side of the path.

What if the TCP evidence shows the network problem is ours?

Then you have found your bug in minutes rather than after an escalation, which is the more common outcome and the more valuable one. It is worth saying plainly because the framing of this question, proving it is not us, quietly assumes the answer. Running the checks honestly means accepting the result in either direction, and the direction that points inward is the one you can actually fix today.

The three findings that point back at your own host are specific and easy to read. Your side sent the first RST, which means your application or your kernel ended the connection. A response arrived in good time and your client rejected it, which is a parsing, deadline or validation bug. Or the request that left your interface was malformed or incomplete, which is almost always a serialisation or truncation problem above the socket.

None of those is visible in a dashboard that reports your error rate, which is why teams escalate first and discover the internal cause later. Reading the exchange before you open the ticket inverts that order. The cost of checking is one filter; the cost of not checking is an escalation that ends with the other team showing you your own packets.

The bottom line: show who sent the first RST, not who is blamed

These disputes stall because both teams arrive with outcome metrics and no shared view of the exchange. The way through is to stop arguing about whose service was healthy and start reading what happened on the wire, where the facts are the same for everyone.

Run the one-line check first: whether your host sent the first abort. It resolves the question in either direction, and it resolves it privately. If the answer is that you did not, you now have a packet with a direction, a source address and a timestamp, and that is evidence a stranger can verify. If the answer is that you did, you have saved everyone an escalation and found your bug.

Frequently asked questions

How do I prove a network issue is not caused by my application?

Collect evidence about the exchange rather than about your service. Show that your host sent a complete, well-formed request, record the timestamp it left the interface, show that your side did not send the first FIN or RST, and capture what the remote returned. Those are packet-level facts that another team can check, unlike error rates and latency percentiles, which only establish that you were affected.

How do I tell which side closed a TCP connection first?

Read the direction of the first FIN or RST in the connection's final segments. A FIN is an orderly close from the side that finished writing; an RST is an abort. Because each packet has a direction, filtering for tx rst versus rx rst on your host answers whether you or the peer ended it, with no inference involved.

What evidence should I give a vendor for a network problem?

A specific timestamp in UTC, the request your host sent, the response or lack of one that came back, and which side terminated the connection. Quoting the status line and headers the endpoint returned is far more useful than reporting your client's exception, because client libraries collapse many distinct remote behaviors into one generic error.

Can I capture traffic without deploying a proxy or sidecar?

Yes. An eBPF program on the kernel's traffic control hook observes frames on the interface without being in the path, so nothing is re-routed, no TLS is terminated, and the service is not restarted. That matters for disputes because a proxy substitutes its own network stack for your client's and can make the failure you are investigating disappear.

Why do both teams' dashboards show everything is fine?

Because each team is measuring its own endpoints rather than the exchange between them. Your dashboard shows requests failing; theirs shows handlers succeeding. Both can be accurate at once when the failure occurs in the path, at a load balancer, a NAT, or a firewall that neither team's application instrumentation can see. The packet record is the only view that covers the space in between.

Does a timeout mean the remote server never responded?

Not necessarily. A client timeout means no usable response arrived in time, which covers a request that was never answered, a response that arrived too late, and a response that arrived but failed to parse. These look identical in application logs and completely different on the wire, which is why reading the actual bytes changes what the ticket says.

Can I prove the network path was healthy from my host?

No, and claiming it weakens your case. From one end you can show that a packet left your interface and that nothing came back, which does not distinguish a drop in the path from a remote that declined to answer. The defensible claim is about your own behavior: what you sent, when, and that you did not terminate the connection.

What is the difference between a FIN and an RST in a dispute?

A FIN indicates an orderly close by a side that finished writing, which is usually normal and sometimes premature. An RST is an abort that discards in-flight data, and it is a stronger signal that something went wrong. In an escalation, an inbound RST with a timestamp is the most direct evidence that the other end stopped the conversation.

How precise do timestamps need to be for a cross-team investigation?

Precise enough to match a single request, which in practice means sub-second and expressed in UTC, which assumes the hosts involved agree on the time, so NTP synchronisation is a precondition for this evidence rather than a detail. The purpose is to let the other team find the same event in their logs, so a timestamp to the second, paired with the source port of the connection, usually identifies the exact transaction on both sides without ambiguity.

Should I use tcpdump or a live analyzer for this?

Both, for different jobs. A live analyzer with protocol decoding is faster for diagnosis, because you read the decoded fields and direction as it happens rather than collecting a file and analyzing it afterward. tcpdump -w remains useful when your dispute process requires a pcap artifact to attach to a ticket, since a live view is not a file you can hand over.

Sources

  1. RFC 9293, the current TCP specification, defining the FIN and RST control bits whose direction settles which side ended a connection. https://www.rfc-editor.org/rfc/rfc9293.html
  2. RFC 9110, HTTP semantics, for the status codes a third-party endpoint returns and the Retry-After header. https://www.rfc-editor.org/rfc/rfc9110.html
  3. RFC 5905, the Network Time Protocol, since cross-team timestamp evidence depends on synchronised clocks. https://www.rfc-editor.org/rfc/rfc5905.html
  4. tcpdump, for writing a capture file when a dispute process requires an artifact to attach. https://www.tcpdump.org/manpages/tcpdump.1.html
  5. Linux kernel BPF program types, for the traffic control attach point that observes without entering the path. https://docs.kernel.org/bpf/libbpf/program_types.html
  6. pktscope, a terminal packet analyzer built as a yeet script on a TCX (clsact) eBPF tap, with TCP flag and direction filters and a decoded protocol tree. https://github.com/yeet-src/pktscope
  7. yeet documentation, the runtime the packet tap runs on. https://yeet.cx/docs/

Related resources

  1. Who is resetting my TCP connections on Linux, attributing a reset to the process that owned the socket. https://yeet.cx/topical-takes/who-is-resetting-my-tcp-connections
  2. Why is my TLS handshake failing on Linux, reading a negotiation that fails before your application logs it. https://yeet.cx/topical-takes/why-is-my-tls-handshake-failing
  3. What the TC layer sees, on observing traffic at the traffic control hook without a sidecar or a proxy. https://yeet.cx/topical-takes/what-the-tc-layer-sees