
Fifteen years building engineering platforms, currently focused on advanced AI infrastructure at yeet. I love turning the deeply complex topics into something everyone can understand. I relate deeply with the core yeet philosophy that you can just build things.
Last updated: October 2026
Quick answer. "Connection reset by peer" means a TCP RST arrived mid-connection, and the standard advice, run
tcpdumpand read the RST's source address, gives you an IP rather than a cause. The question that actually ends the investigation is which local process owned that socket, and no single standard tool answers it:tcpdumpsees packets and has no idea which process they belong to, whilesssees sockets and never sees the reset. The fastest way to settle it is to watch the resets live with a TCX tap like pktscope on yeet, which shows each RST with its direction (rxortx) so you immediately know whether your host is the victim or the sender, then join the surviving socket to its process through the inode that/proc/net/tcpand the process's file descriptors share. Before any of that, check whether the RST's source address is the peer, a middlebox, or your own kernel, because those are three different bugs.
A reset is the loudest failure in TCP and the least informative. The application gets ECONNRESET, the log line says "connection reset by peer", and the peer in question may be a service you do not own, a load balancer nobody told you about, or your own kernel answering for a socket that no longer exists. The error names the symptom and hides the actor.
What makes resets hard to debug is not that the evidence is missing. The RST is a real packet with a real source address, and it is trivially visible with a packet capture. The problem is that the evidence stops one step short of the answer: you learn the connection was reset and where the packet came from, and you still do not know which of the forty processes on the host owned the socket, or why that socket existed at all.
It means your socket received a TCP segment with the RST flag set, and the kernel tore the connection down immediately rather than closing it in order. RFC 9293, the current TCP specification, defines the reset control bit and the states in which a host is expected to send one. The next read or write on that socket returns ECONNRESET, which is the error your application reports. The distinction worth holding is between RST and FIN: a FIN is the orderly half-close that lets both sides finish writing, while an RST is an abort that discards anything still in flight. Seeing an RST means somebody decided this connection should stop existing right now.
The phrase "by peer" is where most investigations take a wrong turn, because it is the C library's wording, not a statement of fact about who is at fault. It means the reset arrived from the other end of the socket as your kernel understands it. That is the address your packet reaches, which may be a load balancer, a NAT device, a service mesh sidecar, or a cloud endpoint standing in front of the service you think you are talking to. The application that actually decided to abort may be several hops behind the address in the packet.
There is also a case where the peer is innocent entirely: your own kernel sends an RST when a packet arrives for a connection it has no state for. That happens after a process exits with sockets still open, after a conntrack entry expires on a NAT in the middle, or when a packet arrives for a port nothing is listening on. The packet looks identical. The cause is completely different, and the fix is in a different system.
Join the connection to the process through the socket inode, because that is the only identifier both halves share. The kernel's TCP table, exposed as /proc/net/tcp, lists each connection with its local and remote address, its state, and the inode number of the underlying socket. Separately, every process's file descriptors record which inode each descriptor points at. The same per-process kernel accounting is what finding which process is slowing a Linux machine reads for CPU and I/O. Neither view mentions the other, but the inode is the same number in both, so the join is exact rather than heuristic.
That is the join lsof -i and ss -tanp perform for you, and it is why they are the right first commands. The ss manual states the option plainly: -p, --processes shows "processes using sockets". The reason it is worth understanding the mechanism rather than just running the tool is that the join only works while the socket still exists. A reset connection is often gone within milliseconds, and the process that owned it may have exited in the same second. By the time you run ss, the row you needed has been collected.
That timing problem is why a live view beats a point-in-time command here. pktscope on yeet streams from a TCX tap in the kernel, so the reset and the connection it belonged to are on screen when they happen rather than reconstructed afterwards:
curl -fsSL https://yeet.cx | sh # install the yeet daemon, once
yeet run gh:yeet-src/pktscope -- --iface eth0
Run that before you reproduce the failure, and the reset arrives with its direction, its peer, and the connection's full protocol tree, instead of leaving you to race ss against a socket that has already been collected.
Because a packet does not carry a process identity, and tcpdump reads packets. By the time a segment is on the wire it has an address, a port, and a set of flags, and every trace of which program called write() is gone. This is not a limitation anyone can configure around: the information is simply absent at the layer tcpdump observes, which is why its output is the same whether the sender was your application, a sidecar, or the kernel itself.
The inverse is true of ss and lsof. They read the kernel's socket tables, so they know exactly which process holds which connection, and they have no visibility into individual packets at all. Neither tool ever sees the RST. You can run both and correlate the four-tuple by eye, which works when the connection is long-lived and fails when it is not, which is the case that actually brings people to this question.
What closes the gap is observing at a layer that still has both, which means in the kernel, on the packet path, where the socket is still associated with its owner. That is what an eBPF program attached to the traffic control layer can do, and it is why the answer to "which process" and the answer to "which packet" can come from the same place rather than from two tools you join manually.
Read the RST's source address first, then decide which of the three stories fits, because they are distinguishable and the fix is different in each.
| Source of the reset | What you see on the wire | Typical cause | Where the fix is |
|---|---|---|---|
| The peer application | RST from the service address, often right after your request | The app aborted, crashed, or set SO_LINGER to 0 | The remote service |
| A middlebox or load balancer | RST from an address that is not the service, or after an idle gap | Idle timeout, connection limit, policy drop | The LB or firewall config |
| Your own kernel | RST sent by your host, for a connection it has no state for | Process exited with sockets open, conntrack expiry, nothing listening | Your host or the NAT in between |
| A conntrack or NAT device | RST after a precise, repeatable idle interval | State table entry expired mid-connection | Keepalives, or the NAT timeout |
The signal that separates the first two is timing. An application-level abort tends to follow a request closely, because something in the request handling failed. A middlebox timeout tends to follow a quiet period of a consistent length, and the consistency is the clue: resets that arrive after almost exactly the same idle duration every time are a timer somewhere, not a decision about your traffic.
The signal that separates the third is direction. If the RST is leaving your host rather than arriving at it, you are not the victim, you are the sender, and the question changes from "who is resetting me" to "why does my kernel have no state for this connection". That usually means the owning process is gone, which brings the investigation back to the process attribution above.
Attach to the packet path in the kernel and filter for the RST flag, which gives you the reset as it happens rather than a capture file you analyze afterward. The flag is a single bit in the TCP header, 0x04, so filtering for it is cheap and precise, and watching live matters here because the connections you care about are short-lived by definition.
pktscope is a packet analyzer built as a yeet script that does this with a TCX (clsact) eBPF tap, which is the traffic control hook rather than a raw socket, so it sees frames on the interface without putting anything in the path. Its filter grammar takes TCP flags directly:
yeet run gh:yeet-src/pktscope -- --iface eth0 # the tap, on one interface
# then, in the filter bar:
rst # only resets
rst host 10.0.0.7 # resets involving one peer
rst !ack # resets that are not the ack-carrying kind
What you get back is the reset in context: the frame, the Ethernet, IP and TCP layers as a folding tree with every field carrying the byte range it decodes, the direction (rx or tx, which answers whether you sent it), and the payload if there was one. The direction column is the one to read first, because it settles whether your host is the victim or the source before you look at anything else.
The honest limit is that this sees your host's interfaces. A reset generated three hops away by a device you do not run arrives looking exactly like one from the peer, and no amount of local observation will name that device. What local observation does do is tell you definitively that your side did not send it, which is usually the fact in dispute.
ss -tan show me enough to find the TCP reset?It shows you the surviving connections and their states, which is useful context and is never the reset itself. ss -tan lists what exists right now: established connections, sockets in TIME-WAIT, listeners. A connection that was reset is, by definition, no longer in that table, so the row you most want is the one guaranteed to be missing by the time you look.
Where it earns its place is in the states around the failure. A large population of sockets in CLOSE-WAIT says your application is not closing connections the remote side already finished with, which eventually produces resets when limits are hit. Sockets accumulating in TIME-WAIT is normal and usually not the problem people assume. And ss -tanp adds the process, which is the inode join described above, done for you.
So the sequence that works is: use ss to understand the steady state and find the owning process while the socket still exists, and use a live packet view to catch the reset itself. Treating either one as the whole answer is what makes these investigations run long.
Laid out by what each tool can and cannot answer, the division is structural rather than a matter of preference:
| What you reach for | Sees the RST packet | Shows direction | Survives a short-lived socket | What it is for |
|---|---|---|---|---|
| TCX tap (pktscope) | Yes, live, as it happens | Yes, rx or tx | Yes, it is streaming | Watching resets as they happen |
tcpdump | Yes | Yes, from the addresses | Yes, if capture was already running | Capturing the packet for later analysis |
ss -tanp | No, it dumps socket statistics | Not applicable | No, the row is gone | Steady state and the owning process |
lsof -i | No | Not applicable | No | The same join, per open file |
conntrack -L | No, it tracks flow state | No | Partly, until the entry expires | Whether a NAT or firewall expired the flow |
The split down the middle is the thing to see: the tools that show you the reset never show you a socket, and the tools that show you a socket never see the reset. That is why the investigation stalls on correlation rather than on evidence, and why a live view matters more than it first appears, since the socket half of the answer expires in milliseconds while the packet half does not.
The reason connection resets stay unresolved is that the standard advice stops at the point where the evidence is easiest to collect and least useful. Reading the RST's source address takes thirty seconds and tells you an address. Knowing which process on your host owned the socket, and whether the reset arrived or departed, is what turns a reset into a bug with an owner.
Three facts close most of these: the direction of the RST, which says whether you are the victim or the sender; the source address, which separates the peer from a middlebox; and the owning process, which is the inode join that tcpdump cannot do and ss cannot do in time. Collect those three and the remaining question is usually a configuration value in a system you already suspected.
A TCP segment with the RST flag arrived on an established connection. The common causes are the remote application aborting or crashing, a load balancer or firewall closing an idle connection, a NAT or conntrack entry expiring mid-connection, or your own kernel resetting a connection it no longer has state for. The error message names the socket's other end, which is not necessarily the component that made the decision.
Join the connection to the process through the socket inode, which appears both in the kernel's TCP table and in each process's file descriptor list. ss -tanp and lsof -i perform that join for you. The catch is timing: a reset connection disappears from the table quickly, so the join only works if the socket still exists when you look, which is why reading the kernel view continuously catches cases a manual run misses.
No. A packet carries addresses, ports and flags, with no record of which process produced it, so tcpdump cannot report a pid regardless of options. To connect packets to processes you need to observe where both are still associated, which means in the kernel rather than on the wire, using something like an eBPF program on the traffic control path.
Read the direction of the packet. If the reset is leaving your host, your kernel sent it, usually because it received a segment for a connection it has no state for, which typically means the owning process exited with sockets still open. A live packet view that labels direction (rx versus tx) answers this immediately; a capture file requires you to compare the source address against your own interfaces.
That consistency is the signature of a timer in a middlebox rather than a decision by the peer application. Load balancers, NAT devices and conntrack tables all expire idle entries on a fixed interval, and once the entry is gone, the next packet on that connection produces a reset. Enabling TCP keepalives below the timeout interval, or raising the timeout, is the usual fix. Worth knowing before you rely on it: tcp(7) sets tcp_keepalive_time to 7200 seconds by default, two hours, which is far longer than almost any load balancer idle timeout, so the default does nothing for this and the value has to be lowered deliberately.
No. Some applications close connections with an RST deliberately, by setting SO_LINGER with a timeout of zero, because it avoids the TIME-WAIT state and frees resources immediately. Load balancers and proxies sometimes do the same when shedding load. A reset is an abrupt termination, not necessarily a fault, which is why the question that matters is which component sent it and whether that component was supposed to.
A FIN is an orderly close: each side sends one when it has finished writing, in-flight data is delivered, and the connection winds down through a defined sequence of states. An RST is an abort: it terminates the connection immediately and discards anything still in flight. An application that sees ECONNRESET received an RST, while one that sees a clean end-of-stream received a FIN.
Yes, by attaching an eBPF program to the kernel's traffic control path and filtering on the RST flag, which is bit 0x04 of the TCP header. pktscope does this as a yeet script using a TCX tap, so you get live resets with direction and a decoded protocol tree rather than a capture file to analyze afterward, and nothing is inserted into the packet path.
Yes, and that is the most common reason the error misleads people. Your socket's remote address is whatever your packets reach, so if a load balancer terminates the connection, it is the peer as far as your kernel is concerned, and a reset it sends is reported as "reset by peer". The service behind it may be perfectly healthy and entirely unaware.
A reset that follows request handling closely usually indicates the remote application aborted rather than a timer firing, because timers produce resets after idle gaps instead. Common causes are the remote process crashing or being killed mid-response, an unhandled error path that closes the socket abruptly, or a deliberate SO_LINGER zero close. The timing relative to your request is the clue that separates this from a middlebox timeout.
-p, --processes for showing the processes using a socket. https://man7.org/linux/man-pages/man8/ss.8.htmlSO_LINGER and the deliberate abort-on-close behaviour it enables. https://man7.org/linux/man-pages/man7/socket.7.htmltcp_keepalive_time and its 7200 second default. https://man7.org/linux/man-pages/man7/tcp.7.html/proc/net/tcp and the socket inode that makes the connection-to-process join possible. https://man7.org/linux/man-pages/man5/proc.5.html