You're Not Crazy, They Never Sent FIN

Jacob Pradels
Jacob Pradels
·8 min read

Founding engineer at yeet, working on kernel-side observability and the tooling around it. I write about eBPF, Linux internals, and why your telemetry bill looks the way it does.

There's a paragraph in the AWS docs, under the Network Load Balancer, under "Connection idle timeout", that goes like this:

If no data is sent through the connection by either the client or target for longer than the idle timeout, the connection is no longer tracked. If a client or target sends data after the idle timeout period elapses, the client receives a TCP RST packet to indicate that the connection is no longer valid.

Three hundred and fifty seconds, by default. The NAT gateway page says the same thing and adds, in parentheses, the part that matters: it does not send a FIN packet.

You will find this paragraph on roughly your fourth month of chasing a 502 that happens a few times a night, never during the day, never while you're watching, from a service whose logs are spotless. You will find it because someone on a forum found it in 2019, and you'll read it twice, and you'll say "oh" out loud. And as fun and easy as it is to search AWS documentation for a sentence you don't know you're looking for, it would've been nicer to just catch the thing when it happened.

So that's this post. What the paragraph means, how to reproduce it on one box in fifteen seconds instead of four months, and the 110 lines of eBPF that make the kernel tell you who reset you and how long you'd been idle when it happened.

12:27:52 RST from 127.0.0.1:8091 -> python[220588] 127.0.0.1:58330  idle 355.0s before the write that got reset

A quick refresher on TCP

You know SYN and ACK. Meet FIN and RST.

Everyone has seen the handshake. Your process calls connect, the kernel sends a SYN, the other side answers SYN-ACK, you ACK, and now there's a connection: two sockets, one on each host, agreeing on sequence numbers. One round trip before you can send a byte. Put TLS on top and it's two more round trips plus some math. That's the cost a connection pool exists to skip, and the reason every HTTP client worth using keeps one.

checkout and payments open a connection with SYN, SYN-ACK and ACK, then send one request on it

The other two flags are less discussed but are important nonetheless.

FIN is the polite close. It means "I have nothing more to send." Each side sends its own FIN when it's done and the other side acknowledges it. When your kernel receives a FIN it marks the socket, and the next read on it returns end-of-stream.

payments closes an idle connection with a FIN, checkout's kernel marks the socket, and the pool dials a fresh connection for the next request

RST is the rude close. It's what your high school girlfriend did when she hung up without saying bye. It means "there is no such connection here, stop." A kernel sends it when a packet arrives for a socket it doesn't have, or when a program closes with data still unread. Nothing is negotiated. Your next read or write fails with connection reset by peer. You probably shouldn't have told her you thought that girl at the party was cute.

Now the part that matters. A socket that has received neither looks exactly like a healthy socket. There is no flag for "the thing in the middle forgot you." If a load balancer drops the connection without sending a FIN, your kernel has nothing to mark, your pool's liveness check has nothing to read, and the first write after the gap is the first anyone hears about it, as an RST from a box that no longer knows your name.

Ok so what is a connection pool actually doing

The connection pool exists to skip the handshake. For a service that calls payments two hundred times a second, that's the difference between a fast service and a service that is mostly SYNs. So when a request finishes, the connection goes back on a shelf, and the next request takes it off the shelf instead of dialing.

The shelf has a timer. requests keeps a connection for as long as the server says it will, urllib3 by default forever, Go's transport for 90 seconds, your Java client for whatever somebody set in 2019. The thing on the shelf is a socket the kernel still believes is open.

The question nobody asks is whether anyone else on the path agrees.

Everyone on the path has their own timer

A TCP connection between checkout and payments is a pair of sockets, one on each end. But the packets go through things, and the things keep state:

In the middleGives up afterWhat it sends when it does
nginx, HAProxy, Envoykeep-alive timeout, 75s by default in nginxFIN
AWS Application Load Balancer60sFIN
AWS Network Load Balancer350snothing, RST on your next write
AWS NAT gateway350snothing, RST on your next write
Azure Load Balancer4 minutesnothing, unless you turn on "TCP reset on idle"
A stateful firewall or conntrack tablewhatever it shipped withusually nothing

Each one forgets the connection after its own timeout. The polite ones send a FIN and the kernel marks the socket so that the next time the pool goes to reuse it, the library notices and dials fresh. Each lib handles this slightly differently, urllib3 does a zero-timeout readability check before every reuse for exactly this reason. Go keeps a goroutine reading each idle connection. But generally, the polite close is handled.

The impolite ones send nothing. The gateway doesn't close the connection, it just forgets about it. As a result, until the socket is used again all anyone can do is assume it's healthy. The pool's readability check sees nothing to read, because there is nothing to read. Then you write, and the gateway, which has no idea who you are anymore, answers with a reset.

checkout and a load balancer exchange one request, sit idle for 350 seconds, the load balancer forgets the flow without a FIN, and the next write gets an RST that checkout logs as a 502 while payments logs nothing

So the bug has a signature. It only happens on the first request after a gap longer than the middlebox's timeout. When the traffic is low and everything seems fine. It looks random, and a retry fixes it, because the retry dials a new connection. Which is how it lives in production for years with a retry wrapped around it and a comment that says # payments is flaky.

Let's take this puppy for a spin

I don't have four months, so I built the smallest thing that reproduces the bug on purpose. One box running Fedora 42 with two Python services and nginx in the middle playing the load balancer:

The test setup: checkout with a requests pool, nginx on 8091 standing in for the load balancer, payments on 8092, an nftables rule that drops nginx's FIN toward checkout, and rstwatch watching the kernel

payments answers POST /v1/charges. checkout uses a plain requests session, makes a call, sleeps 15 seconds, makes another. nginx's keep-alive is set to 10 seconds so I don't have to scroll X waiting for the bug to happen. The nftables rule and rstwatch come in a minute.

First, with nginx being polite:

$ .venv/bin/python checkout.py 15 2
10:56:40 checkout: POST /v1/charges -> 200 {"charge_id": "ch_1791568600", "status": "ok"}
         (idle for 15s)
10:56:55 checkout: POST /v1/charges -> 200 {"charge_id": "ch_1791568615", "status": "ok"}

Two 200s. nginx closed the idle connection at 10 seconds and sent a FIN, urllib3 saw it, opened a new one, nobody noticed. This is the case the happy case.

Now I make nginx into a NAT gateway. One nftables rule, which drops the FIN nginx sends on its client-facing port:

nft add rule inet lbdemo output tcp sport 8091 'tcp flags & fin == fin' drop

nginx still closes its side on schedule. The client just never hears about it, which is the documented behavior we see on AWS:

$ .venv/bin/python checkout.py 15 2
11:07:25 checkout: POST /v1/charges -> 200 {"charge_id": "ch_1791569245", "status": "ok"}
         (idle for 15s)
11:07:40 checkout: POST /v1/charges -> 502 upstream failure: ('Connection aborted.', RemoteDisconnected('Remote end closed connection without response'))

Payments' log has one request in it. If there were a proxy in front of checkout it would have logged a real 502 with upstream connect error or disconnect/reset before headers, which is the string you have grepped for at 3am and learned nothing from.

What the kernel knew the whole time

Nothing arrived during the idle gap, so the kernel had nothing to know. Then checkout wrote, the RST came back, and for one moment the kernel on the checkout host held the whole story: which socket the reset was for, which process owned it, and when it last carried data. It used the first fact to set an error on the socket and threw the rest away.

Linux has a tracepoint, tcp:tcp_receive_reset, that fires every time a socket receives an RST, with the socket's address pair and a pointer to the socket itself. For "how long idle" I hook tcp_sendmsg and tcp_recvmsg and keep, per socket, the last two moments it carried data. The reason to keep two is because the write that provokes the reset is itself activity. The gap I want is between that write and whatever came before it.

struct owner {
	__u64 last_ns;
	__u64 prev_ns;
	__u32 pid;
	char comm[16];
};

static __always_inline void touch(struct sock *sk)
{
	__u64 key = (__u64)sk;
	__u64 now = bpf_ktime_get_ns();
	struct owner *o = bpf_map_lookup_elem(&owners, &key);
	if (o) {
		o->prev_ns = o->last_ns;
		o->last_ns = now;
		return;
	}
	struct owner fresh = {};
	fresh.last_ns = now;
	fresh.pid = bpf_get_current_pid_tgid() >> 32;
	bpf_get_current_comm(fresh.comm, sizeof(fresh.comm));
	bpf_map_update_elem(&owners, &key, &fresh, BPF_ANY);
}

SEC("fentry/tcp_sendmsg")
int BPF_PROG(on_sendmsg, struct sock *sk, struct msghdr *msg, size_t size)
{
	touch(sk);
	return 0;
}

And the reset itself:

SEC("tracepoint/tcp/tcp_receive_reset")
int on_receive_reset(struct trace_event_raw_tcp_event_sk *ctx)
{
	__u64 key = (__u64)ctx->skaddr;
	struct owner *o = bpf_map_lookup_elem(&owners, &key);
	struct rst_event *e = bpf_ringbuf_reserve(&events, sizeof(*e), 0);
	if (!e)
		return 0;
	*e = (struct rst_event){};
	e->family = ctx->family;
	e->lport = ctx->sport;
	e->rport = ctx->dport;
	bpf_probe_read_kernel(e->laddr, 4, ctx->saddr);
	bpf_probe_read_kernel(e->raddr, 4, ctx->daddr);
	if (o) {
		e->known = 1;
		e->pid = o->pid;
		__builtin_memcpy(e->comm, o->comm, sizeof(e->comm));
		if (o->prev_ns)
			e->idle_ns = o->last_ns - o->prev_ns;
	}
	bpf_ringbuf_submit(e, 0);
	return 0;
}

That's most of the C. The JavaScript side is a ring buffer subscription and a console.log, which in most eBPF tooling is the part that takes forty files of Go. yeet is a JavaScript runtime that sits on top of the Linux kernel, one daemon per machine, and it loads the object, attaches the hooks and hands the events to a script. Run it next to the second experiment:

$ yeet run . -- --port 8091
11:07:22 rstwatch: watching for TCP resets on port 8091 (ctrl-c to stop)
11:07:40 RST from 127.0.0.1:8091 -> python[192121] 127.0.0.1:38866  idle 15.0s before the write that got reset
11:07:40 RST from 127.0.0.1:38866 -> 127.0.0.1:8091  owner unknown (no data seen on this socket)

The first line tells the whole story. The reset came from port 8091, the load balancer. It landed on a socket owned by python[192121], checkout. The connection had been idle for 15.0 seconds before the write that got reset, longer than the 10 seconds nginx was told to wait. Code, infrastructure and the number, from the host where it happened, with nothing installed in either service.

The second line is the echo. After the reset, checkout tears its socket down, and nginx's half-closed socket on 8091 gets that reset back. It has no owner because nginx already closed it. If you ever wanted proof that "both sides blame each other" is a protocol feature, there it is.

Or skip the C and have your agent build it. Paste this into Claude Code, Codex or whatever you run on a Linux box, and it writes the tracer, proves it on a real reset, leaves it running as a service you can curl, and puts the reset count on a /metrics route for Prometheus so the wall at 350 shows up in Grafana:

Build a tool on this host that reports every TCP reset it receives, with who
sent it, which process owned the socket, and how long the connection sat idle
before the write that got reset. Build it as a yeet script and keep it running
as a yeet service.

Setup, if yeet is not already on this host. yeet runs on Linux only; on
anything else, stop and tell me.

  curl -fsSL https://yeet.cx | sh
  yeet status -w

Read https://yeet.cx/docs/raw/scripts/ebpf.md, cli/services.md and
scripts/yeet-telemetry.md, and skim https://github.com/yeet-src/tcpsnoop for
the API. Scaffold with `yeet new rstwatch` and build with `make`. No terminal
UI.

What to build:

1. fentry hooks on tcp_sendmsg and tcp_recvmsg that keep, per socket, the
   owning pid and comm and the last two times it carried data, in an LRU hash
   keyed by the socket pointer. Delete the entry in tcp_close.
2. The tcp:tcp_receive_reset tracepoint: look the socket up and send the
   address pair, the owner, and last_ns minus prev_ns to a ring buffer. Read
   the address arrays with bpf_probe_read_kernel, not memcpy; the verifier
   rejects pointer arithmetic on the tracepoint context.
3. One line per reset to the console: time, sender address, owner comm and
   pid, local address, and the idle gap in seconds. Add a --port filter.
4. A yeet:telemetry registry in a shared worker with one counter,
   tcp_resets_received_total, labeled by sender address and owner comm, and a
   histogram of the idle gap in seconds. The watcher keeps the worker alive
   with keep(); a separate lazy scrape script renders the worker's snapshot
   and exits.
5. A yeet service with the daemon as the server: an isolate unit for the
   watcher, the lazy scrape unit, and a web-server unit on 0.0.0.0 on a free
   port above 1024 that you pick and report. Mount the watcher's console at
   /events, which serves a chunked text/plain stream, one line per reset,
   that stays open until the client leaves. Mount the scrape unit at /metrics
   with tenancy per-connection so each scrape runs it once.

Verify without sudo, before the service: a Python server that closes with
SO_LINGER set to 0 sends a real RST to its client. Run `yeet run .`, run
that, and the line must name the client process. Then start the service and
confirm `curl -sN http://127.0.0.1:<port>/events` streams, and
`curl -s http://127.0.0.1:<port>/metrics` returns the counter. Never run yeet
with sudo, and do not send resets to my services to make it fire. A 403 from
a route means the host is not signed in; `yeet login` fixes that.

End with: the `yeet service start` command, the two URLs with host and port,
the curl for /events, a Prometheus scrape_config for /metrics, the Grafana
query `rate(tcp_resets_received_total[5m])` by sender, and how to stop it.

The end

Here's the thing. There are a thousand paragraphs like the one at the top of this post. The NLB forgets you at 350 seconds. Node statically links its own OpenSSL. Linux sends an RST, not a FIN, when you close a socket with unread data. HTTP/1.1 keep-alive and a proxy that downgrades to HTTP/1.0. Every protocol, every library and every box in the middle has a handful of these, and they are all documented somewhere, by someone, in a sentence you will only recognize after you already know the answer.

You cannot know them all. Nobody can. What you can do is have a way to see what actually happened on the host when it happened, instead of reconstructing it from four logs that each saw a different piece. The reset was right there in the kernel, with all of the relevant info included. The only thing missing was the infrastructure to ask the question.

That's what yeet is for, and it's why we built the API debugging use case around exactly this question: is it the code or the infrastructure. Your agent reads the kernel on the host where it failed and says which, with the evidence. Next time it's 3am and payments is "flaky," you're not crazy. They never sent FIN.