Why Is My TLS Handshake Failing?

Necco Ceresani
Necco Ceresani
·16 min read

Fifteen years building engineering platforms, currently focused on advanced AI infrastructure at yeet. I love turning the deeply complex topics into something everyone can understand. I relate deeply with the core yeet philosophy that you can just build things.

Last updated: October 2026

Quick answer. A TLS handshake fails before your application has a session to log about, so the error surfaces as something generic like certificate verify failed or handshake failure while the actual cause sits in the exchange that just happened. The useful thing about that exchange is that its opening is cleartext: the ClientHello carries the hostname your client asked for in the SNI extension, and the alert that ends the handshake is a TLS record of content type 21, both readable on the wire without any key material. On Linux you can read them with a TCX eBPF tap like pktscope on yeet, which needs no private key, no proxy in the path, and no restart of the process you are debugging. Check the SNI first, because a surprising share of certificate errors are a client asking for a name nobody expected.

Handshake failures are frustrating in a specific way: the error message is produced by a library that knows the negotiation failed and often very little about why, and the application around it knows even less. You get certificate verify failed with no indication of which certificate, which name, or which side objected. Meanwhile the thing that would answer all three went past on the wire a few milliseconds earlier.

The reason this is recoverable is that TLS does its setup in the open. Encryption begins after the handshake, not before it, so the early records are readable by anyone on the path. RFC 8446, the TLS 1.3 specification, defines the record layer and the handshake messages this post reads. That includes the single most useful field for debugging, the server name your client asked for, and the alert record that says the negotiation is over. Neither requires a key to read.

Why does my TLS handshake fail before my app logs anything?

Because the handshake happens underneath your application, in the TLS library, before there is a connection for your code to log about. Your application calls something that amounts to "connect and secure this", and the library performs several round trips of negotiation on its own. If any step fails, the library returns an error and your code never had a session, a request, or a response to record. What reaches your log is the library's summary, which is often a single enum rendered as a string.

That is why the error names a category rather than a cause. certificate verify failed tells you verification did not succeed, which could mean the chain was incomplete, the name did not match, the certificate had expired, or your trust store lacks the issuer. All four produce similar messages from the client's point of view, and the client is frequently the only side you have logs for.

The practical consequence is that adding more application logging rarely helps. The information that distinguishes those four cases exists in the handshake, not in your process, and the way to get it is to observe the handshake rather than to instrument the code that sits above it.

How do I see which hostname (SNI) my app actually sent?

Read the SNI extension out of the ClientHello, which is the first message of the handshake and is sent in cleartext. Server Name Indication, defined as the server_name extension in RFC 6066, exists because one address can host many certificates, so the client has to say which name it wants before the server can choose a certificate. That makes it both the server's routing key and your best debugging signal, because it is the client's actual intent stated plainly rather than inferred from config.

This matters more than it sounds. A large share of certificate errors are not certificate problems at all: the client asked for a name the operator did not expect, so the server presented a certificate for a different name and verification correctly failed. Causes include a hostname built at runtime from configuration, a proxy environment variable rewriting the destination, a stale DNS answer pointing at the wrong vhost, or a client library that omits SNI entirely on an IP-literal connection. RFC 6066 is explicit that literal IPv4 and IPv6 addresses are not permitted in server_name.

Reading it live takes a tap on the interface and a filter. pktscope decodes TLS records and extracts the SNI from the ClientHello, so the name appears as a field you can filter on:

yeet run gh:yeet-src/pktscope -- --iface eth0 --port 443   # a TCX tap
# then, in the filter bar:
tls                                      # TLS records only
sni example.com                          # one name

The field carries the byte range it decodes, so you can see the name in the hex pane and confirm you are reading what was actually sent rather than what you expected to be sent. When the SNI is absent, that absence is itself the finding, and it usually explains a server presenting its default certificate.

Why do I get certificate verify failed, and which side rejected it?

The client rejected it, in essentially every case where you see that specific message, because verification is the client's job. The server presents a chain; the client checks that the chain terminates in a trusted root, that the leaf covers the name it asked for, and that nothing in the chain has expired, the path validation rules of RFC 5280. certificate verify failed is the client reporting that one of those checks did not pass.

Which of those checks failed is the real question, and the handshake separates them. The server's certificate message is part of the cleartext handshake, so the chain it actually sent is observable, and comparing it against what you believed was deployed settles the most common cases quickly.

What you observeLikely causeWhere to fix it
Chain is short, missing an intermediateServer not sending the full chainServer certificate bundle
Leaf name does not cover the SNI sentWrong vhost, or unexpected client hostnameDNS, client config, or server routing
Chain looks correct, client still rejectsTrust store lacks the issuerThe client's CA bundle, or the container image
Dates outside the validity windowExpired certificate, or host clock skewRenewal, or NTP on the client

The case worth calling out is the third, because it produces the most wasted time: the server is configured perfectly and the client is missing a root, which happens constantly in minimal container images that ship without ca-certificates. Nothing on the server side will reveal it, and the handshake shows a chain that looks fine.

Which side sent the TLS alert, my client or the server?

Read the direction of the alert record. TLS signals failure with an alert, which is a record of content type 21, distinct from handshake records (22) and application data (23). RFC 8446 lists these as alert(21), handshake(22) and application_data(23) in the ContentType enum. Whichever side sends it is the side that decided the negotiation was over, and that single fact redirects the whole investigation.

An alert from the server usually means it objected to something the client offered: a protocol version it will not accept, a cipher suite list with no overlap, or a name it has no certificate for. An alert from the client usually means verification failed, which is the certificate case above. Knowing the direction before you start changing configuration saves you from tuning the wrong end.

A packet view gives you both the direction and the position in the exchange, which is the other half of the signal. An alert immediately after the ClientHello says the server rejected the opening terms, long before certificates mattered. An alert after the server's certificate message says the client did not accept what it was shown. The timing distinguishes a version or cipher mismatch from a trust problem without any further testing.

Being precise about the limit: the record type and direction are decoded, so you can see that an alert occurred, from which side, and at what point in the handshake. The alert's specific description code is not broken out as a named field, so for the exact reason string you still read the client library's error, from openssl s_client or your language's TLS stack, alongside the position in the exchange. In practice the direction and timing narrow it to one or two candidates and the library message confirms which.

Can I debug a TLS handshake on Linux without the private key?

Yes, for the handshake itself, because the parts that fail are the parts sent before encryption starts. The ClientHello with its SNI and offered versions, the ServerHello with the chosen parameters, the certificate chain, and any alert record are all cleartext on the wire. None of them need a key to read, which is what makes handshake debugging fundamentally different from debugging the traffic that follows it.

What you cannot read without key material is the application data after the handshake succeeds. That is the correct behavior and it is also not what you need here: a failing handshake never produces application data, so the bytes you want are exactly the ones still readable. The common instinct to reach for a TLS-terminating proxy such as mitmproxy is usually unnecessary for this class of problem, and it changes the thing you are measuring by substituting the proxy's TLS stack for your client's.

This is also why the approach works against a process you cannot modify. There is no agent to install in the application, no certificate to add to its trust store, and no restart required, because the observation happens on the packet path rather than inside the process.

The alternatives differ on what they cost you rather than on what they can see, which is the comparison worth making before you reach for one:

ApproachReads the SNIReads the alertNeeds key materialChanges what you are measuring
TCX tap (pktscope)Yes, liveType and direction, not the codeNoNo, it never enters the path
tcpdump plus WiresharkYesYesNoNo, but it is a file to analyse afterwards
openssl s_clientYes, you set itYes, verboselyNoYes, it is a different client than yours
TLS proxy (mitmproxy)YesYesInstalls its own CAYes, it substitutes its TLS stack for yours
Application loggingOnly if the library exposes itAs a generic error stringNoRequires a restart to add

The row that catches people is openssl s_client. It is the most common first move and it answers a subtly different question: whether a client can complete a handshake with that endpoint, not whether your client did. When your application fails and s_client succeeds, the difference between the two clients is the bug, and only observing the real one will show it.

Can I debug TLS on a process I cannot restart, with a TCX tap?

Yes, and that is usually the reason to do it this way. A TCX tap attaches at the kernel's traffic control hook, so it observes frames on the interface without being in the path and without the application knowing. Nothing is injected, no library is preloaded, and the process keeps running with the same TLS stack it had before you started looking.

That matters for the failures that are hardest to reproduce. Handshake errors that appear only under load, only from one host, or only against one upstream are exactly the ones that disappear when you restart the process to add logging or route it through a proxy. Observing without touching keeps the failing condition intact.

The boundary is the same one that applies to any host-local observation: you see your host's interfaces. A handshake that fails at a load balancer in front of the real service shows you the conversation with the load balancer, which is what your client actually experienced, but it will not tell you what happened on the far side of it. For that you need the same observation at the other end, which is a different host rather than a different technique.

The bottom line: the handshake is cleartext, so the failure is visible

The reason TLS errors feel opaque is that people look for the answer in the layer that reports the problem rather than the layer that produced it. The library gives you a category. The exchange gives you the hostname that was asked for, the certificate chain that was offered, which side gave up, and when.

Start with the SNI, because it is the cheapest check and it invalidates the most assumptions. Then read the direction of the alert to decide which end to investigate. Then compare the chain actually sent against the one you think is deployed. Those three observations resolve the large majority of handshake failures, and none of them require a key, a proxy, or a restart.

Frequently asked questions

Why does my TLS handshake fail with certificate verify failed?

The client rejected the server's certificate chain during verification. The usual causes are a missing intermediate certificate in what the server sent, a leaf certificate that does not cover the hostname the client requested, a trust store on the client without the issuing root, or dates outside the validity window including client clock skew. The message is the same for all four, so you separate them by reading the chain the server actually sent.

How do I see the SNI a client is sending on Linux?

Capture the ClientHello and read the Server Name Indication extension, which is cleartext. A TCX eBPF tap such as pktscope decodes TLS records and exposes the SNI as a filterable field, so tls plus sni example.com isolates the handshakes for a given name. This is often the fastest way to discover that a client is requesting a hostname nobody expected.

Can I debug TLS without the server's private key?

Yes for the handshake, no for the application data. Everything that fails during negotiation, the ClientHello and its SNI, the server's chosen parameters, the certificate chain, and the alert that ends it, travels before encryption begins and is readable on the wire. Traffic after a successful handshake needs key material or a terminating proxy, but a failing handshake never reaches that point.

What is a TLS alert record?

It is a TLS record with content type 21, as distinct from handshake records (22) and application data (23), used to signal that something went wrong or that the connection is closing. When a handshake fails, one side sends an alert, and the direction of that record tells you which side made the decision. Reading the direction and its position in the exchange is usually enough to tell a version or cipher mismatch from a certificate trust problem.

Why does my handshake fail only inside a container?

The most common reason is that the container image has no CA bundle, so the client cannot build a path to a trusted root even though the server's configuration is correct. Minimal base images frequently omit ca-certificates. Other container-specific causes are a clock that drifted from the host, and a proxy environment variable in the image that redirects the connection to a different endpoint than you expect.

Does an IP address connection send SNI?

Usually not. SNI carries a hostname, so clients connecting to a bare IP address often omit the extension entirely, which leaves the server no basis for choosing between multiple certificates. The server then presents its default certificate, which commonly does not match what the client expects to verify, producing a certificate error that looks like a server misconfiguration but is really a missing name.

How do I tell a cipher mismatch from a certificate problem?

Read where in the handshake the alert appears. An alert arriving immediately after the ClientHello, before any certificate has been sent, points to the server rejecting the offered versions or cipher suites. An alert arriving after the server's certificate message points to the client rejecting what it was shown. The position separates the two without changing any configuration.

Can a load balancer cause TLS handshake failures?

Yes, and it is easy to misattribute, because your client's TLS session terminates at the load balancer rather than at the service behind it. The certificate you are verifying is the load balancer's, the protocol versions and ciphers are its, and a failure there never reaches the backend. Observing from your host shows you the conversation your client actually had, which is with the load balancer.

Why does the error say handshake failure instead of naming the cause?

Because the TLS alert protocol signals a small set of categories rather than a diagnostic message, and libraries render those categories as short strings. The protocol was not designed to explain failures to operators, partly to avoid leaking information to an attacker. The detail you want is in the records exchanged before the alert, which is why reading the handshake is more productive than parsing the error text.

Does observing the handshake require changing my application?

No. A TCX eBPF tap attaches at the kernel's traffic control hook and reads frames on the interface, so there is no library to preload, no proxy in the path, no certificate to install in the application's trust store, and no restart. That is what makes it usable against a running process whose failure disappears when you restart it.

Sources

  1. RFC 8446, the TLS 1.3 specification, defining the record layer and the ContentType values alert(21), handshake(22) and application_data(23). https://www.rfc-editor.org/rfc/rfc8446.html
  2. RFC 6066, defining the server_name extension that carries SNI, and excluding literal IP addresses from it. https://www.rfc-editor.org/rfc/rfc6066.html
  3. RFC 5280, the certificate path validation rules a client applies when it verifies a chain. https://www.rfc-editor.org/rfc/rfc5280.html
  4. openssl s_client, for reproducing a handshake by hand and reading the library's own error. https://www.openssl.org/docs/man3.0/man1/openssl-s_client.html
  5. Linux kernel BPF program types, for the traffic control attach point used to read the handshake without entering the path. https://docs.kernel.org/bpf/libbpf/program_types.html
  6. pktscope, a terminal packet analyzer built as a yeet script on a TCX (clsact) eBPF tap, decoding TLS records including the SNI from the ClientHello. https://github.com/yeet-src/pktscope
  7. yeet documentation, the runtime the tap runs on. https://yeet.cx/docs/

Related resources

  1. Who is resetting my TCP connections on Linux, the connection-layer failure underneath a handshake that never completes. https://yeet.cx/topical-takes/who-is-resetting-my-tcp-connections
  2. What the TC layer sees, on observing traffic at the traffic control hook without a sidecar or a proxy. https://yeet.cx/topical-takes/what-the-tc-layer-sees
  3. How to monitor HTTP traffic on Linux, the request-level view once the handshake succeeds. https://yeet.cx/topical-takes/monitor-http-traffic-linux