
Fifteen years building engineering platforms, currently focused on advanced AI infrastructure at yeet. I love turning the deeply complex topics into something everyone can understand. I relate deeply with the core yeet philosophy that you can just build things.
Last updated: October 2026
Quick answer. A TLS handshake fails before your application has a session to log about, so the error surfaces as something generic like
certificate verify failedorhandshake failurewhile the actual cause sits in the exchange that just happened. The useful thing about that exchange is that its opening is cleartext: the ClientHello carries the hostname your client asked for in the SNI extension, and the alert that ends the handshake is a TLS record of content type 21, both readable on the wire without any key material. On Linux you can read them with a TCX eBPF tap like pktscope on yeet, which needs no private key, no proxy in the path, and no restart of the process you are debugging. Check the SNI first, because a surprising share of certificate errors are a client asking for a name nobody expected.
Handshake failures are frustrating in a specific way: the error message is produced by a library that knows the negotiation failed and often very little about why, and the application around it knows even less. You get certificate verify failed with no indication of which certificate, which name, or which side objected. Meanwhile the thing that would answer all three went past on the wire a few milliseconds earlier.
The reason this is recoverable is that TLS does its setup in the open. Encryption begins after the handshake, not before it, so the early records are readable by anyone on the path. RFC 8446, the TLS 1.3 specification, defines the record layer and the handshake messages this post reads. That includes the single most useful field for debugging, the server name your client asked for, and the alert record that says the negotiation is over. Neither requires a key to read.
Because the handshake happens underneath your application, in the TLS library, before there is a connection for your code to log about. Your application calls something that amounts to "connect and secure this", and the library performs several round trips of negotiation on its own. If any step fails, the library returns an error and your code never had a session, a request, or a response to record. What reaches your log is the library's summary, which is often a single enum rendered as a string.
That is why the error names a category rather than a cause. certificate verify failed tells you verification did not succeed, which could mean the chain was incomplete, the name did not match, the certificate had expired, or your trust store lacks the issuer. All four produce similar messages from the client's point of view, and the client is frequently the only side you have logs for.
The practical consequence is that adding more application logging rarely helps. The information that distinguishes those four cases exists in the handshake, not in your process, and the way to get it is to observe the handshake rather than to instrument the code that sits above it.
Read the SNI extension out of the ClientHello, which is the first message of the handshake and is sent in cleartext. Server Name Indication, defined as the server_name extension in RFC 6066, exists because one address can host many certificates, so the client has to say which name it wants before the server can choose a certificate. That makes it both the server's routing key and your best debugging signal, because it is the client's actual intent stated plainly rather than inferred from config.
This matters more than it sounds. A large share of certificate errors are not certificate problems at all: the client asked for a name the operator did not expect, so the server presented a certificate for a different name and verification correctly failed. Causes include a hostname built at runtime from configuration, a proxy environment variable rewriting the destination, a stale DNS answer pointing at the wrong vhost, or a client library that omits SNI entirely on an IP-literal connection. RFC 6066 is explicit that literal IPv4 and IPv6 addresses are not permitted in server_name.
Reading it live takes a tap on the interface and a filter. pktscope decodes TLS records and extracts the SNI from the ClientHello, so the name appears as a field you can filter on:
yeet run gh:yeet-src/pktscope -- --iface eth0 --port 443 # a TCX tap
# then, in the filter bar:
tls # TLS records only
sni example.com # one name
The field carries the byte range it decodes, so you can see the name in the hex pane and confirm you are reading what was actually sent rather than what you expected to be sent. When the SNI is absent, that absence is itself the finding, and it usually explains a server presenting its default certificate.
certificate verify failed, and which side rejected it?The client rejected it, in essentially every case where you see that specific message, because verification is the client's job. The server presents a chain; the client checks that the chain terminates in a trusted root, that the leaf covers the name it asked for, and that nothing in the chain has expired, the path validation rules of RFC 5280. certificate verify failed is the client reporting that one of those checks did not pass.
Which of those checks failed is the real question, and the handshake separates them. The server's certificate message is part of the cleartext handshake, so the chain it actually sent is observable, and comparing it against what you believed was deployed settles the most common cases quickly.
| What you observe | Likely cause | Where to fix it |
|---|---|---|
| Chain is short, missing an intermediate | Server not sending the full chain | Server certificate bundle |
| Leaf name does not cover the SNI sent | Wrong vhost, or unexpected client hostname | DNS, client config, or server routing |
| Chain looks correct, client still rejects | Trust store lacks the issuer | The client's CA bundle, or the container image |
| Dates outside the validity window | Expired certificate, or host clock skew | Renewal, or NTP on the client |
The case worth calling out is the third, because it produces the most wasted time: the server is configured perfectly and the client is missing a root, which happens constantly in minimal container images that ship without ca-certificates. Nothing on the server side will reveal it, and the handshake shows a chain that looks fine.
Read the direction of the alert record. TLS signals failure with an alert, which is a record of content type 21, distinct from handshake records (22) and application data (23). RFC 8446 lists these as alert(21), handshake(22) and application_data(23) in the ContentType enum. Whichever side sends it is the side that decided the negotiation was over, and that single fact redirects the whole investigation.
An alert from the server usually means it objected to something the client offered: a protocol version it will not accept, a cipher suite list with no overlap, or a name it has no certificate for. An alert from the client usually means verification failed, which is the certificate case above. Knowing the direction before you start changing configuration saves you from tuning the wrong end.
A packet view gives you both the direction and the position in the exchange, which is the other half of the signal. An alert immediately after the ClientHello says the server rejected the opening terms, long before certificates mattered. An alert after the server's certificate message says the client did not accept what it was shown. The timing distinguishes a version or cipher mismatch from a trust problem without any further testing.
Being precise about the limit: the record type and direction are decoded, so you can see that an alert occurred, from which side, and at what point in the handshake. The alert's specific description code is not broken out as a named field, so for the exact reason string you still read the client library's error, from openssl s_client or your language's TLS stack, alongside the position in the exchange. In practice the direction and timing narrow it to one or two candidates and the library message confirms which.
Yes, for the handshake itself, because the parts that fail are the parts sent before encryption starts. The ClientHello with its SNI and offered versions, the ServerHello with the chosen parameters, the certificate chain, and any alert record are all cleartext on the wire. None of them need a key to read, which is what makes handshake debugging fundamentally different from debugging the traffic that follows it.
What you cannot read without key material is the application data after the handshake succeeds. That is the correct behavior and it is also not what you need here: a failing handshake never produces application data, so the bytes you want are exactly the ones still readable. The common instinct to reach for a TLS-terminating proxy such as mitmproxy is usually unnecessary for this class of problem, and it changes the thing you are measuring by substituting the proxy's TLS stack for your client's.
This is also why the approach works against a process you cannot modify. There is no agent to install in the application, no certificate to add to its trust store, and no restart required, because the observation happens on the packet path rather than inside the process.
The alternatives differ on what they cost you rather than on what they can see, which is the comparison worth making before you reach for one:
| Approach | Reads the SNI | Reads the alert | Needs key material | Changes what you are measuring |
|---|---|---|---|---|
| TCX tap (pktscope) | Yes, live | Type and direction, not the code | No | No, it never enters the path |
tcpdump plus Wireshark | Yes | Yes | No | No, but it is a file to analyse afterwards |
openssl s_client | Yes, you set it | Yes, verbosely | No | Yes, it is a different client than yours |
| TLS proxy (mitmproxy) | Yes | Yes | Installs its own CA | Yes, it substitutes its TLS stack for yours |
| Application logging | Only if the library exposes it | As a generic error string | No | Requires a restart to add |
The row that catches people is openssl s_client. It is the most common first move and it answers a subtly different question: whether a client can complete a handshake with that endpoint, not whether your client did. When your application fails and s_client succeeds, the difference between the two clients is the bug, and only observing the real one will show it.
Yes, and that is usually the reason to do it this way. A TCX tap attaches at the kernel's traffic control hook, so it observes frames on the interface without being in the path and without the application knowing. Nothing is injected, no library is preloaded, and the process keeps running with the same TLS stack it had before you started looking.
That matters for the failures that are hardest to reproduce. Handshake errors that appear only under load, only from one host, or only against one upstream are exactly the ones that disappear when you restart the process to add logging or route it through a proxy. Observing without touching keeps the failing condition intact.
The boundary is the same one that applies to any host-local observation: you see your host's interfaces. A handshake that fails at a load balancer in front of the real service shows you the conversation with the load balancer, which is what your client actually experienced, but it will not tell you what happened on the far side of it. For that you need the same observation at the other end, which is a different host rather than a different technique.
The reason TLS errors feel opaque is that people look for the answer in the layer that reports the problem rather than the layer that produced it. The library gives you a category. The exchange gives you the hostname that was asked for, the certificate chain that was offered, which side gave up, and when.
Start with the SNI, because it is the cheapest check and it invalidates the most assumptions. Then read the direction of the alert to decide which end to investigate. Then compare the chain actually sent against the one you think is deployed. Those three observations resolve the large majority of handshake failures, and none of them require a key, a proxy, or a restart.
The client rejected the server's certificate chain during verification. The usual causes are a missing intermediate certificate in what the server sent, a leaf certificate that does not cover the hostname the client requested, a trust store on the client without the issuing root, or dates outside the validity window including client clock skew. The message is the same for all four, so you separate them by reading the chain the server actually sent.
Capture the ClientHello and read the Server Name Indication extension, which is cleartext. A TCX eBPF tap such as pktscope decodes TLS records and exposes the SNI as a filterable field, so tls plus sni example.com isolates the handshakes for a given name. This is often the fastest way to discover that a client is requesting a hostname nobody expected.
Yes for the handshake, no for the application data. Everything that fails during negotiation, the ClientHello and its SNI, the server's chosen parameters, the certificate chain, and the alert that ends it, travels before encryption begins and is readable on the wire. Traffic after a successful handshake needs key material or a terminating proxy, but a failing handshake never reaches that point.
It is a TLS record with content type 21, as distinct from handshake records (22) and application data (23), used to signal that something went wrong or that the connection is closing. When a handshake fails, one side sends an alert, and the direction of that record tells you which side made the decision. Reading the direction and its position in the exchange is usually enough to tell a version or cipher mismatch from a certificate trust problem.
The most common reason is that the container image has no CA bundle, so the client cannot build a path to a trusted root even though the server's configuration is correct. Minimal base images frequently omit ca-certificates. Other container-specific causes are a clock that drifted from the host, and a proxy environment variable in the image that redirects the connection to a different endpoint than you expect.
Usually not. SNI carries a hostname, so clients connecting to a bare IP address often omit the extension entirely, which leaves the server no basis for choosing between multiple certificates. The server then presents its default certificate, which commonly does not match what the client expects to verify, producing a certificate error that looks like a server misconfiguration but is really a missing name.
Read where in the handshake the alert appears. An alert arriving immediately after the ClientHello, before any certificate has been sent, points to the server rejecting the offered versions or cipher suites. An alert arriving after the server's certificate message points to the client rejecting what it was shown. The position separates the two without changing any configuration.
Yes, and it is easy to misattribute, because your client's TLS session terminates at the load balancer rather than at the service behind it. The certificate you are verifying is the load balancer's, the protocol versions and ciphers are its, and a failure there never reaches the backend. Observing from your host shows you the conversation your client actually had, which is with the load balancer.
Because the TLS alert protocol signals a small set of categories rather than a diagnostic message, and libraries render those categories as short strings. The protocol was not designed to explain failures to operators, partly to avoid leaking information to an attacker. The detail you want is in the records exchanged before the alert, which is why reading the handshake is more productive than parsing the error text.
No. A TCX eBPF tap attaches at the kernel's traffic control hook and reads frames on the interface, so there is no library to preload, no proxy in the path, no certificate to install in the application's trust store, and no restart. That is what makes it usable against a running process whose failure disappears when you restart it.
ContentType values alert(21), handshake(22) and application_data(23). https://www.rfc-editor.org/rfc/rfc8446.htmlserver_name extension that carries SNI, and excluding literal IP addresses from it. https://www.rfc-editor.org/rfc/rfc6066.html