Interview with Julian Goldstein

Necco Ceresani
Necco Ceresani
·24 min read

Fifteen years building engineering platforms, currently focused on advanced AI infrastructure at yeet. I love turning the deeply complex topics into something everyone can understand. I relate deeply with the core yeet philosophy that you can just build things.

Julian Goldstein, founder and CEO of yeet, sat down with our team to talk about what he is building, why eBPF has stayed out of reach for almost everyone, and what changes when agents can finally see and operate the machines they are writing code for. This is a transcript of the conversation, edited lightly for clarity.

What yeet is

Let's start with the obvious one. What is yeet?

Yeah, so I guess the one-liner is we make machines readable by machines.

And what I mean by that is, right now, if you want to know what's happening inside a running Linux server, you've basically got three options and they're all the same option. Logs, metrics, traces. The three golden pillars. And the problem with all of them is you had to know the question before the question showed up. You had to instrument for it ahead of time. So you end up shoveling everything into a database forever, just in case, and then building really clever query engines on top to mask how dumb the capture was.

That's the whole industry. It's smart data retrieval on top of dumb capture. yeet is just smart capture. You manufacture the instrument you need, at the moment you need it, and you get the actual answer.

So how it works. There's this subsystem that's been getting phased into modern Linux kernels called eBPF, and it lets you stick probes basically anywhere in the operating system and pull real data out. It's one of the coolest primitives I've ever come across. But almost nobody uses it, because it's brutally hard. There's like no documentation, you kind of have to learn it by reading kernel source, and everyone who does learn it goes and builds one point solution with it. A profiler. A security agent. One fixed thing that can't be pointed anywhere else.

What I did instead was go underneath all of that and build the layer that makes the primitive generally usable. Brick by brick from the kernel upward. And the trick, the part I'm actually proud of, is we took V8, the JavaScript engine out of Chrome, and remixed it into a system daemon. So instead of pointing it at the DOM and rendering web pages, it points at the seven thousand trace points in the Linux kernel.

Which means two things. One, you get browser-grade isolation for free, so untrusted code running in a trusted area is a solved problem, it's the same thing as an ad on a website. And two, and this is the part that really matters right now, the API looks like the web. So when an AI goes to write against it, it goes "hey, I've seen this before," even though it really hasn't. You get everything the model learned from the entire web, pointed at the kernel, for free.

So you end up with these things I call instruments. You just kind of jab one into the system and the data starts shooting out like a firehose. And an agent can write about a hundred lines of JavaScript to shape that into whatever it actually needed.

The reason I think this matters more every month is that agents are writing an absurd amount of code now, and all of that code has to run somewhere, and that somewhere is Linux. But agents are completely blind and handless on the box. They can write anything and operate nothing. So yeet is kind of the oven mitts. It's the thing that lets an agent reach into production infrastructure and actually touch it without getting burned.

And what's cool is you can build wildly different things on it and it's all the same engine. One of our enterprise customers built the first HTTP/2 L7 firewall on it, decodes HPACK in the kernel, blocks at line rate, sub-microsecond. Their own kernel guy told them it couldn't be done. We did it in about four hours. Same runtime also profiles their ancient Java app across 3,000 machines so nobody gets paged at 3am. Same runtime, a prediction markets company uses it to catch API drift and ship the PR automatically, fully closed loop.

Firewall, profiler, API monitoring. Those look like three different products, three different companies even. They're all manifestations of the same thing. Observability, networking, security. Those were always the same thing wearing different hats, and I don't think anybody's gone a layer below all three and built the thing underneath the thing.

So that's yeet. It's a runtime. It's a thing that makes things. And the version of this I'm actually chasing is that it becomes the substrate every agent swims in whenever it needs to understand or operate a Linux machine.


You said "a thing that makes things" just now, and I want to pin that down. Is yeet actually a platform, or is it really a very sophisticated Linux toolkit?

Platform. And I'd push back a little on the premise, because the toolkit is the output, not the thing.

Every other eBPF company is a point solution. They pick one thing, say I'm going to build the best firewall in the world, and go sell that firewall. I went the other direction. I went underneath all of those and built the thing that builds all of them.

So the tools you see, the profiler, the firewall, the packet sniffers, those are manifestations of one runtime. They're outputs. If we'd built them as products we'd be a toolkit. We built the engine, and the tools fall out of it.

The way I'd put it is yeet is a thing that makes things. Ephemeral instruments that live for a second, real tool sets like that sub-microsecond firewall, and eventually entire products. Same substrate all the way up.

And the reason that matters right now is there's a huge shift happening in enterprise software, from buying tools to building them. Engineering teams have really capable agents at their disposal now, so they want these highly specialized tools that used to mean either an expensive third-party contract or hundreds of hours of somebody's time. yeet is how you do that for infrastructure. Instead of buying another specialized tool, or spending months building one from scratch, you build exactly the thing you need on yeet, in hours.

So if I had to say it in one breath: we're an agent-native infrastructure platform. You build instead of buy, and you save hundreds of engineering hours and millions of dollars in software you never had to purchase.


Say I'm an engineer and I already live in Claude Code all day. What does yeet give my agent that it doesn't already have?

Claude Code is brilliant at writing software. It just has no idea what's actually happening on the machine.

You can generate enormous amounts of software now, very quickly. But there's a wall, and the wall is that you can't see your own system at runtime. If you ask a coding agent to go figure out why this node is behaving weird, it'll write you a script that shells out to top and greps some logs, and it'll be approximately right, and it won't tell you the actual answer.

What we give it is hands. I call it oven mitts, it's the thing that lets an agent actually reach into production and touch it without getting burned.

Concretely, there are about 7,000 trace points in the Linux kernel, and agents turn out to be perfect at exactly this kind of problem. It's trivia plus a hundred lines of last-mile code. Show me the entry point to the USB stack and snap a probe on it. That's a terrible task for a human and a great one for a model.

And the trick that makes it work is that we made the API look like the web. The model already knows JavaScript, already knows how reactivity works, already knows this shape. So you get everything it learned from the web for free, pointed at the kernel. We go further than that, honestly: if we watch agents consistently hallucinate an API that should exist, we just go make it real. Tilt the table toward fast, accurate generation.

Agents don't want the Food Network of use cases. They want the labeled pantry. We label the pantry.


Let's get concrete. If you had to narrow it down, what is it actually really good at right now? Give me the short list.

So there's a pattern we found, and it's held up every single time we ship something. The stuff that wins is the stuff where the surface area is huge, always changing, always dirty, and nobody could see inside it before.

The five, roughly:

One, investigation that actually gets you to root cause. Something's slow or broken in production and every dashboard says everything's fine, because nobody instrumented for this particular problem ahead of time. With yeet you drop an instrument exactly where the problem is, right then, and you get the actual cause instead of a guess. No redeploy, no waiting for it to happen again. People go from question to answer in about a minute on problems their existing stack couldn't see at all.

Two, building the specialized infrastructure you can't buy. This is the build-versus-buy thing I was talking about, in practice. Somebody puts agents into production and needs a security layer for them, something that can see what the agent's actually doing on the box and block it from underneath. Somebody needs to profile a fifteen-year-old Java app nobody's allowed to touch. Somebody's got a C program that's slower than it should be and wants to know exactly where the cycles are going. Somebody just wants to see inside a third-party app they run but didn't write and can't get answers about. You can buy a generic version of some of those. You can't buy the version that fits your system, and until now that meant an expensive contract or months of somebody's time. On yeet it's an afternoon, and it's exactly the thing you needed.

Three, profiling without touching the code. Everyone has that fifteen-year-old Java app the CEO wrote that nobody can touch and nothing can see inside. We go underneath it. No instrumentation, no APIs, no databases, none of that silliness. One of our enterprise customers runs that across 3,000 machines so their people stop getting paged at 3am about crappy Java code.

Four, networking at the packet level. Anything that grabs network traffic, organizes it, and does something with it just wins. The big proof point here is the first-ever HTTP/2 L7 firewall, built on yeet for a large enterprise customer. It decodes HPACK in the kernel, matches on user agents, and blocks at line rate. Sub-microsecond. Their existing vendors couldn't provide that capability, their own kernel engineer said it couldn't be done, and it came together in about four hours.

Five, closing the loop. Seeing the problem is half of it. The other half is doing something about it. A prediction markets platform uses yeet so that the second an upstream exchange changes an API, a signal fires, yeet sniffs the actual traffic, sees exactly what changed, and ships a PR. Full loop, no human. The same team turned that live traffic into integration test fixtures instead of mocks, so their tests are testing reality.

And the thing to notice is those are five totally different products, and they're all the same runtime. That's the whole thesis.


How it spreads

So how does this spread? When a team first picks it up, what's the way in? What's the wedge?

A production question somebody can't answer.

That's it, that's the wedge. Something is broken or slow or weird, and the existing stack can't tell them why, because you had to know the question before the question showed up in order to have instrumented for it. They ask yeet, and they get an answer cheaply and in about a minute.

That's the way in every single time. Then they realize they can ask another question. And another. And then it clicks that this is a runtime and they start building things they don't have.

There's a specific version of this that's becoming our best channel, which is that everybody building AI SRE products and sandbox products hits the same wall. It stops being an AI problem and becomes a Linux problem. They've got great models and no way to safely touch the box. They throw their hands up, and that's exactly where we come in. We're the layer those companies are missing, and there are a lot of them right now.


Walk me through that, then. Someone's kicking the tires on a free tool on their laptop, and at some point it's running across their production fleet. How does that actually happen? And is this a PLG company or an enterprise company, because it sounds a little like you're trying to be both.

Yeah, so the loop is pretty simple.

Somebody has a question. Why is my API slow, what's actually going over this socket, what's this Java app doing. They find one of our open-source tools, they run it, they get the answer. That's minute one.

Then the thing that always happens, and this is the part I love, is they realize the tool is like 90% of what they wanted. And they have this reflex now where they just go, oh, I'll ask Claude to change it. That's just how people work now. So instead of one size fits all, it's kind of all sizes fit one. We ship the 99% and they add their 1%.

And then the third beat is they want it somewhere real. On their AWS, across a private subnet, on 40 boxes instead of their laptop. That's where it stops being a toy. That's where we gate it, and that's where it turns into money.

So it's not two motions, it's one motion with a gate in the middle. The PLG thing isn't marketing for the enterprise thing, it's the same thing earlier.

And for retention, the honest mechanic is that once that data is flowing into their integration tests or their CI or their alerting, we don't come out. One customer took one of our traffic tools and started generating real test fixtures off live traffic instead of mocks. You're not ripping that out. It's load-bearing now.


You mentioned these open-source tools people find. Who's actually writing them at this point, you or the community?

Mostly agents, at this point.

We started by building a small library of tools ourselves. Pick a protocol, build the tool, ship the repo. That was the seed.

But what our users actually do is point their AI models at yeet and let it go. The agent picks up the API, figures out the shape of the system, and starts building. Dozens of tools, for fixing things, for optimizing things, for whatever's going on with their servers that day. Stuff we never would have thought to write.

So the honest framing is we seeded the archetypes and the agents took it from there. Every archetype we ship spawns hundreds of variants we never wrote and never will.

Where this goes is a marketplace. People build tools on yeet, publish them, and get paid for them. Engineers love building tools and they love showing them off, and now their agents can do most of the building. We're going to hand them every resource we have to do it. That's the flywheel, and it's the thing that gets us to a million machines.


Let's put some numbers on it. How many people are actually using this today? Users, customers, machines, however you count it.

Growing really, really fast, and faster than I expected honestly.

We've seen over 20,000 installs of yeet in the last six months, and that number is growing exponentially MoM as more people and agents discover our new capabilities. But the number I care about most out of all that is that 82% of those are live servers. Not laptops, not people kicking tires. Real production infrastructure. That's people putting an unfamiliar daemon on machines that matter, without a sales call, which tells you something.

There are names in that install base I didn't expect this early. Bloomberg's in there. A couple of Microsoft. F5, Snyk, Cisco, Scale AI, GitHub, Adobe. Just a ton of hardcore infra companies getting deep into yeet right now.

And the number we're driving at is a million machines in twelve months. That's the goal, and the reason I think it's real is that agents install us themselves. A model hits a Linux problem it can't solve, finds yeet, and curls it down without a human ever being in the loop. That's a new distribution channel we're very excited about.


Buyers and business model

Who's actually buying this? When yeet lands inside a company, whose budget does it come out of, and where does it sit in the org?

It's whoever owns the infrastructure. That's the real answer and it's deliberately a little fuzzy because the title moves around depending on company size.

At a big company it's the VP of Infrastructure, the person who has security and networking and observability teams sitting underneath them and actually holds the budget. At a 50-person startup it's just the CTO, or whoever got handed the AWS account.

But here's the part I think is more interesting. We get in through the developers. Every single time. Somebody's got a production problem, they find one of our tools, they curl it down, they get their answer in 60 seconds. The infra lead is the person who signs, but they're never the person who discovers us.

And the people who do discover us are a very specific kind of engineer. Elite, and they tend to live in two places. In the enterprise it's the platform engineers, the people who own the infrastructure everybody else builds on. And then there's this new generation of AI-native companies, where it's what I'd call agentic engineers, people whose whole job is basically directing agents. Those are our power users. Right now we're entirely focused on finding them, making them wildly successful, and then getting into the room with the rest of their org.

And I think that role changes pretty dramatically over the next few years. Right now that team is a bunch of people doing manual work. What I want is for that same person to be running a fleet of agents instead, where they're the one directing the whole thing and the agents are the ones with their hands on the machines. Same buyer, way bigger job, way bigger budget.


And is that new money, or are you taking it from somebody? Is yeet additive budget, or is it rip-and-replace?

We start additive and we end up replacing. That's what the data shows us so far.

We come in on an incident. Something's broken in production, nobody can see it, and we answer the question in an afternoon for basically nothing. That's new budget, that's not a line item anyone's defending.

But then, because we're a runtime, we just keep building things. And at some point the customer does the math themselves. One of them said it to us almost word for word: why would I string another Datadog dashboard, you're not charging me per event.

And I think the deeper thing is that this cannibalizes the whole existing business model. Pay-per-event stops making sense the second you can manufacture exactly the events you need at the moment you need them. Why pay to store everything forever when you can create the exact signal you need at the moment you need it? So yeah, I think we have the potential to mess a lot of things up, and I'm kind of excited about it.


Okay, and how do you actually charge for it? What's the business model?

Per host, per month. A host is a Linux kernel.

We have a free tier so you can evaluate it and get real work done. The paywall kicks in when you want it somewhere real: across a modern AWS or GCP setup where you're bridging public and private subnets, across a fleet. That's the natural gate, because that's exactly where our networking layer is doing the work.


Autonomy and maintainability

I want to come back to the agent side, because "agents touching production" is going to make some people nervous. How much of this is actually running on its own?

As much as you want it to be. That's the honest answer, and it's on purpose.

You can run it fully manual. You ask a question, it drops an instrument, you look at the answer, you decide what happens next. Or you can run it fully autonomous. Agents set their own triggers. They wake themselves up later. They send each other mail, so when one wakes up it picks up where the conversation left off. They can spawn whole organizations of sub-agents underneath them to work across a cluster. At that point the dashboard is just for us humans to peek at.

And most people live somewhere in between. Let the agents investigate on their own, but a human signs off before anything changes. Or hands-off on staging and hands-on in production. You set the dial wherever your trust is today, and you turn it up as that trust builds.

For the people who want to go all the way, it's self-operating infrastructure. Something breaks and the PR to fix it is merged in a second. Look at what people are already doing, someone decompiled and recompiled an entire Super Smash Bros game for macOS with an agent. The version I'd describe is an AI fire department for 300,000 servers. Get in the truck, go where the fire is, fix it, come back. And you go check the fishbowl when you feel like it.


Here's the skeptic's question, though. If agents are generating all this infrastructure, what does it look like in a year? Does any of it stay maintainable, or do you wake up with a pile of stuff nobody understands?

Yeah, and I think this is where being a runtime instead of a code generator actually matters a lot.

The failure mode everyone's worried about is right: you let agents generate a mountain of infrastructure and in six months nobody knows what any of it does. That's real.

Two things save you here. First, these aren't big programs. The generated part is the last mile, like a hundred lines of JavaScript on top of primitives we built and we maintain. You're not maintaining a sprawling codebase, you're maintaining a small script against a stable API. The 99% underneath is our problem, not yours.

Second, we built a reproducible build format for this specifically. It didn't exist, so we made it. What runs on your laptop is what runs in production, deterministically. That's the actual thing that keeps generated infrastructure from rotting.


Competition and moat

Let's talk about competition. If yeet didn't exist tomorrow, what would your customers be using instead?

Honestly, for some of the problems people bring us, there isn't a clean second option today.

I'll give you the real example. One of our enterprise customers wanted an HTTP/2 firewall that could match user agents at line rate. Before they came to us they went and asked their own head BPF guy, like, the guy whose whole job is this, can you build this, yes or no. And he said no. That's not a knock on him, he's great. It's just that nobody had built the layer underneath that makes it possible.

So the honest answer is one of three things. Either they don't do it at all, which is most of the time. Or they go hire a really, really expensive systems person and hope, and that person spends eight months on it. Or they buy a point solution that does one fixed thing and can't be pointed anywhere else.

And that's kind of the thing I want people to get. A lot of the existing stack is built around collecting and storing a broad set of data ahead of time, then querying it later. They shovel everything into a database and then build really clever query engines to mask how dumb the capture was. yeet is just smart capture. You manufacture the instrument you need, at the moment you need it, and you get the actual answer.

Historically, you either instrument everything in advance or accept that some questions will be expensive to answer. We're trying to create a third option.


Some of these customers have serious systems people on staff, though. Why can't they just build this themselves?

Because it's not one hard problem. It's a stack of hard problems that normally don't live in the same person's head.

You need to really deeply know BPF and the verifier. You need to know V8 internals well enough to pull it out of Chrome and repurpose it. You need to build peer-to-peer networking over QUIC with all the PKI. You need to invent a file format for reproducible builds of AI-generated instruments, because that didn't exist, we had to make it. And there's no APIs for any of this in Linux. You start from literally the middle and build up.

I had to build all of it brick by brick from the kernel upward. That took years, and I'd been doing systems programming since I was ten.

And the thing is, the payoff is nonlinear. Until you've done all of it, the juice isn't worth the squeeze. So nobody does it. Everyone stops and makes a point solution, because a point solution pays off right away. Nobody's gone a layer below that and built the thing underneath the thing.

There's just not that many people in the world who understand Linux at this level and also actually want to make it usable for everybody else. That second part is rarer than the first.


Let me ask the open-source question directly, because I'm sure you get it a lot. eBPF itself is open, it's in every kernel. So what's actually defensible beyond the primitive? And why does the daemon stay closed?

eBPF being open is exactly why this works, it's not a threat. It's in every modern Linux kernel already, which means our install has no dependency, no agent to negotiate, nothing to convince anyone to adopt. That's a gift.

The defensibility was never the primitive. Everyone can read the same BPF man page I read. The defensibility is everything I had to build to make that primitive usable: the abstraction layer, the JavaScript runtime tuned for code generation, the peer-to-peer network, the reproducible build format, the file formats that didn't exist. That's years of work that doesn't look like anything until the day it all clicks.

The daemon stays closed right now for a pretty simple reason: it's the thing that makes the whole thing coherent. But I've deliberately de-risked that. Even if it were open tomorrow, the business doesn't change, because what actually matters is the network between the nodes and the certificates that secure it. We'd still run that.

And over time I want to open more, not less. A closed thing you can't inspect is a worse product, and open forces you to actually be the best. When it's not behind the glass anymore you have to win on merit.


Say someone did decide to come after you. What about this gets harder to copy the longer you're around?

Three things, and they all compound.

First is the network. Not the network effect, I mean the actual peer-to-peer fabric between the nodes. That was like digging under a sea cable. Enormous amount of unglamorous logistics about how does this get over there, how do you do the certs, how do you bridge a public and private subnet. And here's the thing I'd say to anyone worried about the open-source question: it wouldn't actually matter if the engine were open tomorrow, because what matters is the network between the nodes, and we control that.

Second is the community we're building. Every instrument, every tool, every agent that gets built on yeet is another thing in circulation that everyone else can pull from. More machines means more instruments means more agents means more places to dispatch the fire department. It's just more miles on the car.

Third, and this is the one I'm most excited about looking out a few years, is the data. A load balancer falling over looks basically the same at company A and company B. There's like ten statistical reasons an AWS load balancer dies. Once we're sitting on enough machines, we can train models against that and fine-tune them to identify infrastructure problems faster than any human. Almost like online reinforcement learning against reality. Nobody else is in a position to collect that, because nobody else is down at this layer.

So year one the network moat is that it's really hard. Year three the moat is the network and the corpus. Those are different moats and that's kind of the point.


Where it's going

Okay, last one, and it's the big-picture one. When you zoom all the way out, is yeet a foundational runtime, or is it a really powerful set of primitives for engineers? Which company are you building?

It's the foundational runtime. That's the company.

The primitives are real, and engineers do use them directly, and that's a great business on its own. But it's not what I'm building.

What I think happens is that as agents write more and more of the world's software, all of that software has to run somewhere, and that somewhere is Linux. And right now agents are essentially blind and handless on that box. They can write anything and operate nothing. That gap doesn't shrink as models get better, it actually gets worse, because there's exponentially more code running that nobody wrote by hand and nobody fully understands.

So the thing I want yeet to be is the layer an agent reaches for whenever it needs to understand or operate a Linux machine. Not a tool you go pick. The substrate all of them swim in. In the same way Unity means a game developer never thinks about physics, we take the brutal kernel machinery and abstract it so the agent can just focus on the thing it's actually trying to do.

The one-liner I keep coming back to is: we make machines readable by machines.

If that works, we're not competing with the observability companies. We're the thing underneath all of them, and underneath the security companies, and underneath the networking companies, because those three were always the same thing wearing different hats. That's the version of the company I'm building toward.