Hey there! Welcome to Platform Weekly. Your weekly scan of the platform engineering radar. Every week, we round up what the community is building, breaking, and arguing about.
Plus…
Datadog broke down for us yesterday exactly how good incident management is done in platform engineering. One of our most in-depth webinars yet, on exactly what you need to do to succeed.
If you’ve also overspent on AI infrastructure without getting the promised value, you’re not alone. Kelsey Hightower from Google is joining us to break down what enterprise AI infrastructure actually looks like in production. You can register for the free webinar here.
AI is changing observability forever
For years I’ve said that observability is a quality problem wearing a scale problem’s outfit. But for a long time, we never got much opportunity to deep dive into proving whether that was the case. Now, with observability becoming a platform problem basically everywhere, I got the chance to go deep.
A few weeks ago, we published our first ever market guide, Observability Trends in Platform Engineering. It was based on survey data and EXTENSIVE interviews. And I mean extensive. Let’s just say there’s a lot of observability folks who aren’t going to hop on a “casual” coffee chat with me anytime soon. We spoke to every vendor under the sun, big enterprises, practitioners, team leads with 100m budgets. Everybody.
Now there’s no way to summarise 11,000 words of gold in one short Platform Weekly. But let me give you some of the juice so you know it’s worth diving into.
28 to 40% annual growth in telemetry volume, while IT budgets stay flat across the board.
Every LLM query throws off 2 to 5x the telemetry of a standard app log event, while an AI SRE chasing 10 to 12 hypotheses at once is generating 10 to 100x the query load we’re used to.
57% of you say your setup is either too noisy or does not point at root cause
As a response to this, 80% are on OpenTelemetry now, with 37% having fully replaced proprietary agents and 43% running hybrid
BUT 58% name the skill gap between developers and SRE as the biggest cultural blocker to owning observability as a platform capability
Dilek Altin, Sam Barlien, and I unpacked a lot of this in a community webinar. Going through the highlights and detailing how observability is one of the clearest domains becoming a fundamental platform capability.
Let’s dive into some of that data.
OTel has gone massively vertical. And for good reason. if its agents reading your traces, standardization is do-or-die. Basically all LLMs are trained on OTel semantic conventions, which means an agent investigating an incident knows exactly where to look. But only if those conventions are actually enforced. Get that right at the Collector and your devs never have to think about instrumentation again (hopefully). That’s a massive massive boon. So it’s no surprise that OTel is becoming the de facto standard.
But this is not happening in isolation. Nobody is going to teach every dev cardinality management when agents deploy thirty services before the first smoke break. One dev adding a user ID to a metric can multiply your time series by orders of magnitude and nobody is going to catch that in a pull request at the speed that agents move. Either the load gets shifted down and the platform absorbs it by design, or nobody does.
This is a massive change from just a short time ago, when observability was still treated as a separate universe to platform engineering.
Which brings us back to that 58%. And to be clear, the skill gap isn’t a hiring problem. No one we talked to was saying they just need to hire new people with these skills. Basically, no one has the skills yet. We’re in a new frontier. It’s a paved road you haven’t paved yet. But it’s also something that’s very surmountable. We’ve got one course out on this already. And another course specifically on the OTel Collector coming soon.
Key thing to understand in all this. The cost blowout, the noise, the agent load and the skill gap are not four different problems. They are 1 problem wearing four hats, and the fix to all of them is the same thing. The platform. So here’s what you need to do on Monday.
Pick one service. Count how many observability decisions a developer had to make to ship it. That number is your backlog.
Then go and read the guide. All 11,000 words of it. Or… you watch me break down in 45 minutes. Depends how much this all matters to you 😉
And as always… stay crunchy 🥐
Quick bites
Highlight of the week
Coder executive forum: Building the agentic enterprise - If you’re in the Bay Area or AI Infra Summit, join us at Levi Stadium. This is going to be the highest-density information event we've ever done!
From the community
Every PlatformCon talk, past and present, is on the Platform Engineering YouTube channel.








