Hey there! Welcome to Platform Weekly. Your weekly drizzle of platform engineering maple syrup. Every week, we pour the best of the community’s thinking over your Friday stack (no pancakes required).
Plus…
AgenticCon is coming! It is our new conference focused fully on the AI platform teams building, shipping, and delivering INSANE value with agents and agentic engineering. Join us in San Francisco on February 25 if you own your company’s agent infrastructure. Get super early bird tickets here. I am so excited for this!
Did you miss my webinar with Kelsey Hightower, and Dan Ciruli from Nutanix? It was an awesome conversation where I unpacked with them why AI is just another workload, how to secure agents, and what it all means for platform engineers.
Your agents are about to blow up your telemetry bill
For the last few years, telemetry volume has creeped its way up and up and up, but we just about managed to optimize our way around it. That has all changed. Fast.
The economics of telemetry are breaking. Agents aren’t just creepin, they’re stomping all over a problem that was already getting harder to manage. Just look at these numbers.
28-40% annual growth in telemetry volumes, while IT budgets stay flat
2-5x more telemetry from each LLM query than from a standard application log event
10-100x the query load AI SRE tools can put on observability endpoints when they chase parallel hypotheses
And at the same time, 57% of practitioners say their observability setup is too noisy or doesn’t help them find the root cause… ouch.
These stats are from our Market Guide for Observability Trends in Platform Engineering and were shared in Tuesday’s community webinar (alongside how to combat them) by James Conway & Bill Emmett from Cribl.
Agents investigate differently. A human mostly works through one hypothesis at a time. An agent spins up ten, spawns sub-agents, and fires queries at all of them at once. Every architecture and pricing model we’ve built over the previous decades assumed curiosity capped by human-speed. That just ain’t it anymore.
Every one of those agent queries burns tokens and hits your observability backend, and often every one of those queries spits out more telemetry that needs to be absorbed. This is a huge increase in noise and cost, often without enough extra value to compensate.
So what is the answer? Well, James from the webinar and I feel the same thing. You can’t just “buy more storage” your way out of this one. It’s twofold.
You need to stop treating all data the same. He called it caviar versus peanut butter. Your detection tier gets the premium, indexed, expensive treatment. Your compliance and archive data gets spread cheap, but kept at full fidelity so you can replay it later. Only pay to analyze what’s actually worth analyzing.
Then of course… you can’t solve Agent problems with human solutions. There is a reason 47.5% of teams are now (finally) embracing shift down. Shifting this down into the platform is the best way to manage the deluge of telemetry, to get actual actionable results, and keep costs down.
There was a final question in the webinar, I’d love to get your answer to - how many different data engines in your environment are storing the same log right now? Two? Three? Four?
If you’re not sure, go watch the webinar and learn how to figure it out.
There is a good chance you’re one of the around 50% who say cost or wrestling with AI is their biggest observability challenge, so if you’re looking for some big wins… I’d start there ;)
Tell us your about experience:
As always… stay crunchy 🥐
Quick bites
Highlight of the week
We host webinars every week. On Oct 13, Port’s Matar Peles shows how to build an enterprise-grade AI software factory, from ticket to PR. A week later, on Oct 20, Chainguard’s Erika Heidi explains why SHA pinning in GitHub Actions is not the whole story. Both are free, both are live, and both are worth your hour. Come hang out.
From the community
Missed a webinar? Catch up on the Platform Engineering YouTube channel!






