The headline dropping this week out of Reuters was as predictable as it will be catastrophic: Mark Zuckerberg had a bold plan to replace Meta staff with AI. Here’s how it imploded.
Across the industry, organizations are fixated on AI adoption. In a rush to cut costs and chase the AI hype cycle, companies are laying off their human operations and reliability teams, replacing them with “autonomous” AI agents.
From an operational standpoint, this isn’t just misguided. It’s dangerous.
The early data is already quietly proving it: many of the organizations that eliminated their human reliability teams are experiencing a massive spike in both the sheer number of incidents and the severity of those incidents. They aren’t preventing outages; they are just accelerating them.
The problem starts with the vendors. The AI operations market is currently flooded with point-solutions claiming you no longer need people to maintain your systems. They sell the dream of full autonomy—a magic button that fixes everything.
But as Ciroos Distinguished Engineer Niall Murphy puts it, selling blind autonomy to an enterprise is like adding a “10x faster” button to your keyboard that only works if you close your eyes.
It highlights a fundamental truth about this market: most vendors don’t actually have respect for—let alone a basic understanding of—the reliability discipline.
Reliability is not a feature you can just toggle on with an LLM. It is a rigorous engineering discipline. Understanding that discipline is critical, both internally for your culture, and externally for the tools you choose to trust. If the core philosophy of a vendor is broken, the tool they sell you will become a liability.
If we take this obsession with “replacing humans with AI” to its logical conclusion, we end up in a terrifying operational state. If AI builds the architecture, and AI monitors the architecture, and AI is trusted to automatically remediate the architecture… what happens when it breaks in a way the AI hasn’t seen before?
We are barreling toward a reality where an engineer is on call for a system nobody wrote, responding to an alert nobody configured, to execute a remediation nobody understands.
That is not reliability. That is a black box of doom.
To be clear, autonomous operations is a future state that we must aspire for. There has to be a bridge from the current state to that future. Identify high-confidence, low risk actions that can be taken by the AI system after trust in that AI system is established. Check the guardrails. Inspect the workflow. Understand the blast radius of actions taken. Ensure that audit logs are in place. What is the rollback plan if anything goes wrong? A well-functioning operations team asks these tough questions (and more) of their staff; the same should be done of an AI system. If your vendor is unable to articulate any of these steps with specificity as it pertains for your enterprise, they are merely riding the hype train.
You don’t close your eyes and push the button. You keep your eyes wide open. And most importantly, you never outsource judgement.