
The enterprise AI industry has a math problem. Cisco data shows 85% of companies They are testing AI agents, but only 5% have shipped them to production. In VB Transformation 2026 on Tuesday, Bryan Silverthorndirector of AGI Autonomy at Amazon, explained why that gap persists and why better benchmarks aren’t the answer.
Silverthorn, who joined Amazon through its acquisition of Adept AI and now leads multimodal agent training within the company’s AGI lab, argued that reliability should be broken down into four distinct dimensions: consistency, robustness, predictability and security, a framework he attributes to Princeton research.
"It breaks down different factors that I see intertwined in almost every evaluation I’ve seen." said.
Why AI Agents Pass Internal Assessments But Fail Real Customers in Production
The framework is important because agents routinely pass internal evaluations and then collapse into the wild. Silverthorn described a customer who implemented an agent for software quality control that involved extracting serial numbers from screens. It worked perfectly for two months and then started reading incorrect numbers intermittently. The culprit: The underlying vision encoder behaved differently depending on where the serial number appeared on the screen, and a software change imperceptible to humans caused the failure.
The lesson, Silverthorn said, is about measurement, not just models. "The models have to be better. Obviously, we are working hard to improve the models." said. But the deeper takeaway, he added, is that teams need to identify their dimensions of variability and match measurement rigor to application challenges. VentureBeat’s own research, presented before the session, reinforces the point: Half of the companies surveyed sent agents who passed internal assessments but failed real customers, and companies overwhelmingly track uptime while ignoring accuracy: checking the pulse without checking the diagnosis. A related finding underscored how few barriers exist: Most companies default to model makers’ own assessments and little else, leaving their testing strategy, as I described it in the scenario, as a toss-up between trusting the supplier and trusting nothing.
Inside Amazon’s ‘intern’ framework for managing autonomous AI agents
Silverthorn’s most memorable recipe was cultural, not technical. Inside Amazon’s AGI Lab, Researchers Literally Call in Their Agents "interns" – as in, "I’ll have my intern talk to your intern." The joke carries a serious operating philosophy. Agents, like interns, are powerful but sometimes clueless, capable of amazing work and spectacular derailments.
Managing them, he argued, requires management skills more than software skills: asking what could go wrong, adding backup and undo capabilities, and consciously deciding what risk can be accepted. "You can ask the intern, ‘Hey, what could you do wrong here? How could you mitigate your negative results?" said. Amazon’s lab has adopted that trade-off, accepting agents occasionally running the wrong experiment in exchange for research speed, including an agent running experiments 24 hours a day in its own high-level research plan.
What business leaders should do before deploying agents at scale
Silverthorn was candid about the limits of current technology. AI keeps improving "a loaded term," He said: Amazon uses AI to constantly improve its models, but fully autonomous self-improvement is far away. Computer use remains a core focus of his lab, and one commercial trucking customer already uses browser automation to stitch together warranty claims across fragmented systems**, although he emphasized that no future agent will rely solely on computer use – it will work alongside MCP, APIs and other tools to complete end-to-end workflows**. And LLM’s judging techniques, while promising, are just one of several strategies for aligning agent capability with acceptable risk.
For companies stuck in pilot purgatory, the path forward begins with a mindset shift: Stop asking if your agent can do something impressive once and start asking if they can do it right a thousand times in a row.
In other words, the companies that escape the 85% ceiling will not be the ones with the smartest agents. They will be the ones with the best managers.





