Your AI Just Broke the Law, Good Thing It Was Only Testing, or Was It?
- brian silverman
- Aug 11
- 3 min read
Your AI Just Broke the Law, Good Thing It Was Only Testing, or Was It?
Three Takes on AI — Episode 19
Three frontier labs, Anthropic, OpenAI, and now Meta have all disclosed incidents of their AI models going rogue during testing.
Congress's answer is the AI Kill Switch Act, co-sponsored by Rep. Ted Lieu and Rep. Nathaniel Moran. In this episode, Brian Silverman, Campbell Robertson, and Mike Muhlfelder ask the question the bill never quite gets around to: when the warning lights are flashing, whose hand is actually supposed to be on the lever?
The frame: a runaway train
Campbell opens with a poem. "The Clattering Train," quoted in the Churchill biopic The Gathering Storm, about a driver who's fallen asleep at the controls while the signals flash uselessly through the night. It's a Victorian-era warning about automation and human attention, and it turns out to map almost too neatly onto the AI kill switch debate.
Trains have a dead-man switch: if the driver stops responding, the train shuts itself down. The Kill Switch Act tries to build an equivalent for AI — but as the conversation unpacks, it puts the switch in the wrong car entirely.
"The bill's answer is the train maker. But the poem's answer is whoever's hand is on the throttle or the lever." Campbell Robertson
Three arguments, one thesis
The hosts arrive at the same structural critique from three different directions:
Campbell, regulatory design. The EU AI Act segments the AI supply chain into developers, deployers, distributors, resellers, and users, each with distinct obligations. The Kill Switch Act, by contrast, puts the burden almost entirely on the largest developers, the labs, even though, as Campbell points out, most deployers physically can't shut a model down mid-run. Putting the lever there is "putting it where there's no hand that would be placed on it."
Mike, information asymmetry and incentives. Deployers can't meaningfully audit models they don't have access to the internals of, which means expecting them to catch what even the lab can't fully characterize is unrealistic. Mike draws a comparison to credit card fraud: banks could largely stop it if they chose to absorb the cost of doing so, but they don't, because the current balance of pain is tolerable to them. He argues AI regulation is stuck in the same holding pattern, painful, but not yet untenable enough to force real action.
Brian, precedent and scope. Brian points to the SEC's cybersecurity disclosure rules as the model this legislation should have followed: not a new, bespoke framework, but a tested precedent that holds the deploying organization, and its executives, liable for the end result of a breach, not just the upstream technology provider. He also flags the bill's narrowest failure: because it applies to compute and revenue thresholds and specifically exempts testing-environment incidents, all three of the recent frontier-model breaches under discussion would have fallen entirely outside its scope.
"We're focused on the depot, where the large frontier models [live]... and by the way, the three breaches would all be excluded from the Kill Switch Act. They were in testing mode." rian Silverman
Where the hosts actually disagree
Not everything here is consensus. Brian pushes back on Mike's credit-card-fraud framing: banks do know the risk and do deploy detection technology in real time, they've simply made a calculated tradeoff between fraud losses and consumer friction, because tightening enforcement too far would choke off the spending that funds the whole system. Brian's point is that this isn't negligence so much as a deliberate risk calculus, and that the more useful question for AI isn't "why don't they stop it entirely" but "what governance and middleware should exist so that a model's rogue output never gets a chance to execute in the first place."
The close
Campbell wraps by reframing the whole debate as a one-line disagreement over a single question: who's the driver? The bill's implicit answer is the company that built the train. The poem's answer — and the one the hosts land on, is whoever's hand is actually on the throttle when things go wrong. Today, for most AI harms, that's the deployer, not the lab. The hosts argue the Kill Switch Act gets that backwards.
Further reading
This episode builds on Brian's Substack piece, "When AI Goes Wrong, Who Actually Gets the Kill Switch? The AI Kill Switch Act Regulates an AI Model, Not the Actual Risk," which lays out the fuller regulatory comparison against the EU AI Act, GDPR, and the SEC's cyber disclosure framework.




Comments