Back to blog
REGULATIONCOMPLIANCEETHICS

The AI That May Never Break a Rule

AI can obey every narrow rule and still create harm at a much larger scale.

10 min read
A pristine rulebook and green compliance light overlooking a structurally fractured city system
Following every rule does not guarantee a good outcome.

One of the foundational assumptions of regulatory governance is that rules can be written that, if followed, will prevent the harms the rules are designed to address. The history of regulation is a history of discovering the limits of this assumption: the financial products that complied with every capital requirement while creating the systemic risk that destroyed the global economy; the pharmaceutical compounds that cleared every safety trial while producing the off-label harms that the trials were not designed to detect; the environmental practices that met every permitted emission level while contributing to the diffuse, cumulative damage that the permit system was not designed to aggregate. In each case, the rules were followed. The harm occurred anyway. The governance failure was not one of enforcement but of design: the rules captured the measurable correlates of the harm without capturing the harm itself.

AI governance is in the early stages of developing its rules, and the pressure to make those rules specific, measurable, and enforceable is entirely understandable. Vague principles do not create compliance obligations, do not give regulated entities clear guidance, and do not give regulators tools for enforcement. The turn toward specific requirements, mandatory impact assessments, prohibited use cases, transparency obligations, is the natural response of governance systems trying to operationalise general commitments into specific constraints. But the history of rules-based governance in complex domains suggests that the most significant harms are often the ones that the rules were not written to address, and that compliance with whatever rules exist can create a false sense of security that makes those harms more rather than less likely.

The specific version of this concern for AI is that the most consequential AI-related risks may be diffuse, emergent, and structural in ways that specific rules about specific AI systems are not designed to address. The AI that breaks the world may never produce discriminatory outputs in the way that anti-discrimination rules prohibit. It may never violate a privacy regulation in any instance that is individually actionable. It may never produce false information in a way that any specific misinformation rule captures. It may do something much larger and harder to govern: gradually transform the epistemic environment, the labour market, the distribution of power, the capacity for democratic self-governance, in ways that are consequential at the civilisational scale but not traceable to any specific rule violation.

A small gear measured precisely while the larger mechanism strains
Narrow rules can miss system-wide harm.

Rules-based governance is genuinely good at something: creating clear, enforceable obligations that prevent specific, identifiable, well-understood harms. It works best when the harm mechanism is specific and well-understood, when individual instances of the harm can be identified and attributed, when the causal link between the prohibited behaviour and the harm is clear, and when the population of potential violators can be monitored. Anti-discrimination law works reasonably well for individual discriminatory decisions that can be identified and attributed. Food safety regulation works reasonably well for specific contamination risks that can be tested for. These are domains where the rules can be written to address the actual harm because the harm is specific, identifiable, and causally linked to specific behaviours.

The AI-related harms that do not fit this profile are numerous and may be the most important ones. The concentration of information power in a small number of AI providers, and the epistemic consequences of that concentration, is not addressable through rules about specific AI system behaviours because it is a structural feature of the information environment rather than a property of any individual system. The erosion of institutional trust that AI-generated misinformation contributes to is not addressable through rules about specific false statements because the harm is cumulative, diffuse, and does not attach to any specific violation. The labour market disruption that AI creates at scale is not addressable through rules about specific AI systems because it is an aggregate consequence of many individual deployment decisions, none of which individually causes the harm.

The governance approaches that complement rules-based regulation and address the harms that rules cannot capture are less familiar and less institutionalised. They include structural governance that addresses the conditions under which harms arise rather than the specific behaviours that produce them: competition policy that prevents the concentration of AI capability in ways that create systemic dependence, investment in the public goods that provide resilience against AI-driven disruption, and institutional design that maintains the human capacities and oversight structures that AI deployment tends to erode. These approaches are politically harder to design and implement than rules, because they address causes rather than symptoms and require sustained commitment over time scales longer than typical political cycles.

Green compliance lights hiding tangled and risky machinery
Compliance can become a polished performance.

There is a specific dynamic that rules-based governance creates in domains where the rules do not adequately capture the relevant risks: compliance theatre. Compliance theatre occurs when regulated entities direct resources toward demonstrating compliance with the existing rules, in ways that consume resources that could otherwise be directed toward genuinely reducing the relevant risks. The organisation that spends heavily on mandatory impact assessments and transparency reporting while not investing in the structural changes that would actually reduce its AI systems' potential for harm is engaged in compliance theatre. The compliance is genuine, the theatre is the substitution of compliance for genuine risk reduction.

Compliance theatre is not simply the result of bad faith. It is the rational response of regulated entities to a regulatory environment in which compliance is what is measured and rewarded, and genuine risk reduction is not measurable by the available tools. The organisation that reduces its AI systems' genuine risks in ways that are not captured by the regulatory framework gets no credit for doing so and bears the cost of the risk reduction. The organisation that complies with the regulatory framework while not reducing genuine risks gets credit for compliance and saves the cost of the risk reduction. The incentive structure systematically rewards compliance over genuine risk reduction when the two diverge.

Addressing compliance theatre requires developing regulatory frameworks that are better at measuring genuine risk reduction rather than only compliance with specific rules. This is harder than it sounds, because the rules exist precisely because genuine risk reduction is hard to measure directly. The path toward better governance is iterative: use rules as the initial approximation of what genuine risk reduction requires, develop better measurement of genuine risk reduction over time, and update the rules to reflect that better measurement. This requires regulatory capacity for learning and adaptation that existing AI governance institutions do not yet consistently demonstrate, and building that capacity is as important as writing better initial rules.

Public institutions connected in a resilient network around an AI core
Resilience takes more than another checklist.

The governance approaches that are adequate to the scale of potential AI harm include but go beyond rules-based regulation. They include the development of strong public AI research capacity that is not dependent on private AI developers for its understanding of what AI systems are doing; the maintenance of diversity in AI capability so that no single failure mode or concentration of power can produce systemic harm; the investment in the social and institutional infrastructure, education, journalism, democratic deliberation, that provides resilience against the harms that AI can produce but that specific rules about specific AI systems cannot prevent; and the development of international coordination mechanisms adequate to harms that cross national boundaries.

These governance investments do not look like AI regulation in the conventional sense, and they are not the kinds of things that AI governance discussions typically focus on. They look more like social policy, competition policy, educational policy, and international diplomacy. But they are the kinds of things that determine whether societies have the structural resilience to navigate the AI transition without catastrophic harm, and they are systematically underinvested relative to the rules-based governance that is easier to design, easier to enforce, and easier to point to as evidence that something is being done.

The AI that breaks the world may never break a rule. Governing against that possibility requires governance that extends beyond rules about AI to include the structural conditions that make catastrophic harm possible or impossible. This is not an argument against rules. Rules matter. Compliance matters. But the limits of rules-based governance in this domain are significant, and the harms that fall outside those limits are among the most serious ones. Taking those harms seriously requires governance ambition that the current AI policy moment has not yet fully risen to.

REGULATIONCOMPLIANCEETHICSARTIFICIAL INTELLIGENCESAHIR MAHARAJ

Topics in this article