Back to blog
ALIGNMENTVALUESETHICS

Whose Values, Exactly?

AI alignment is not just technical. It is a question of whose values count.

10 min read
A branching network of different value paths converging on an AI system
Alignment starts with deciding whose values count.

The AI alignment problem is usually framed as a technical challenge: how do we build AI systems that reliably pursue human values? But embedded in this framing is a prior question that receives less attention and that is, in some ways, more fundamental: whose values, exactly, are we trying to align AI to? The technical challenge assumes that there is a set of human values that AI could in principle be aligned with, and that the challenge is figuring out how to do the alignment once the values are specified. The prior question is whether the values to be aligned to can be specified in a way that is acceptable to the diversity of humans whose lives the AI will affect. The answer, when examined carefully, reveals a political problem inside the technical challenge that is as difficult as the technical problem itself.

The disagreement about values among humans is not a temporary condition that will resolve itself as AI development advances. It is a persistent feature of human existence that reflects genuine differences in how people understand what matters, what kind of life is worth living, what obligations people have to each other, and what trade-offs between competing goods are acceptable. These differences exist across cultures, across political traditions, across religious and philosophical frameworks, and across individuals within any given community. They are not, in most cases, resolvable by appeal to facts or by better reasoning: they reflect different starting points about what is ultimately important, and those different starting points are not straightforwardly wrong in ways that could be corrected.

The AI systems that will be most consequential are not systems with narrow, technical objectives. They are systems that will make or inform decisions about who gets jobs, who receives healthcare, who is released from prison, who is targeted for advertising, who is flagged as a security risk, and how information is filtered and presented to billions of people. In each of these domains, the values embedded in the AI's decision-making are not neutral technical choices. They reflect judgments about fairness, about desert, about the appropriate trade-offs between efficiency and equity, about who bears the costs of uncertainty. These are exactly the domains where human value disagreement is most persistent and most consequential.

Different value systems connected to one black-box AI model
Every model carries someone’s choices.

One of the most important points about the values-in-AI question is that there is no such thing as a value-neutral AI system. Every design choice that goes into an AI system reflects some values, and the appearance of value neutrality is itself a value choice, usually one that reflects the values of the people making the design decisions and that advantages the interests of those people over the interests of others. The choice to optimise for accuracy as the primary metric reflects a value judgment that accuracy is the most important property of the system, a judgment that is not obviously correct in domains where accuracy and equity trade off against each other. The choice of training data reflects judgments about which kinds of outputs are desirable, and those judgments embed the values of the people who labelled the data.

The hidden value choices in AI systems are not the result of bad faith or insufficient care. They are the inevitable result of making design decisions that require judgment, in a context where the people making the decisions reflect a particular set of backgrounds, values, and interests. The homogeneity of the teams that have built most consequential AI systems, which has been well documented and is gradually improving but remains significant, means that the value choices embedded in those systems reflect a relatively narrow range of human experience and perspective. The systems are built to work for people like the people who built them, and the ways in which they work less well for others are not always visible to their builders.

Making the value choices in AI systems explicit is the first step toward making them accountable. The system that optimises for accuracy at the expense of equity reflects a value judgment that has consequences for particular groups of people. Making that trade-off explicit allows it to be contested, alternatives to be considered, and the people who bear the costs of the choice to have standing to object. The system whose value choices are hidden behind apparent technical neutrality does not offer these affordances, and the people who bear the costs of those choices have less standing to contest them.

A diverse set of choices revealing hidden assumptions inside an AI system
Hidden values still shape real outcomes.

The governance question of whose values get aligned into consequential AI systems is being answered right now, mostly by the companies and research institutions that are building those systems, mostly in ways that reflect the interests and values of those organisations and the regulatory environments in which they operate. This is not a satisfactory answer from the perspective of democratic legitimacy or from the perspective of the billions of people whose lives will be shaped by these systems and who have had no meaningful input into the value choices embedded in them.

The alternatives to this default are not straightforwardly better. International governance of AI values faces the same collective action and sovereignty problems that have made international human rights governance so difficult: the states that most need their populations' values represented in global governance are least positioned to participate effectively in it, and the states that are most positioned to participate have incentives to embed their own values in ways that serve their interests. Democratic governance of AI values within states faces the problem that democratic majorities do not always protect the interests of minorities, and that the technology evolves faster than democratic processes can keep pace with.

Pluralistic approaches to AI alignment, which try to build AI systems that are sensitive to value diversity rather than attempting to resolve it, are an intellectually interesting response to this challenge. The AI system that behaves differently in different value contexts, deferring to local norms in ways that respect local autonomy, solves some problems and creates others: it risks amplifying harmful local norms rather than checking them, and it may create AI systems that are manipulable by whoever controls the definition of local norms. There is no clean solution to the political problem inside the alignment challenge. There are only choices with different trade-offs, and the most important governance task is making those choices explicitly rather than having them made by default.

Open governance pathways surrounding a transparent AI core
Better alignment needs room for disagreement.

The recognition that AI alignment cannot resolve human value disagreement does not make alignment less important. It changes what alignment means and what success looks like. If alignment cannot mean embedding a single value system that all humans share, it can mean building AI systems that are transparent about their value choices, that are accountable to the people affected by those choices, that include meaningful mechanisms for those people to contest choices that harm them, and that are designed to be modified when value choices prove inadequate.

These are political and institutional design challenges as much as technical ones. The AI system that is transparent about its value choices requires interpretability techniques that allow those choices to be examined. The system that is accountable to affected communities requires governance structures that give those communities standing and mechanisms to use it. The system that can be modified when value choices prove inadequate requires governance processes that are fast enough to keep pace with the technology. Building these features into AI systems requires the kind of sustained political will and institutional investment that technical AI development has not typically attracted.

The question in the title, whether we can align AI when humans cannot agree on values, does not have a simple yes or no answer. We can build AI systems that navigate human value disagreement with more or less transparency, accountability, and responsiveness to the people most affected. We can make value choices that are better or worse at including the perspectives of people who are not well-represented in the institutions that build AI. We can create governance structures that are more or less adequate to the ongoing challenge of managing systems whose value choices have real consequences for real people. Whether we do these things well or poorly is a political question as much as a technical one, and the answer will be shaped by the political choices that are being made now about how AI is governed.

ALIGNMENTVALUESETHICSARTIFICIAL INTELLIGENCESAHIR MAHARAJ

Topics in this article