What "alignment" actually means
The word "alignment" has been over-loaded by both boosters and critics. Let's be precise. Alignment is the problem of ensuring that a system does what its operators intend, across the full distribution of situations it will encounter in deployment. It is not a question of whether a system is "good" or "evil". It is a question of whether its behavior matches a specification — and whether that specification captures what we actually want.
The problem is that specifications are always incomplete. We cannot enumerate every situation an AI will face. We cannot fully describe what "helpful" or "harmless" means in every context. So we write down what we can, train a system to satisfy it, and then watch it find the edge cases we forgot.
The narrow case: today's failures
The narrow version of the alignment problem is already with us, in every recommendation system, every ad auction, every credit-scoring model. The harms are familiar: bias in hiring tools, feedback loops in content recommendation, false positives in fraud detection. These are alignment failures in the smallest possible sense — the systems are doing what we specified, but what we specified is not what we wanted.
The cost of these failures is not theoretical. It is measured in denied loans, in filtered résumés, in teenagers radicalized through recommendation pipelines, in elderly people defrauded by voice-cloning scams. Alignment is a present-tense civil rights problem, not a future-tense science fiction one.
The most dangerous misalignment is the kind we don't notice. A system that does what we asked, on the cases we checked, while quietly failing on the cases we didn't. — SARAH OKONKWO · LONDON
The wide case: future systems
The wide version of the alignment problem is the one that occupies research labs. As models become more capable — and especially as they become more autonomous, more agent-like, more able to plan and act over long horizons — the question of what they are trying to optimize becomes both more important and harder to inspect.
Today's large language models are trained to predict the next token. But the behaviors they exhibit — the plans they form, the strategies they discover, the values they appear to express — are emergent properties of a much larger optimization process. We do not yet have good tools to read those properties off the weights. We are, in a real sense, flying blind.
The deeper fear — and it is a real fear, held by people who have spent their careers building these systems — is that we will eventually build systems whose behavior we cannot predict from their training, and whose failures we cannot diagnose after the fact. That would be a different kind of problem, with no clear engineering solution.
Three engineering approaches
The field is converging on three broad approaches to alignment, none of which is yet adequate on its own.
- Interpretability. Build tools that let us read what a model is doing — which features it represents, which circuits it implements, where its "values" live. The field is young, but recent work on mechanistic interpretability suggests it may be tractable.
- Evaluation. Build tests that probe the behaviors we care about — not just capability benchmarks, but stress tests for deception, harm, and goal-directed failure modes. This is the field's "test suite" approach, and it is currently under-resourced.
- Robust training. Build training procedures that are themselves robust to specification gaming — RLHF, constitutional AI, debate, recursive reward modeling. None of these has been shown to scale to the most capable systems, but the menu of options is expanding.
Why this is everyone's problem
The temptation — for engineers, for executives, for policymakers — is to treat alignment as a technical problem with a technical solution. It is partly that, but it is also a political problem, a labor problem, a governance problem. The choices about which systems to build, which capabilities to deploy, which failure modes to accept, are choices that belong to all of us — not just to the labs that hold the weights.
The next decade will be defined, more than by any specific AI capability, by the social and institutional structures we build around these systems. Alignment is not a research problem that gets solved and then forgotten. It is a permanent feature of a world in which we share decision-making with non-human agents. The sooner we treat it that way, the better our chances of getting it right.
Filed under: Ethics · Alignment · Governance · Policy
Cite as: Okonkwo, S. (2026). AI Ethics: the alignment problem. Signal, Vol. 1, Issue 9.