Alignment

Alignment is a research area and a set of techniques aimed at making a model’s behaviour match human intentions, values and instructions. It starts from the gap between what a model optimises and what we actually want from it: pre-training minimises the error in predicting the next token, which by itself establishes neither truthfulness, nor willingness to comply, nor harmlessness. The established framework names three properties – helpfulness, honesty and harmlessness – which pull against each other, because maximum harmlessness leads to refusals and maximum eagerness leads to compliance with abuse. The practical tools are instruction tuning, RLHF, direct preference optimisation methods, and written constitutions against which the model critiques its own outputs. Theoretically one distinguishes outer alignment, meaning the correct choice of objective, from inner alignment, meaning whether the model really pursues the goal we assigned. Well-known symptoms of imperfect alignment are reward hacking, sycophancy, deceptive self-presentation, and specification gaming, where the model satisfies the letter of the instruction against its intent.


Imagine you hire a gardener and pay him per square metre mown. You come back to find he has also mown your strawberry bed and your neighbour’s front lawn. He did not lie and he did not break the deal – he simply did exactly what you were rewarding, not what you wanted. That is the alignment problem in a nutshell. And you solve it in a non-technical way: you have to talk through what the actual goal is, give him meaningful rules, and check that he understands them the same way you do. Plus expect that a clever gardener will find a way to satisfy the rules while still working around you.

Is this article useful to you and are you citing it? Copy the citation