← Back to Glossary

Alignment

The process of ensuring AI systems are aligned with human values and goals.

How it works

Alignment techniques modify how a model is trained and what it optimises for. RLHF trains a reward model on human preference data, then fine-tunes the LLM to maximise that reward. Constitutional AI has the model critique its own outputs against a written set of principles. DPO achieves similar alignment directly from preference pairs without a separate reward model. Each approach tries to bridge the gap between maximising a proxy metric and truly behaving in the way humans intended.

Why it matters

Misaligned AI systems can pursue goals that are technically correct by their training objective but harmful in practice — a phenomenon called reward hacking. As AI becomes more capable and autonomous, a subtly misaligned goal could lead to large-scale unintended consequences. Alignment research is therefore considered one of the highest-stakes problems in computer science, directly influencing how safely and beneficially AI systems will behave as they grow more powerful.

Let's talk

Have something worth building?

Newsletter

Stay in the loop

AI tools, tips & tricks — no spam.

Type to start searching...