AI Alignment
Whether a system is doing what was intended, and how you would know.
Intent is easy to state and hard to specify. A system optimises what it was given, and what it was given is never quite what was meant. The distance between the two is not a bug to be found; it is a property of writing things down.
The gap stays invisible while the system is weak, because a weak optimiser cannot reach the corner of the space where the specification and the intent disagree. Capability is what makes the difference legible, and by then it is expensive.
The question is whether that distance can be measured before it matters, on a system whose behaviour is not readable from its description.
That makes it an empirical question rather than a philosophical one. Not what the system wants, which nobody has agreed how to define, but which pressures move its behaviour, in which direction, and by how much.
All researchUndecidable Research
Undecidable Research is an informal research collaboration. There is no registered legal entity behind the name.