A loss function is a preference over error distributions
It’s a preference (ordering over options) because:
- any two candidate models produce different error distributions
- the loss assigns each a scalar, which lets them be ordered. The ordering across all possible models IS the preference.