Test the edges, not the average.
Fairness isn't a feeling. It's a set of simple checks you run before launch and again whenever something changes.
Practical checks
Before Launch
- Compare outcomes across groups
- Swap one detail, like a name, and compare
- Check accuracy per language and accent
- Have people review borderline cases
Keep checking
After Launch
- After every model or prompt change
- When complaints cluster in one group
- On a regular sample, quarterly
- When you enter a new market
Check who it works worst for.