Bogachov, M. (2026). Fair synthetic data is not about fairness. Big Data & Society, 13(3). https://doi.org/10.1177/20539517261473353
"Data" usually means evidence of what the world is like. Modern AI is trained on unfathomably large datasets, and this makes its impressive capabilities easy to explain: when AI works well, it is because it has picked up meaningful patterns in records of the real world. There are, however, many skills that AI cannot learn from existing data. This is because either the data describing them does not exist, or the skills are so complex that humanity cannot produce enough data for AI to learn them successfully. In response, the AI industry increasingly turns to so-called synthetic training data, one that does not represent actual events or phenomena, but depicts how the world could or should look. Such data is generated using AI itself and is usually very similar to "real" data. Training on it often allows AI companies to cut costs and improve benchmark results at the same time. Still, there are also big risks. Over-reliance on synthetic training data can make AI systems perform poorly in the real world, leak private information, and amplify biases.
Bias in AI systems is studied in a research field, known as algorithmic fairness. There, researchers focus on developing measures that help AI systems, especially those used for decision-making, adhere to the political value of fairness. Some researchers believe that synthetic training data, if used right, can make AI systems less biased than they would be if trained on real data. In my paper, I analyse several proposed approaches to improving fairness in AI through synthetic training data. It turns out that synthetic data engineers tend to misunderstand the nature of the political problem they aim to solve, and end up achieving something else instead. Fair synthetic data promises benefits to society, but in practice it serves the interests of AI developers and deployers, allowing for cheaper training or deployment of AI, just as other synthetic data techniques that do not aim for fairness. So, many of the approaches I've looked into should not be expected to make AI meaningfully fairer. Still, some of the underlying conceptual and technical innovations are genuinely promising, and I discuss how they might be used to make AI substantively fairer in the future.