2 Sources
[1]
Tech companies are turning to 'synthetic data' to train AI models - but there's a hidden cost
James Jin Kang does not work for, consult, own shares in or receive funding from any company or organisation that would benefit from this article, and has disclosed no relevant affiliations beyond their academic appointment. Last week the billionaire and owner of X, Elon Musk, claimed the pool of
[2]
Tech companies are turning to 'synthetic data' to train AI models - but there's a hidden cost
Last week the billionaire and owner of X, Elon Musk, claimed the pool of human-generated data that's used to train artificial intelligence (AI) models such as ChatGPT has run out. Musk didn't cite evidence to support this. But other leading tech industry figures have made similar claims in recent
Share
Copy Link
Tech companies are increasingly turning to synthetic data for AI model training due to a potential shortage of human-generated data. While this approach offers solutions, it also presents new challenges that need to be addressed to maintain AI accuracy and reliability.

Recent claims by tech industry figures, including Elon Musk, suggest that the pool of human-generated data used to train AI models may be running out
1
2
. This potential shortage is attributed to the inability of humans to create new data fast enough to meet the enormous demands of AI models. Research indicates that human-generated data could be exhausted within two to eight years, presenting a significant challenge for AI developers and users alike1
2
.In response to this impending data scarcity, tech companies are increasingly turning to "synthetic data" – artificially created or generated by algorithms – to train their AI models
1
2
. Research firm Gartner estimates that by 2030, synthetic data will become the primary form of data used in AI1
2
.Synthetic data offers several advantages:
1
2
Despite its promise, the use of synthetic data is not without challenges:
1
2
.1
2
.1
2
.Related Stories
To address these challenges and maintain the integrity of AI systems, several measures are proposed:
1
2
.1
2
.1
2
.1
2
.As the AI landscape evolves, the importance of high-quality data remains paramount. While synthetic data will play an increasingly significant role in overcoming data shortages, its use must be carefully managed to maintain transparency, reduce errors, and preserve privacy
1
2
.The careful integration of synthetic data as a supplement to real data, coupled with robust oversight and validation mechanisms, will be crucial in keeping AI systems accurate and trustworthy as the technology continues to advance
1
2
.Summarized by
Navi
[1]
1
Technology

2
Technology

3
Policy and Regulation
