10 Sources
[1]
Explainer: What is AI model distillation and why is it becoming a US-China flashpoint?
BEIJING, July 31 (Reuters) - A technique that allows developers to shrink powerful artificial intelligence models into cheaper, more efficient systems has become the latest battleground in the intensifying U.S.-China race for AI dominance. Known as model distillation, the method uses the outputs of a powerful AI system to train a smaller model that can perform some of the same tasks with fewer computing resources. Long regarded as a standard tool of AI research, distillation is now at the centre of a growing dispute over whether advanced AI capabilities can be transferred without the consent of the companies that created them. Washington and leading U.S. AI firms have accused Chinese rivals of using the technique to extract capabilities from proprietary models, opening a new front in an increasingly bitter competition over technological leadership. Below are key facts about the practice at the centre of the debate. WHAT IS MODEL DISTILLATION? The largest AI models, known as frontier models, require enormous amounts of computing power, data and investment to train. Model distillation offers a way to create smaller systems by using a large "teacher" model to train a smaller "student" model. The teacher generates examples, such as answers and computer code, which are then used as training material for the student. The smaller model is not a replica of the teacher. It does not inherit the teacher's weights, architecture or full capabilities. Instead, it learns selected behaviours that enable it to perform specific tasks more efficiently. WHY DOES DISTILLATION MATTER? The appeal of distillation is that it can make AI cheaper and easier to deploy. A frontier model may require large data centres and expensive chips to operate. Distilled models can run on less powerful hardware and be tailored for specific tasks. That makes them appealing to companies and governments looking to deploy AI more widely, from devices and factories to vehicles and private networks. WHY ARE REASONING TRACES IMPORTANT? Recent AI systems have increased interest in transferring not only final answers but also the steps used to reach them. These "reasoning traces" can show a smaller model how to approach a difficult problem rather than simply what answer to produce. Florian Tramèr, an assistant professor at ETH Zurich who researches machine-learning security, compared the process to human learning. "If I give you a book of complicated math problems with final solutions, you will have a much harder time learning how to solve problems than if I gave you detailed solutions that describe all steps to take," he said. As reasoning traces have become more valuable, access to AI outputs has become more sensitive because they may expose some of the methods advanced systems use to tackle complex problems. WHO USES DISTILLATION? Distillation is a widely used AI training technique, not an inherently improper practice. U.S. researchers and companies have long used it, including Stanford University's Alpaca project and Microsoft's Orca research, which relied on outputs from more advanced models to improve smaller ones. Chinese researchers have also used outputs from U.S. models in public research projects, including efforts to create Chinese-language instruction models. The key difference lies in access. Open-weight models let researchers inspect and modify underlying parameters. Closed models, such as OpenAI's ChatGPT and Anthropic's Claude, remain under company control and are typically accessed through proprietary interfaces or APIs. WHY HAS DISTILLATION BECOME A U.S.-CHINA ISSUE? The controversy is less over distillation itself and more about unauthorised extraction. AI companies argue there is a distinction between legitimate research and systematically harvesting outputs from proprietary models to replicate commercially valuable capabilities. Anthropic has accused Chinese entities including DeepSeek, Moonshot and MiniMax of conducting large-scale campaigns to obtain capabilities from Claude models. The company said those efforts targeted capabilities including software engineering and advanced reasoning. OpenAI has also said it has detected attempts by Chinese actors to use its models for distillation-related purposes. No Chinese companies have accused U.S. rivals of distilling closed-source models so far. Reporting by Eduardo Baptista Editing by Shri Navaratnam Our Standards: The Thomson Reuters Trust Principles., opens new tab * Suggested Topics: * Artificial Intelligence Eduardo Baptista Thomson Reuters Eduardo Baptista is a Senior Correspondent for Reuters based in Beijing, covering China's technology, space, and automotive industries. He has led enterprise and investigative reporting on China's military-linked companies, artificial intelligence and semiconductor supply chains, as well as macroeconomic and industrial policy. Baptista has reported from China for nearly a decade and holds a BA in History from the University of Cambridge.
[2]
EXCLUSIVE: Chinese military researchers tap US AI models to train defence systems
BEIJING, July 31 (Reuters) - Chinese military researchers have used outputs from leading U.S. artificial intelligence models developed by OpenAI and Anthropic to train domestic AI systems to advance China's defence capabilities, according to a Reuters review of more than 80 Chinese academic papers and patents. The previously unreported findings offer a rare glimpse into how military and security-linked institutions in China are leveraging cutting-edge U.S. AI models as a shortcut to developing specialised systems of their own, despite Washington's efforts to restrict Beijing's access to advanced chips and other strategic technologies. The documents show widespread use of a technique known as "model distillation," in which outputs from a powerful AI system are used to train smaller, specialised models that can be deployed locally without the enormous computing requirements needed to build frontier AI systems from scratch. Reuters' review, which included research compiled by the Washington-based Jamestown Foundation and shared exclusively with the news agency, showed distillation is widely used by researchers linked to the People's Liberation Army and other military institutions. The papers suggest Chinese defence institutions see leading U.S. AI models as both a source of technical insight and a way to close the gap with American rivals. The dispute centres on unauthorised extraction, not distillation itself, a widely used industry practice. The issue has emerged as a major flashpoint ahead of U.S.-China talks on AI governance and safety. U.S. officials have accused some Chinese entities of using distillation to extract capabilities from American AI models, potentially undermining export controls and infringing intellectual property rights. China has rejected the accusations, saying Washington is pursuing AI "hegemonism" while arguing that U.S. firms have engaged in similar practices. Chinese developers have also disputed claims that their AI advances rely on foreign models. AI startup Moonshot last week denied allegations by the Trump administration that its Kimi K3 model was built using distillation, saying it was driven by proprietary innovations. Sunny Cheung, a Jamestown fellow who analysed over 60 of the papers, said Chinese military scientists are systematically capturing the reasoning steps of Western models to adapt them for surveillance, cyber warfare and tactical decision-making. "Teaching a model the right answer is one thing but teaching it the reasoning behind the answer is much harder," said Cheung. "These papers show Chinese military-linked researchers are trying to transfer that expensive, proprietary reasoning from Western models into smaller systems they can control and deploy locally." Reuters verified the academic literature and identified an additional two dozen military-linked case studies. One paper published last year by researchers in PLA Unit 96941, a military intelligence and cyber-warfare unit in Beijing, described using OpenAI's GPT-3.5 to process sensitive military source code. The researchers said third-party models were unsuitable for handling classified information. To overcome that limitation, they used GPT-3.5 to summarize software code and trained a domestic model on those summaries to run entirely within Chinese military networks. The White House, Pentagon, China's foreign ministry, the PLA and OpenAI did not respond to requests for comment. WIDE USE, FROM MONITORING TO MILITARY Chinese researchers have used distillation for purposes ranging from content monitoring to military deployment, Reuters' and Jamestown's review of the papers showed. At the North University of China, which has close links to the country's weapons industry, researchers used Anthropic's Claude 3 Haiku to generate synthetic training data for a text classification model for social media monitoring and content moderation. Anthropic said it does not provide commercial access to Claude in China or to Beijing-controlled firms and uses monitoring systems to detect policy violations. The company added that distilled models may lose the original systems' safety safeguards, potentially allowing sensitive capabilities to be transferred to models beyond its control. A 2024 paper from the PLA's National University of Defense Technology described using distillation to shrink an image-processing model for deployment on unmanned aerial vehicles, allowing drones to analyse live video and support navigation and targeting decisions in real time even when communications are cut. Similarly, researchers at China's Academy of Military Sciences used distillation to run a target-recognition model on tactical hardware during simulated maritime operations involving drones, ships and unmanned submarines, a study published earlier this year showed. BENEFITS AND LIMITS China has embraced distillation as it seeks to compete with the U.S. in frontier AI while facing constraints on advanced computing resources due to Washington's export controls on high-end chips. Central and local governments have promoted "model lightweighting" and edge computing, directing subsidies and research funding toward technologies that enable AI models to run on drones, satellites and other devices with limited processing power. Experts caution, however, that distillation has significant limitations. As Chinese AI models close the gap with their U.S. counterparts, military researchers are also examining distillation as a potential security risk. In January, researchers at the Army Engineering University published a paper on the threat of "data-free distillation," a method of reverse-engineering a model's capabilities without direct access to its core parameters. To counter that vulnerability, they proposed defense mechanisms designed to mask the hidden logical information exposed in a model's public outputs. Distilled models also inherit only selected capabilities and cannot fully replicate the broad intelligence of frontier systems. Trevor Koverko, co-founder of AI data company Sapien, said distilled models remain less capable than their teacher models. "It is best understood as transferring selected capabilities into a cheaper, locally controlled system, not achieving independence from frontier AI." Reporting by Eduardo Baptista; Editing by Miyoung Kim and Shri Navaratnam Our Standards: The Thomson Reuters Trust Principles., opens new tab * Suggested Topics: * Disrupted Eduardo Baptista Thomson Reuters Eduardo Baptista is a Senior Correspondent for Reuters based in Beijing, covering China's technology, space, and automotive industries. He has led enterprise and investigative reporting on China's military-linked companies, artificial intelligence and semiconductor supply chains, as well as macroeconomic and industrial policy. Baptista has reported from China for nearly a decade and holds a BA in History from the University of Cambridge.
[3]
Chinese military reportedly uses American AI models to train its defense systems - tools from OpenAI and Anthropic reportedly among those affected
AI-enhanced defense systems from the East, leveraging closed-weight models from the West * Distillation sidesteps the logic of US export controls, which restrict the chips needed to train frontier models but cannot restrict the text those models produce * Review of more than 80 Chinese papers and patents found PLA-linked researchers distilling outputs from OpenAI and Anthropic models into small systems they can run locally * The Trump administration is increasingly critical of the approach and has claimed Chinese research lab-created Kimi K3 distilled Anthropic's Fable model A new investigation has claimed Chinese military researchers have used outputs from American AI models built by OpenAI and Anthropic to train domestic systems intended to advance the country's defense capabilities. The findings are based on a Reuters review of more than 80 Chinese academic papers and patents, incorporating research compiled by the Washington-based Jamestown Foundation shared exclusively with the news agency. Reuters says it independently verified the literature and found a further two dozen military-linked case studies of its own. A distillation problem in an AI race that continues to heat up The mechanism at issue is model distillation, in which the outputs of a large, capable system are used as training data for a smaller one. The smaller model inherits selected behaviors at a fraction of the compute cost and, crucially, can run on modest local hardware. That is the strategic point. Washington's export controls are built to deny China the chips needed to train frontier models. Distillation does not need those chips, because the expensive part has already been paid for by someone else. It sits alongside Beijing's other route around the controls: substituting domestic silicon for American designs, a shift that carries its own long-term threat to Nvidia and AMD. Sunny Cheung, the Jamestown fellow who analyzed more than 60 of the papers, frames the value as reasoning rather than answers. Getting a model to produce the right output is comparatively easy, he told Reuters, but "teaching it the reasoning behind the answer is much harder." The papers, on his reading, show Chinese military-linked researchers trying to move that expensive proprietary reasoning into small systems they can control and deploy themselves. There is an irony in the pattern. Beijing has been moving to restrict consumer access to Western AI models on security grounds, and has pushed its own labs away from US-designed chips, even as its military researchers mine the outputs of American systems. The asymmetry points at what is wrong with the US instrument. Export controls regulate objects: chips, tools, hardware that crosses a border and can be counted. The thing being transferred here is text a model produced- which crosses no border in any customs sense and cannot be enumerated. The White House has registered the problem at both the chip and distillation end, but registering it is not the same as having a lever. For now, the debate over whether distillation is theft or standard practice may matter less than the fact that neither answer offers a workable control. Follow TechRadar on Google News and add us as a preferred source to get our expert news, reviews, and opinion in your feeds.
[4]
Report claims China is distilling U.S. frontier models to power military AI applications
An exclusive report by Reuters today has surfaced evidence that suggests Chinese artificial intelligence firms have been leveraging the outputs of American frontier models developed by OpenAI Group PBC and Anthropic PBC to train their own AI systems for defense applications. The review by Reuters included a detailed examination of more than 80 academic papers and patent applications published by Chinese researchers, and found that multiple institutions linked to the military have been harvesting the outputs of U.S. models. The report claims that China is doing this to get around Washington's aggressive export controls, which have prevented its startups from accessing the advanced chips needed to create more sophisticated AI models from scratch. Reuters said it conducted the review with the Washington D.C.-based defense policy thinktank Jamestown Foundation. Its findings suggest that Chinese researchers are increasingly hijacking the work of America's best known AI startups in order to build advanced systems that power military applications. The report accuses Chinese researchers of misusing a technique known as "model distillation," which refers to a process that involves taking the outputs, reasoning steps and data from powerful "teacher models" and using them to train smaller, more specialized "student" models. Through model distillation, it's possible to create smaller models that can run more efficiently on less powerful hardware. The result is that China has been able to develop numerous models that match the performance of OpenAI's and Anthropic's most advanced frontier models despite using far less computational resources. Although distillation does have legitimate uses, such as in optimizing models for specific workloads, there's a lot of controversy around the technique, especially as Chinese models continue to surprise with their frontier-level capabilities. Last month, Alibaba Group Holding Ltd. launched a preview of Qwen3.8, releasing benchmarks that suggest its performance is second only to Anthropic's Claude Fable 5, surpassing OpenAI's best model, GPT-5.6. That came after the Chinese startup Moonshot AI made headlines when it released the world's largest open-weights model to date in Kimi K3, with similar claims about its performance. In response to those developments, U.S. officials and AI startups have cried foul, accusing Chinese firms of distilling American models to advance the capabilities of their own. U.S. Treasury Secretary Scott Bessent has even threatened to sanction any Chinese AI firms it deems guilty of the distillation practice, even though others, such as Microsoft Corp. Chief Executive Satya Nadella, have pointed out the irony of the accusations, given how Anthropic's models were trained on massive amounts of internet data without asking permission. Reuters provided a number of examples of how China has been distilling U.S. models. For instance, it claimed that PLA Unit 96941, the main cyberwarfare group in the People's Liberation Army, distilled OpenAI's GPT-3.5 to process and summarize sensitive military source code. From those outputs, the unit was able to develop a lightweight model that specializes in safely processing sensitive data on Chinese military networks. The report also points fingers at researchers at the North University of China, an institution believed to have close ties to the country's weapons manufacturing industry. It said they used Anthropic's Claude 3 Haiku to create synthetic training data to build a text classification system for social media monitoring. Meanwhile, at the National University of Defense Technology, researchers are alleged to have distilled a U.S.-made image processing model so it can be installed directly on unmanned aerial vehicles. Using this lightweight version of the model, drones can analyze live video feeds in real time to aid their navigation and weapons targeting systems. The revelations are likely to increase tensions ahead of upcoming talks between U.S. and Chinese officials on AI governance and safety. The White House is concerned that distillation undermines its export controls and infringes on the intellectual property rights of U.S. model makers, while Beijing accuses the Americans of pursuing "AI hegemonism." However, any concerns that China might use model distillation to one day surpass the U.S. in AI are likely to be unfounded, said SapienX Inc. co-founder Trevor Koverko. "It is best understood as transferring selected capabilities into a cheaper, locally controlled system," he explained. "It doesn't mean achieving independence from frontier AI." In other words, while Chinese researchers might be able to devise systems that can match the capabilities of U.S. models in a narrow range of applications, they'll still have to come up with their own original breakthroughs to outperform their U.S. rivals.
[5]
Exclusive-Chinese Military Researchers Tap US AI Models to Train Defence Systems
BEIJING, July 31 (Reuters) - Chinese military researchers have used outputs from leading U.S. artificial intelligence models developed by OpenAI and Anthropic to train domestic AI systems to advance China's defence capabilities, according to a Reuters review of more than 80 Chinese academic papers and patents. The previously unreported findings offer a rare glimpse into how military and security-linked institutions in China are leveraging cutting-edge U.S. AI models as a shortcut to developing specialised systems of their own, despite Washington's efforts to restrict Beijing's access to advanced chips and other strategic technologies. The documents show widespread use of a technique known as "model distillation," in which outputs from a powerful AI system are used to train smaller, specialised models that can be deployed locally without the enormous computing requirements needed to build frontier AI systems from scratch. Reuters' review, which included research compiled by the Washington-based Jamestown Foundation and shared exclusively with the news agency, showed distillation is widely used by researchers linked to the People's Liberation Army and other military institutions. The papers suggest Chinese defence institutions see leading U.S. AI models as both a source of technical insight and a way to close the gap with American rivals. The dispute centres on unauthorised extraction, not distillation itself, a widely used industry practice. The issue has emerged as a major flashpoint ahead of U.S.-China talks on AI governance and safety. U.S. officials have accused some Chinese entities of using distillation to extract capabilities from American AI models, potentially undermining export controls and infringing intellectual property rights. China has rejected the accusations, saying Washington is pursuing AI "hegemonism" while arguing that U.S. firms have engaged in similar practices. Chinese developers have also disputed claims that their AI advances rely on foreign models. AI startup Moonshot last week denied allegations by the Trump administration that its Kimi K3 model was built using distillation, saying it was driven by proprietary innovations. Sunny Cheung, a Jamestown fellow who analysed over 60 of the papers, said Chinese military scientists are systematically capturing the reasoning steps of Western models to adapt them for surveillance, cyber warfare and tactical decision-making. "Teaching a model the right answer is one thing but teaching it the reasoning behind the answer is much harder," said Cheung. "These papers show Chinese military-linked researchers are trying to transfer that expensive, proprietary reasoning from Western models into smaller systems they can control and deploy locally." Reuters verified the academic literature and identified an additional two dozen military-linked case studies. One paper published last year by researchers in PLA Unit 96941, a military intelligence and cyber-warfare unit in Beijing, described using OpenAI's GPT-3.5 to process sensitive military source code. The researchers said third-party models were unsuitable for handling classified information. To overcome that limitation, they used GPT-3.5 to summarize software code and trained a domestic model on those summaries to run entirely within Chinese military networks. The White House, Pentagon, China's foreign ministry, the PLA and OpenAI did not respond to requests for comment. WIDE USE, FROM MONITORING TO MILITARY Chinese researchers have used distillation for purposes ranging from content monitoring to military deployment, Reuters' and Jamestown's review of the papers showed. At the North University of China, which has close links to the country's weapons industry, researchers used Anthropic's Claude 3 Haiku to generate synthetic training data for a text classification model for social media monitoring and content moderation. Anthropic said it does not provide commercial access to Claude in China or to Beijing-controlled firms and uses monitoring systems to detect policy violations. The company added that distilled models may lose the original systems' safety safeguards, potentially allowing sensitive capabilities to be transferred to models beyond its control. A 2024 paper from the PLA's National University of Defense Technology described using distillation to shrink an image-processing model for deployment on unmanned aerial vehicles, allowing drones to analyse live video and support navigation and targeting decisions in real time even when communications are cut. Similarly, researchers at China's Academy of Military Sciences used distillation to run a target-recognition model on tactical hardware during simulated maritime operations involving drones, ships and unmanned submarines, a study published earlier this year showed. BENEFITS AND LIMITS China has embraced distillation as it seeks to compete with the U.S. in frontier AI while facing constraints on advanced computing resources due to Washington's export controls on high-end chips. Central and local governments have promoted "model lightweighting" and edge computing, directing subsidies and research funding toward technologies that enable AI models to run on drones, satellites and other devices with limited processing power. Experts caution, however, that distillation has significant limitations. As Chinese AI models close the gap with their U.S. counterparts, military researchers are also examining distillation as a potential security risk. In January, researchers at the Army Engineering University published a paper on the threat of "data-free distillation," a method of reverse-engineering a model's capabilities without direct access to its core parameters. To counter that vulnerability, they proposed defense mechanisms designed to mask the hidden logical information exposed in a model's public outputs. Distilled models also inherit only selected capabilities and cannot fully replicate the broad intelligence of frontier systems. Trevor Koverko, co-founder of AI data company Sapien, said distilled models remain less capable than their teacher models. "It is best understood as transferring selected capabilities into a cheaper, locally controlled system, not achieving independence from frontier AI." (Reporting by Eduardo Baptista; Editing by Miyoung Kim and Shri Navaratnam)
[6]
What is AI model distillation and why is it becoming a US-China flashpoint?
Long regarded as a standard tool of AI research, distillation is now at the centre of a growing dispute over whether advanced AI capabilities can be transferred without the consent of the companies that created them. Washington and leading US AI firms have accused Chinese rivals of using the technique to extract capabilities from proprietary models, opening a new front in an increasingly bitter competition over technological leadership. A technique that allows developers to shrink powerful artificial intelligence models into cheaper, more efficient systems has become the latest battleground in the intensifying US-China race for AI dominance. Known as model distillation, the method uses the outputs of a powerful AI system to train a smaller model that can perform some of the same tasks with fewer computing resources. Long regarded as a standard tool of AI research, distillation is now at the centre of a growing dispute over whether advanced AI capabilities can be transferred without the consent of the companies that created them. Washington and leading US AI firms have accused Chinese rivals of using the technique to extract capabilities from proprietary models, opening a new front in an increasingly bitter competition over technological leadership. Below are key facts about the practice at the centre of the debate. What is model distillation? The largest AI models, known as frontier models, require enormous amounts of computing power, data and investment to train. Model distillation offers a way to create smaller systems by using a large "teacher" model to train a smaller "student" model. The teacher generates examples, such as answers and computer code, which are then used as training material for the student. The smaller model is not a replica of the teacher. It does not inherit the teacher's weights, architecture or full capabilities. Instead, it learns selected behaviours that enable it to perform specific tasks more efficiently. Why does distillation matter? The appeal of distillation is that it can make AI cheaper and easier to deploy. A frontier model may require large data centres and expensive chips to operate. Distilled models can run on less powerful hardware and be tailored for specific tasks. That makes them appealing to companies and governments looking to deploy AI more widely, from devices and factories to vehicles and private networks. Why are reasoning traces important Recent AI systems have increased interest in transferring not only final answers but also the steps used to reach them. These "reasoning traces" can show a smaller model how to approach a difficult problem rather than simply what answer to produce. Florian Tramer, an assistant professor at ETH Zurich who researches machine-learning security, compared the process to human learning. "If I give you a book of complicated math problems with final solutions, you will have a much harder time learning how to solve problems than if I gave you detailed solutions that describe all steps to take," he said. As reasoning traces have become more valuable, access to AI outputs has become more sensitive because they may expose some of the methods advanced systems use to tackle complex problems. Who uses distillation Distillation is a widely used AI training technique, not an inherently improper practice. US researchers and companies have long used it, including Stanford University's Alpaca project and Microsoft's Orca research, which relied on outputs from more advanced models to improve smaller ones. Chinese researchers have also used outputs from US models in public research projects, including efforts to create Chinese-language instruction models. The key difference lies in access. Open-weight models let researchers inspect and modify underlying parameters. Closed models, such as OpenAI's ChatGPT and Anthropic's Claude, remain under company control and are typically accessed through proprietary interfaces or APIs. Why has distillation become a US-China issue? The controversy is less over distillation itself and more about unauthorised extraction. AI companies argue there is a distinction between legitimate research and systematically harvesting outputs from proprietary models to replicate commercially valuable capabilities. Anthropic has accused Chinese entities including DeepSeek, Moonshot and MiniMax of conducting large-scale campaigns to obtain capabilities from Claude models. The company said those efforts targeted capabilities including software engineering and advanced reasoning. OpenAI has also said it has detected attempts by Chinese actors to use its models for distillation-related purposes. No Chinese companies have accused US rivals of distilling closed-source models so far.
[7]
Chinese military researchers tap U.S. AI models to train defense systems
Beijing - Chinese military researchers have used outputs from leading U.S. artificial intelligence models developed by OpenAI and Anthropic to train domestic AI systems to advance China's defense capabilities, according to a review of more than 80 Chinese academic papers and patents. The previously unreported findings offer a rare glimpse into how military and security-linked institutions in China are leveraging cutting-edge U.S. AI models as a shortcut to developing specialized systems of their own, despite Washington's efforts to restrict Beijing's access to advanced chips and other strategic technologies. The documents show widespread use of a technique known as "model distillation," in which outputs from a powerful AI system are used to train smaller, specialized models that can be deployed locally without the enormous computing requirements needed to build frontier AI systems from scratch. The review of the papers and patents, which included research compiled by the Washington-based Jamestown Foundation, showed distillation is widely used by researchers linked to the People's Liberation Army and other military institutions.
[8]
What is model distillation? The AI technique at the centre of the US-China technology battle
Model distillation, a technique that trains smaller AI models using outputs from more powerful systems, has become a key issue in the US-China AI rivalry. While it lowers AI costs and expands deployment, US companies allege some Chinese firms are using the method to extract capabilities from proprietary models without authorisation. Model distillation, a long-established artificial intelligence (AI) training technique, has emerged as the latest flashpoint in the growing technology rivalry between the United States and China. While the method helps developers build smaller, cheaper AI models, US companies have accused some Chinese firms of using it to extract capabilities from proprietary AI systems without permission. Here is what model distillation is and why it has become controversial. What is model distillation?Training the world's most advanced AI models, often called frontier models, requires massive computing power, data and investment. Model distillation is a process in which a powerful AI system, known as the "teacher" model, helps train a smaller "student" model. Instead of copying the original model, the teacher generates outputs such as answers, explanations or computer code. These outputs then become training material for the smaller model. The student model does not inherit the teacher's architecture, parameters or full capabilities. Instead, it learns selected behaviours that allow it to perform specific tasks more efficiently. Why is it important?The biggest advantage of model distillation is lower cost. Large AI models often require expensive chips and large data centres to operate. Distilled models can run on less powerful hardware, making them easier to deploy across smartphones, factories, vehicles and private enterprise networks. This allows businesses and governments to use AI more widely without bearing the cost of running frontier models. What are reasoning traces?The latest generation of AI systems has increased interest in transferring not just the final answer but also the reasoning process used to reach it. These "reasoning traces" show how an AI system solves a problem step by step, helping a smaller model learn the approach instead of simply memorising answers. Florian Tramèr, assistant professor at ETH Zurich, compared it to learning mathematics. "If I give you a book of complicated math problems with final solutions, you will have a much harder time learning how to solve problems than if I gave you detailed solutions that describe all steps to take," Tramèr told Reuters. As reasoning traces become more valuable, companies have become more protective of AI outputs because they may reveal how advanced systems solve complex tasks. Who uses model distillation?Distillation is widely used across the AI industry and is not considered improper by itself. Researchers and companies in the US have used it for years, including Stanford University's Alpaca project and Microsoft's Orca research, which relied on outputs from more advanced models to improve smaller AI systems. Chinese researchers have also used outputs from US models in public research, including projects aimed at developing Chinese-language instruction models. A key distinction is between open-weight and closed AI models. Open-weight models allow researchers to inspect and modify their parameters. Closed models such as OpenAI's ChatGPT and Anthropic's Claude remain under company control and are generally accessed through proprietary application programming interfaces (APIs). Why has it become a US-China issue?The dispute is not over model distillation itself but over whether companies can use outputs from proprietary AI systems without permission. US AI companies argue there is a difference between legitimate research and systematically collecting outputs from closed models to reproduce commercially valuable capabilities. Anthropic has accused Chinese companies, including DeepSeek, Moonshot and MiniMax, of running large-scale campaigns to obtain capabilities from its Claude models, including software engineering and advanced reasoning functions. OpenAI has also said it detected attempts by Chinese actors to use its models for distillation-related purposes. According to Reuters, no Chinese company has publicly accused US AI firms of distilling capabilities from closed-source models.
[9]
Chinese military researchers tap US AI models to train defence systems
Chinese military researchers are using U.S. AI model outputs to advance defense capabilities. This technique allows them to train specialized domestic AI systems efficiently. Researchers are leveraging powerful U.S. models as a shortcut for their own development. This practice is being used for surveillance, cyber warfare, and tactical decision-making. The findings offer a rare glimpse into China's AI development strategies. Beijing: Chinese military researchers have used outputs from leading U.S. artificial intelligence models developed by OpenAI and Anthropic to train domestic AI systems to advance China's defence capabilities, according to a Reuters review of more than 80 Chinese academic papers and patents. The previously unreported findings offer a rare glimpse into how military and security-linked institutions in China are leveraging cutting-edge U.S. AI models as a shortcut to developing specialised systems of their own, despite Washington's efforts to restrict Beijing's access to advanced chips and other strategic technologies. The documents show widespread use of a technique known as "model distillation," in which outputs from a powerful AI system are used to train smaller, specialised models that can be deployed locally without the enormous computing requirements needed to build frontier AI systems from scratch. Reuters' review, which included research compiled by the Washington-based Jamestown Foundation and shared exclusively with the news agency, showed distillation is widely used by researchers linked to the People's Liberation Army and other military institutions. The papers suggest Chinese defence institutions see leading U.S. AI models as both a source of technical insight and a way to close the gap with American rivals. The dispute centres on unauthorised extraction, not distillation itself, a widely used industry practice. The issue has emerged as a major flashpoint ahead of U.S.-China talks on AI governance and safety. U.S. officials have accused some Chinese entities of using distillation to extract capabilities from American AI models, potentially undermining export controls and infringing intellectual property rights. China has rejected the accusations, saying Washington is pursuing AI "hegemonism" while arguing that U.S. firms have engaged in similar practices. Chinese developers have also disputed claims that their AI advances rely on foreign models. AI startup Moonshot last week denied allegations by the Trump administration that its Kimi K3 model was built using distillation, saying it was driven by proprietary innovations. Sunny Cheung, a Jamestown fellow who analysed over 60 of the papers, said Chinese military scientists are systematically capturing the reasoning steps of Western models to adapt them for surveillance, cyber warfare and tactical decision-making. "Teaching a model the right answer is one thing but teaching it the reasoning behind the answer is much harder," said Cheung. "These papers show Chinese military-linked researchers are trying to transfer that expensive, proprietary reasoning from Western models into smaller systems they can control and deploy locally." Reuters verified the academic literature and identified an additional two dozen military-linked case studies. One paper published last year by researchers in PLA Unit 96941, a military intelligence and cyber-warfare unit in Beijing, described using OpenAI's GPT-3.5 to process sensitive military source code. The researchers said third-party models were unsuitable for handling classified information. To overcome that limitation, they used GPT-3.5 to summarize software code and trained a domestic model on those summaries to run entirely within Chinese military networks. The White House, Pentagon, China's foreign ministry, the PLA and OpenAI did not respond to requests for comment. WIDE USE, FROM MONITORING TO MILITARY Chinese researchers have used distillation for purposes ranging from content monitoring to military deployment, Reuters' and Jamestown's review of the papers showed. At the North University of China, which has close links to the country's weapons industry, researchers used Anthropic's Claude 3 Haiku to generate synthetic training data for a text classification model for social media monitoring and content moderation. Anthropic said it does not provide commercial access to Claude in China or to Beijing-controlled firms and uses monitoring systems to detect policy violations. The company added that distilled models may lose the original systems' safety safeguards, potentially allowing sensitive capabilities to be transferred to models beyond its control. A 2024 paper from the PLA's National University of Defense Technology described using distillation to shrink an image-processing model for deployment on unmanned aerial vehicles, allowing drones to analyse live video and support navigation and targeting decisions in real time even when communications are cut. Similarly, researchers at China's Academy of Military Sciences used distillation to run a target-recognition model on tactical hardware during simulated maritime operations involving drones, ships and unmanned submarines, a study published earlier this year showed. BENEFITS AND LIMITS China has embraced distillation as it seeks to compete with the U.S. in frontier AI while facing constraints on advanced computing resources due to Washington's export controls on high-end chips. Central and local governments have promoted "model lightweighting" and edge computing, directing subsidies and research funding toward technologies that enable AI models to run on drones, satellites and other devices with limited processing power. Experts caution, however, that distillation has significant limitations. As Chinese AI models close the gap with their U.S. counterparts, military researchers are also examining distillation as a potential security risk. In January, researchers at the Army Engineering University published a paper on the threat of "data-free distillation," a method of reverse-engineering a model's capabilities without direct access to its core parameters. To counter that vulnerability, they proposed defense mechanisms designed to mask the hidden logical information exposed in a model's public outputs. Distilled models also inherit only selected capabilities and cannot fully replicate the broad intelligence of frontier systems. Trevor Koverko, co-founder of AI data company Sapien, said distilled models remain less capable than their teacher models. "It is best understood as transferring selected capabilities into a cheaper, locally controlled system, not achieving independence from frontier AI."
[10]
Chinese military researchers tap US AI models to train defence systems
BEIJING, July 31 (Reuters) - Chinese military researchers have used outputs from leading U.S. artificial intelligence models developed by OpenAI and Anthropic to train domestic AI systems to advance China's defence capabilities, according to a Reuters review of more than 80 Chinese academic papers and patents. The previously unreported findings offer a rare glimpse into how military and security-linked institutions in China are leveraging cutting-edge U.S. AI models as a shortcut to developing specialised systems of their own, despite Washington's efforts to restrict Beijing's access to advanced chips and other strategic technologies. The documents show widespread use of a technique known as "model distillation," in which outputs from a powerful AI system are used to train smaller, specialised models that can be deployed locally without the enormous computing requirements needed to build frontier AI systems from scratch. Reuters' review, which included research compiled by the Washington-based Jamestown Foundation and shared exclusively with the news agency, showed distillation is widely used by researchers linked to the People's Liberation Army and other military institutions. The papers suggest Chinese defence institutions see leading U.S. AI models as both a source of technical insight and a way to close the gap with American rivals. The dispute centres on unauthorised extraction, not distillation itself, a widely used industry practice. The issue has emerged as a major flashpoint ahead of U.S.-China talks on AI governance and safety. U.S. officials have accused some Chinese entities of using distillation to extract capabilities from American AI models, potentially undermining export controls and infringing intellectual property rights. China has rejected the accusations, saying Washington is pursuing AI "hegemonism" while arguing that U.S. firms have engaged in similar practices. Chinese developers have also disputed claims that their AI advances rely on foreign models. AI startup Moonshot last week denied allegations by the Trump administration that its Kimi K3 model was built using distillation, saying it was driven by proprietary innovations. Sunny Cheung, a Jamestown fellow who analysed over 60 of the papers, said Chinese military scientists are systematically capturing the reasoning steps of Western models to adapt them for surveillance, cyber warfare and tactical decision-making. "Teaching a model the right answer is one thing but teaching it the reasoning behind the answer is much harder," said Cheung. "These papers show Chinese military-linked researchers are trying to transfer that expensive, proprietary reasoning from Western models into smaller systems they can control and deploy locally." Reuters verified the academic literature and identified an additional two dozen military-linked case studies. One paper published last year by researchers in PLA Unit 96941, a military intelligence and cyber-warfare unit in Beijing, described using OpenAI's GPT-3.5 to process sensitive military source code. The researchers said third-party models were unsuitable for handling classified information. To overcome that limitation, they used GPT-3.5 to summarize software code and trained a domestic model on those summaries to run entirely within Chinese military networks. The White House, Pentagon, China's foreign ministry, the PLA and OpenAI did not respond to requests for comment. WIDE USE, FROM MONITORING TO MILITARY Chinese researchers have used distillation for purposes ranging from content monitoring to military deployment, Reuters' and Jamestown's review of the papers showed. At the North University of China, which has close links to the country's weapons industry, researchers used Anthropic's Claude 3 Haiku to generate synthetic training data for a text classification model for social media monitoring and content moderation. Anthropic said it does not provide commercial access to Claude in China or to Beijing-controlled firms and uses monitoring systems to detect policy violations. The company added that distilled models may lose the original systems' safety safeguards, potentially allowing sensitive capabilities to be transferred to models beyond its control. A 2024 paper from the PLA's National University of Defense Technology described using distillation to shrink an image-processing model for deployment on unmanned aerial vehicles, allowing drones to analyse live video and support navigation and targeting decisions in real time even when communications are cut. Similarly, researchers at China's Academy of Military Sciences used distillation to run a target-recognition model on tactical hardware during simulated maritime operations involving drones, ships and unmanned submarines, a study published earlier this year showed. BENEFITS AND LIMITS China has embraced distillation as it seeks to compete with the U.S. in frontier AI while facing constraints on advanced computing resources due to Washington's export controls on high-end chips. Central and local governments have promoted "model lightweighting" and edge computing, directing subsidies and research funding toward technologies that enable AI models to run on drones, satellites and other devices with limited processing power. Experts caution, however, that distillation has significant limitations. As Chinese AI models close the gap with their U.S. counterparts, military researchers are also examining distillation as a potential security risk. In January, researchers at the Army Engineering University published a paper on the threat of "data-free distillation," a method of reverse-engineering a model's capabilities without direct access to its core parameters. To counter that vulnerability, they proposed defense mechanisms designed to mask the hidden logical information exposed in a model's public outputs. Distilled models also inherit only selected capabilities and cannot fully replicate the broad intelligence of frontier systems. Trevor Koverko, co-founder of AI data company Sapien, said distilled models remain less capable than their teacher models. "It is best understood as transferring selected capabilities into a cheaper, locally controlled system, not achieving independence from frontier AI." (Reporting by Eduardo Baptista; Editing by Miyoung Kim and Shri Navaratnam)
Share
Copy Link
A Reuters investigation uncovered that Chinese military researchers have systematically used outputs from leading US AI models developed by OpenAI and Anthropic to train domestic defence systems. The review of over 80 academic papers and patents reveals widespread use of AI model distillation by People's Liberation Army-linked institutions, exposing a critical gap in US export controls and escalating US-China geopolitical tensions ahead of AI governance talks.

Chinese military researchers have systematically leveraged outputs from frontier AI models developed by OpenAI and Anthropic to train domestic defence systems, according to a Reuters investigation examining more than 80 Chinese academic papers and patents
1
2
. The findings, compiled with research from the Washington-based Jamestown Foundation, reveal how People's Liberation Army-linked institutions are using AI model distillation as a shortcut to develop specialized military applications despite US export controls restricting access to advanced chips1
.The investigation offers rare insight into how Chinese defence institutions view US AI models as both a source of technical knowledge and a mechanism to close the capability gap with American rivals
5
. Reuters independently verified the academic literature and identified an additional two dozen military-linked case studies beyond the Jamestown Foundation's initial analysis2
.AI model distillation involves using outputs from a powerful teacher model to train a smaller student model that can perform specific tasks with significantly fewer computing resources
1
. The technique has become particularly valuable because it transfers not just final answers but reasoning traces—the step-by-step problem-solving approaches that advanced systems use to tackle complex challenges1
.Florian Tramèr, an assistant professor at ETH Zurich specializing in machine-learning security, explained the distinction: "If I give you a book of complicated math problems with final solutions, you will have a much harder time learning how to solve problems than if I gave you detailed solutions that describe all steps to take"
1
. This capability to transfer reasoning rather than just outputs makes distillation particularly powerful for military applications2
.Sunny Cheung, the Jamestown fellow who analyzed over 60 papers, emphasized that Chinese military scientists are systematically capturing reasoning steps from Western models to adapt them for surveillance, cyber warfare and tactical decision-making
5
. "Teaching a model the right answer is one thing but teaching it the reasoning behind the answer is much harder," Cheung noted, adding that the papers show researchers transferring expensive, proprietary reasoning from Western models into smaller systems they can control and deploy locally2
.The Reuters review documented specific military applications across multiple Chinese defence institutions. PLA Unit 96941, a military intelligence and cyber-warfare unit in Beijing, published a paper describing how researchers used OpenAI GPT-3.5 to process sensitive military source code
2
5
. The researchers acknowledged that third-party models were unsuitable for classified information, so they used GPT-3.5 to summarize software code and trained a domestic model on those summaries to run entirely within Chinese military networks1
.At the North University of China, which maintains close links to the country's weapons industry, researchers used Anthropic Claude 3 Haiku to generate synthetic training data for a text classification model designed for social media monitoring and content moderation
5
. Anthropic responded by stating it does not provide commercial access to Claude in China or to Beijing-controlled firms and uses monitoring systems to detect policy violations1
.A 2024 paper from the PLA's National University of Defense Technology detailed using distillation to shrink an image-processing model for deployment on unmanned aerial vehicles, enabling drone navigation and target recognition in real time even when communications are severed
5
. Similarly, researchers at China's Academy of Military Sciences used distillation to run a target-recognition model on tactical hardware during simulated maritime operations involving drones, ships and unmanned submarines1
.The strategic significance of AI model distillation lies in how it sidesteps the logic of US export controls, which restrict the advanced chips needed to train frontier AI models but cannot restrict the text outputs those models produce
3
. Washington's export controls regulate physical objects—chips, tools and hardware that cross borders—but the outputs being transferred through distillation are text that a model produced, which crosses no border in any customs sense3
.This represents a fundamental gap in the US regulatory framework. Export controls are built to deny China the chips needed to train frontier models, but distillation eliminates that requirement because the expensive computational work has already been completed by American companies
3
. The technique allows Chinese researchers to create smaller models that inherit selected behaviors at a fraction of the compute cost and can run on modest local hardware4
.Related Stories
The distillation controversy has emerged as a major flashpoint ahead of US-China talks on AI governance and safety
1
5
. US officials and leading AI firms have accused Chinese entities including DeepSeek, MiniMax and Moonshot of conducting large-scale campaigns to obtain capabilities from proprietary models1
. Anthropic specifically accused these entities of targeting capabilities including software engineering and advanced reasoning from its Claude models1
.China has rejected the accusations, characterizing Washington's position as pursuing AI hegemonism while arguing that US firms have engaged in similar practices
1
5
. Chinese developers have disputed claims that their advances rely on foreign models. AI startup Moonshot denied allegations by the Trump administration that its Kimi K3 model was built using distillation, stating it was driven by proprietary innovations2
5
.US Treasury Secretary Scott Bessent has threatened to sanction Chinese AI firms deemed guilty of unauthorized distillation, though critics point out the irony given that Anthropic's models were trained on massive amounts of internet data without permission
4
. The dispute centers on unauthorized extraction rather than distillation itself, which remains a widely used industry practice employed by US researchers and companies including Stanford University's Alpaca project and Microsoft's Orca research1
.While distillation enables Chinese researchers to match US model capabilities in narrow applications, experts suggest it does not provide a path to surpassing American AI leadership. SapienX co-founder Trevor Koverko characterized distillation as "transferring selected capabilities into a cheaper, locally controlled system" rather than achieving independence from frontier AI
4
. Chinese researchers will still need original breakthroughs to outperform US rivals4
.The White House has registered the problem at both the chip and distillation levels, but identifying the issue differs from having effective control mechanisms
3
. The debate over whether distillation constitutes theft or standard practice may matter less than the fact that neither characterization offers a workable regulatory solution3
.Anthropic warned that distilled models may lose the original systems' safety safeguards, potentially allowing sensitive capabilities to be transferred to models beyond its control
5
. This safety dimension adds another layer of concern beyond the geopolitical and intellectual property issues. The White House, Pentagon, China's foreign ministry, the PLA and OpenAI all declined to comment on the Reuters findings2
5
.Summarized by
Navi
[4]
13 Jul 2026•Policy and Regulation

03 Jul 2026•Policy and Regulation

23 Feb 2026•Technology

1
Technology

2
Policy and Regulation

3
Policy and Regulation
