7 Sources
[1]
Explainer: What is AI model distillation and why is it becoming a US-China flashpoint?
BEIJING, July 31 (Reuters) - A technique that allows developers to shrink powerful artificial intelligence models into cheaper, more efficient systems has become the latest battleground in the intensifying U.S.-China race for AI dominance. Known as model distillation, the method uses the outputs of a powerful AI system to train a smaller model that can perform some of the same tasks with fewer computing resources. Long regarded as a standard tool of AI research, distillation is now at the centre of a growing dispute over whether advanced AI capabilities can be transferred without the consent of the companies that created them. Washington and leading U.S. AI firms have accused Chinese rivals of using the technique to extract capabilities from proprietary models, opening a new front in an increasingly bitter competition over technological leadership. Below are key facts about the practice at the centre of the debate. WHAT IS MODEL DISTILLATION? The largest AI models, known as frontier models, require enormous amounts of computing power, data and investment to train. Model distillation offers a way to create smaller systems by using a large "teacher" model to train a smaller "student" model. The teacher generates examples, such as answers and computer code, which are then used as training material for the student. The smaller model is not a replica of the teacher. It does not inherit the teacher's weights, architecture or full capabilities. Instead, it learns selected behaviours that enable it to perform specific tasks more efficiently. WHY DOES DISTILLATION MATTER? The appeal of distillation is that it can make AI cheaper and easier to deploy. A frontier model may require large data centres and expensive chips to operate. Distilled models can run on less powerful hardware and be tailored for specific tasks. That makes them appealing to companies and governments looking to deploy AI more widely, from devices and factories to vehicles and private networks. WHY ARE REASONING TRACES IMPORTANT? Recent AI systems have increased interest in transferring not only final answers but also the steps used to reach them. These "reasoning traces" can show a smaller model how to approach a difficult problem rather than simply what answer to produce. Florian Tramèr, an assistant professor at ETH Zurich who researches machine-learning security, compared the process to human learning. "If I give you a book of complicated math problems with final solutions, you will have a much harder time learning how to solve problems than if I gave you detailed solutions that describe all steps to take," he said. As reasoning traces have become more valuable, access to AI outputs has become more sensitive because they may expose some of the methods advanced systems use to tackle complex problems. WHO USES DISTILLATION? Distillation is a widely used AI training technique, not an inherently improper practice. U.S. researchers and companies have long used it, including Stanford University's Alpaca project and Microsoft's Orca research, which relied on outputs from more advanced models to improve smaller ones. Chinese researchers have also used outputs from U.S. models in public research projects, including efforts to create Chinese-language instruction models. The key difference lies in access. Open-weight models let researchers inspect and modify underlying parameters. Closed models, such as OpenAI's ChatGPT and Anthropic's Claude, remain under company control and are typically accessed through proprietary interfaces or APIs. WHY HAS DISTILLATION BECOME A U.S.-CHINA ISSUE? The controversy is less over distillation itself and more about unauthorised extraction. AI companies argue there is a distinction between legitimate research and systematically harvesting outputs from proprietary models to replicate commercially valuable capabilities. Anthropic has accused Chinese entities including DeepSeek, Moonshot and MiniMax of conducting large-scale campaigns to obtain capabilities from Claude models. The company said those efforts targeted capabilities including software engineering and advanced reasoning. OpenAI has also said it has detected attempts by Chinese actors to use its models for distillation-related purposes. No Chinese companies have accused U.S. rivals of distilling closed-source models so far. Reporting by Eduardo Baptista Editing by Shri Navaratnam Our Standards: The Thomson Reuters Trust Principles., opens new tab * Suggested Topics: * Artificial Intelligence Eduardo Baptista Thomson Reuters Eduardo Baptista is a Senior Correspondent for Reuters based in Beijing, covering China's technology, space, and automotive industries. He has led enterprise and investigative reporting on China's military-linked companies, artificial intelligence and semiconductor supply chains, as well as macroeconomic and industrial policy. Baptista has reported from China for nearly a decade and holds a BA in History from the University of Cambridge.
[2]
EXCLUSIVE: Chinese military researchers tap US AI models to train defence systems
BEIJING, July 31 (Reuters) - Chinese military researchers have used outputs from leading U.S. artificial intelligence models developed by OpenAI and Anthropic to train domestic AI systems to advance China's defence capabilities, according to a Reuters review of more than 80 Chinese academic papers and patents. The previously unreported findings offer a rare glimpse into how military and security-linked institutions in China are leveraging cutting-edge U.S. AI models as a shortcut to developing specialised systems of their own, despite Washington's efforts to restrict Beijing's access to advanced chips and other strategic technologies. The documents show widespread use of a technique known as "model distillation," in which outputs from a powerful AI system are used to train smaller, specialised models that can be deployed locally without the enormous computing requirements needed to build frontier AI systems from scratch. Reuters' review, which included research compiled by the Washington-based Jamestown Foundation and shared exclusively with the news agency, showed distillation is widely used by researchers linked to the People's Liberation Army and other military institutions. The papers suggest Chinese defence institutions see leading U.S. AI models as both a source of technical insight and a way to close the gap with American rivals. The dispute centres on unauthorised extraction, not distillation itself, a widely used industry practice. The issue has emerged as a major flashpoint ahead of U.S.-China talks on AI governance and safety. U.S. officials have accused some Chinese entities of using distillation to extract capabilities from American AI models, potentially undermining export controls and infringing intellectual property rights. China has rejected the accusations, saying Washington is pursuing AI "hegemonism" while arguing that U.S. firms have engaged in similar practices. Chinese developers have also disputed claims that their AI advances rely on foreign models. AI startup Moonshot last week denied allegations by the Trump administration that its Kimi K3 model was built using distillation, saying it was driven by proprietary innovations. Sunny Cheung, a Jamestown fellow who analysed over 60 of the papers, said Chinese military scientists are systematically capturing the reasoning steps of Western models to adapt them for surveillance, cyber warfare and tactical decision-making. "Teaching a model the right answer is one thing but teaching it the reasoning behind the answer is much harder," said Cheung. "These papers show Chinese military-linked researchers are trying to transfer that expensive, proprietary reasoning from Western models into smaller systems they can control and deploy locally." Reuters verified the academic literature and identified an additional two dozen military-linked case studies. One paper published last year by researchers in PLA Unit 96941, a military intelligence and cyber-warfare unit in Beijing, described using OpenAI's GPT-3.5 to process sensitive military source code. The researchers said third-party models were unsuitable for handling classified information. To overcome that limitation, they used GPT-3.5 to summarize software code and trained a domestic model on those summaries to run entirely within Chinese military networks. The White House, Pentagon, China's foreign ministry, the PLA and OpenAI did not respond to requests for comment. WIDE USE, FROM MONITORING TO MILITARY Chinese researchers have used distillation for purposes ranging from content monitoring to military deployment, Reuters' and Jamestown's review of the papers showed. At the North University of China, which has close links to the country's weapons industry, researchers used Anthropic's Claude 3 Haiku to generate synthetic training data for a text classification model for social media monitoring and content moderation. Anthropic said it does not provide commercial access to Claude in China or to Beijing-controlled firms and uses monitoring systems to detect policy violations. The company added that distilled models may lose the original systems' safety safeguards, potentially allowing sensitive capabilities to be transferred to models beyond its control. A 2024 paper from the PLA's National University of Defense Technology described using distillation to shrink an image-processing model for deployment on unmanned aerial vehicles, allowing drones to analyse live video and support navigation and targeting decisions in real time even when communications are cut. Similarly, researchers at China's Academy of Military Sciences used distillation to run a target-recognition model on tactical hardware during simulated maritime operations involving drones, ships and unmanned submarines, a study published earlier this year showed. BENEFITS AND LIMITS China has embraced distillation as it seeks to compete with the U.S. in frontier AI while facing constraints on advanced computing resources due to Washington's export controls on high-end chips. Central and local governments have promoted "model lightweighting" and edge computing, directing subsidies and research funding toward technologies that enable AI models to run on drones, satellites and other devices with limited processing power. Experts caution, however, that distillation has significant limitations. As Chinese AI models close the gap with their U.S. counterparts, military researchers are also examining distillation as a potential security risk. In January, researchers at the Army Engineering University published a paper on the threat of "data-free distillation," a method of reverse-engineering a model's capabilities without direct access to its core parameters. To counter that vulnerability, they proposed defense mechanisms designed to mask the hidden logical information exposed in a model's public outputs. Distilled models also inherit only selected capabilities and cannot fully replicate the broad intelligence of frontier systems. Trevor Koverko, co-founder of AI data company Sapien, said distilled models remain less capable than their teacher models. "It is best understood as transferring selected capabilities into a cheaper, locally controlled system, not achieving independence from frontier AI." Reporting by Eduardo Baptista; Editing by Miyoung Kim and Shri Navaratnam Our Standards: The Thomson Reuters Trust Principles., opens new tab * Suggested Topics: * Disrupted Eduardo Baptista Thomson Reuters Eduardo Baptista is a Senior Correspondent for Reuters based in Beijing, covering China's technology, space, and automotive industries. He has led enterprise and investigative reporting on China's military-linked companies, artificial intelligence and semiconductor supply chains, as well as macroeconomic and industrial policy. Baptista has reported from China for nearly a decade and holds a BA in History from the University of Cambridge.
[3]
Exclusive-Chinese Military Researchers Tap US AI Models to Train Defence Systems
BEIJING, July 31 (Reuters) - Chinese military researchers have used outputs from leading U.S. artificial intelligence models developed by OpenAI and Anthropic to train domestic AI systems to advance China's defence capabilities, according to a Reuters review of more than 80 Chinese academic papers and patents. The previously unreported findings offer a rare glimpse into how military and security-linked institutions in China are leveraging cutting-edge U.S. AI models as a shortcut to developing specialised systems of their own, despite Washington's efforts to restrict Beijing's access to advanced chips and other strategic technologies. The documents show widespread use of a technique known as "model distillation," in which outputs from a powerful AI system are used to train smaller, specialised models that can be deployed locally without the enormous computing requirements needed to build frontier AI systems from scratch. Reuters' review, which included research compiled by the Washington-based Jamestown Foundation and shared exclusively with the news agency, showed distillation is widely used by researchers linked to the People's Liberation Army and other military institutions. The papers suggest Chinese defence institutions see leading U.S. AI models as both a source of technical insight and a way to close the gap with American rivals. The dispute centres on unauthorised extraction, not distillation itself, a widely used industry practice. The issue has emerged as a major flashpoint ahead of U.S.-China talks on AI governance and safety. U.S. officials have accused some Chinese entities of using distillation to extract capabilities from American AI models, potentially undermining export controls and infringing intellectual property rights. China has rejected the accusations, saying Washington is pursuing AI "hegemonism" while arguing that U.S. firms have engaged in similar practices. Chinese developers have also disputed claims that their AI advances rely on foreign models. AI startup Moonshot last week denied allegations by the Trump administration that its Kimi K3 model was built using distillation, saying it was driven by proprietary innovations. Sunny Cheung, a Jamestown fellow who analysed over 60 of the papers, said Chinese military scientists are systematically capturing the reasoning steps of Western models to adapt them for surveillance, cyber warfare and tactical decision-making. "Teaching a model the right answer is one thing but teaching it the reasoning behind the answer is much harder," said Cheung. "These papers show Chinese military-linked researchers are trying to transfer that expensive, proprietary reasoning from Western models into smaller systems they can control and deploy locally." Reuters verified the academic literature and identified an additional two dozen military-linked case studies. One paper published last year by researchers in PLA Unit 96941, a military intelligence and cyber-warfare unit in Beijing, described using OpenAI's GPT-3.5 to process sensitive military source code. The researchers said third-party models were unsuitable for handling classified information. To overcome that limitation, they used GPT-3.5 to summarize software code and trained a domestic model on those summaries to run entirely within Chinese military networks. The White House, Pentagon, China's foreign ministry, the PLA and OpenAI did not respond to requests for comment. WIDE USE, FROM MONITORING TO MILITARY Chinese researchers have used distillation for purposes ranging from content monitoring to military deployment, Reuters' and Jamestown's review of the papers showed. At the North University of China, which has close links to the country's weapons industry, researchers used Anthropic's Claude 3 Haiku to generate synthetic training data for a text classification model for social media monitoring and content moderation. Anthropic said it does not provide commercial access to Claude in China or to Beijing-controlled firms and uses monitoring systems to detect policy violations. The company added that distilled models may lose the original systems' safety safeguards, potentially allowing sensitive capabilities to be transferred to models beyond its control. A 2024 paper from the PLA's National University of Defense Technology described using distillation to shrink an image-processing model for deployment on unmanned aerial vehicles, allowing drones to analyse live video and support navigation and targeting decisions in real time even when communications are cut. Similarly, researchers at China's Academy of Military Sciences used distillation to run a target-recognition model on tactical hardware during simulated maritime operations involving drones, ships and unmanned submarines, a study published earlier this year showed. BENEFITS AND LIMITS China has embraced distillation as it seeks to compete with the U.S. in frontier AI while facing constraints on advanced computing resources due to Washington's export controls on high-end chips. Central and local governments have promoted "model lightweighting" and edge computing, directing subsidies and research funding toward technologies that enable AI models to run on drones, satellites and other devices with limited processing power. Experts caution, however, that distillation has significant limitations. As Chinese AI models close the gap with their U.S. counterparts, military researchers are also examining distillation as a potential security risk. In January, researchers at the Army Engineering University published a paper on the threat of "data-free distillation," a method of reverse-engineering a model's capabilities without direct access to its core parameters. To counter that vulnerability, they proposed defense mechanisms designed to mask the hidden logical information exposed in a model's public outputs. Distilled models also inherit only selected capabilities and cannot fully replicate the broad intelligence of frontier systems. Trevor Koverko, co-founder of AI data company Sapien, said distilled models remain less capable than their teacher models. "It is best understood as transferring selected capabilities into a cheaper, locally controlled system, not achieving independence from frontier AI." (Reporting by Eduardo Baptista; Editing by Miyoung Kim and Shri Navaratnam)
[4]
What is AI model distillation and why is it becoming a US-China flashpoint?
Long regarded as a standard tool of AI research, distillation is now at the centre of a growing dispute over whether advanced AI capabilities can be transferred without the consent of the companies that created them. Washington and leading US AI firms have accused Chinese rivals of using the technique to extract capabilities from proprietary models, opening a new front in an increasingly bitter competition over technological leadership. A technique that allows developers to shrink powerful artificial intelligence models into cheaper, more efficient systems has become the latest battleground in the intensifying US-China race for AI dominance. Known as model distillation, the method uses the outputs of a powerful AI system to train a smaller model that can perform some of the same tasks with fewer computing resources. Long regarded as a standard tool of AI research, distillation is now at the centre of a growing dispute over whether advanced AI capabilities can be transferred without the consent of the companies that created them. Washington and leading US AI firms have accused Chinese rivals of using the technique to extract capabilities from proprietary models, opening a new front in an increasingly bitter competition over technological leadership. Below are key facts about the practice at the centre of the debate. What is model distillation? The largest AI models, known as frontier models, require enormous amounts of computing power, data and investment to train. Model distillation offers a way to create smaller systems by using a large "teacher" model to train a smaller "student" model. The teacher generates examples, such as answers and computer code, which are then used as training material for the student. The smaller model is not a replica of the teacher. It does not inherit the teacher's weights, architecture or full capabilities. Instead, it learns selected behaviours that enable it to perform specific tasks more efficiently. Why does distillation matter? The appeal of distillation is that it can make AI cheaper and easier to deploy. A frontier model may require large data centres and expensive chips to operate. Distilled models can run on less powerful hardware and be tailored for specific tasks. That makes them appealing to companies and governments looking to deploy AI more widely, from devices and factories to vehicles and private networks. Why are reasoning traces important Recent AI systems have increased interest in transferring not only final answers but also the steps used to reach them. These "reasoning traces" can show a smaller model how to approach a difficult problem rather than simply what answer to produce. Florian Tramer, an assistant professor at ETH Zurich who researches machine-learning security, compared the process to human learning. "If I give you a book of complicated math problems with final solutions, you will have a much harder time learning how to solve problems than if I gave you detailed solutions that describe all steps to take," he said. As reasoning traces have become more valuable, access to AI outputs has become more sensitive because they may expose some of the methods advanced systems use to tackle complex problems. Who uses distillation Distillation is a widely used AI training technique, not an inherently improper practice. US researchers and companies have long used it, including Stanford University's Alpaca project and Microsoft's Orca research, which relied on outputs from more advanced models to improve smaller ones. Chinese researchers have also used outputs from US models in public research projects, including efforts to create Chinese-language instruction models. The key difference lies in access. Open-weight models let researchers inspect and modify underlying parameters. Closed models, such as OpenAI's ChatGPT and Anthropic's Claude, remain under company control and are typically accessed through proprietary interfaces or APIs. Why has distillation become a US-China issue? The controversy is less over distillation itself and more about unauthorised extraction. AI companies argue there is a distinction between legitimate research and systematically harvesting outputs from proprietary models to replicate commercially valuable capabilities. Anthropic has accused Chinese entities including DeepSeek, Moonshot and MiniMax of conducting large-scale campaigns to obtain capabilities from Claude models. The company said those efforts targeted capabilities including software engineering and advanced reasoning. OpenAI has also said it has detected attempts by Chinese actors to use its models for distillation-related purposes. No Chinese companies have accused US rivals of distilling closed-source models so far.
[5]
What is model distillation? The AI technique at the centre of the US-China technology battle
Model distillation, a technique that trains smaller AI models using outputs from more powerful systems, has become a key issue in the US-China AI rivalry. While it lowers AI costs and expands deployment, US companies allege some Chinese firms are using the method to extract capabilities from proprietary models without authorisation. Model distillation, a long-established artificial intelligence (AI) training technique, has emerged as the latest flashpoint in the growing technology rivalry between the United States and China. While the method helps developers build smaller, cheaper AI models, US companies have accused some Chinese firms of using it to extract capabilities from proprietary AI systems without permission. Here is what model distillation is and why it has become controversial. What is model distillation?Training the world's most advanced AI models, often called frontier models, requires massive computing power, data and investment. Model distillation is a process in which a powerful AI system, known as the "teacher" model, helps train a smaller "student" model. Instead of copying the original model, the teacher generates outputs such as answers, explanations or computer code. These outputs then become training material for the smaller model. The student model does not inherit the teacher's architecture, parameters or full capabilities. Instead, it learns selected behaviours that allow it to perform specific tasks more efficiently. Why is it important?The biggest advantage of model distillation is lower cost. Large AI models often require expensive chips and large data centres to operate. Distilled models can run on less powerful hardware, making them easier to deploy across smartphones, factories, vehicles and private enterprise networks. This allows businesses and governments to use AI more widely without bearing the cost of running frontier models. What are reasoning traces?The latest generation of AI systems has increased interest in transferring not just the final answer but also the reasoning process used to reach it. These "reasoning traces" show how an AI system solves a problem step by step, helping a smaller model learn the approach instead of simply memorising answers. Florian Tramèr, assistant professor at ETH Zurich, compared it to learning mathematics. "If I give you a book of complicated math problems with final solutions, you will have a much harder time learning how to solve problems than if I gave you detailed solutions that describe all steps to take," Tramèr told Reuters. As reasoning traces become more valuable, companies have become more protective of AI outputs because they may reveal how advanced systems solve complex tasks. Who uses model distillation?Distillation is widely used across the AI industry and is not considered improper by itself. Researchers and companies in the US have used it for years, including Stanford University's Alpaca project and Microsoft's Orca research, which relied on outputs from more advanced models to improve smaller AI systems. Chinese researchers have also used outputs from US models in public research, including projects aimed at developing Chinese-language instruction models. A key distinction is between open-weight and closed AI models. Open-weight models allow researchers to inspect and modify their parameters. Closed models such as OpenAI's ChatGPT and Anthropic's Claude remain under company control and are generally accessed through proprietary application programming interfaces (APIs). Why has it become a US-China issue?The dispute is not over model distillation itself but over whether companies can use outputs from proprietary AI systems without permission. US AI companies argue there is a difference between legitimate research and systematically collecting outputs from closed models to reproduce commercially valuable capabilities. Anthropic has accused Chinese companies, including DeepSeek, Moonshot and MiniMax, of running large-scale campaigns to obtain capabilities from its Claude models, including software engineering and advanced reasoning functions. OpenAI has also said it detected attempts by Chinese actors to use its models for distillation-related purposes. According to Reuters, no Chinese company has publicly accused US AI firms of distilling capabilities from closed-source models.
[6]
Chinese military researchers tap US AI models to train defence systems
Chinese military researchers are using U.S. AI model outputs to advance defense capabilities. This technique allows them to train specialized domestic AI systems efficiently. Researchers are leveraging powerful U.S. models as a shortcut for their own development. This practice is being used for surveillance, cyber warfare, and tactical decision-making. The findings offer a rare glimpse into China's AI development strategies. Beijing: Chinese military researchers have used outputs from leading U.S. artificial intelligence models developed by OpenAI and Anthropic to train domestic AI systems to advance China's defence capabilities, according to a Reuters review of more than 80 Chinese academic papers and patents. The previously unreported findings offer a rare glimpse into how military and security-linked institutions in China are leveraging cutting-edge U.S. AI models as a shortcut to developing specialised systems of their own, despite Washington's efforts to restrict Beijing's access to advanced chips and other strategic technologies. The documents show widespread use of a technique known as "model distillation," in which outputs from a powerful AI system are used to train smaller, specialised models that can be deployed locally without the enormous computing requirements needed to build frontier AI systems from scratch. Reuters' review, which included research compiled by the Washington-based Jamestown Foundation and shared exclusively with the news agency, showed distillation is widely used by researchers linked to the People's Liberation Army and other military institutions. The papers suggest Chinese defence institutions see leading U.S. AI models as both a source of technical insight and a way to close the gap with American rivals. The dispute centres on unauthorised extraction, not distillation itself, a widely used industry practice. The issue has emerged as a major flashpoint ahead of U.S.-China talks on AI governance and safety. U.S. officials have accused some Chinese entities of using distillation to extract capabilities from American AI models, potentially undermining export controls and infringing intellectual property rights. China has rejected the accusations, saying Washington is pursuing AI "hegemonism" while arguing that U.S. firms have engaged in similar practices. Chinese developers have also disputed claims that their AI advances rely on foreign models. AI startup Moonshot last week denied allegations by the Trump administration that its Kimi K3 model was built using distillation, saying it was driven by proprietary innovations. Sunny Cheung, a Jamestown fellow who analysed over 60 of the papers, said Chinese military scientists are systematically capturing the reasoning steps of Western models to adapt them for surveillance, cyber warfare and tactical decision-making. "Teaching a model the right answer is one thing but teaching it the reasoning behind the answer is much harder," said Cheung. "These papers show Chinese military-linked researchers are trying to transfer that expensive, proprietary reasoning from Western models into smaller systems they can control and deploy locally." Reuters verified the academic literature and identified an additional two dozen military-linked case studies. One paper published last year by researchers in PLA Unit 96941, a military intelligence and cyber-warfare unit in Beijing, described using OpenAI's GPT-3.5 to process sensitive military source code. The researchers said third-party models were unsuitable for handling classified information. To overcome that limitation, they used GPT-3.5 to summarize software code and trained a domestic model on those summaries to run entirely within Chinese military networks. The White House, Pentagon, China's foreign ministry, the PLA and OpenAI did not respond to requests for comment. WIDE USE, FROM MONITORING TO MILITARY Chinese researchers have used distillation for purposes ranging from content monitoring to military deployment, Reuters' and Jamestown's review of the papers showed. At the North University of China, which has close links to the country's weapons industry, researchers used Anthropic's Claude 3 Haiku to generate synthetic training data for a text classification model for social media monitoring and content moderation. Anthropic said it does not provide commercial access to Claude in China or to Beijing-controlled firms and uses monitoring systems to detect policy violations. The company added that distilled models may lose the original systems' safety safeguards, potentially allowing sensitive capabilities to be transferred to models beyond its control. A 2024 paper from the PLA's National University of Defense Technology described using distillation to shrink an image-processing model for deployment on unmanned aerial vehicles, allowing drones to analyse live video and support navigation and targeting decisions in real time even when communications are cut. Similarly, researchers at China's Academy of Military Sciences used distillation to run a target-recognition model on tactical hardware during simulated maritime operations involving drones, ships and unmanned submarines, a study published earlier this year showed. BENEFITS AND LIMITS China has embraced distillation as it seeks to compete with the U.S. in frontier AI while facing constraints on advanced computing resources due to Washington's export controls on high-end chips. Central and local governments have promoted "model lightweighting" and edge computing, directing subsidies and research funding toward technologies that enable AI models to run on drones, satellites and other devices with limited processing power. Experts caution, however, that distillation has significant limitations. As Chinese AI models close the gap with their U.S. counterparts, military researchers are also examining distillation as a potential security risk. In January, researchers at the Army Engineering University published a paper on the threat of "data-free distillation," a method of reverse-engineering a model's capabilities without direct access to its core parameters. To counter that vulnerability, they proposed defense mechanisms designed to mask the hidden logical information exposed in a model's public outputs. Distilled models also inherit only selected capabilities and cannot fully replicate the broad intelligence of frontier systems. Trevor Koverko, co-founder of AI data company Sapien, said distilled models remain less capable than their teacher models. "It is best understood as transferring selected capabilities into a cheaper, locally controlled system, not achieving independence from frontier AI."
[7]
Chinese military researchers tap US AI models to train defence systems
BEIJING, July 31 (Reuters) - Chinese military researchers have used outputs from leading U.S. artificial intelligence models developed by OpenAI and Anthropic to train domestic AI systems to advance China's defence capabilities, according to a Reuters review of more than 80 Chinese academic papers and patents. The previously unreported findings offer a rare glimpse into how military and security-linked institutions in China are leveraging cutting-edge U.S. AI models as a shortcut to developing specialised systems of their own, despite Washington's efforts to restrict Beijing's access to advanced chips and other strategic technologies. The documents show widespread use of a technique known as "model distillation," in which outputs from a powerful AI system are used to train smaller, specialised models that can be deployed locally without the enormous computing requirements needed to build frontier AI systems from scratch. Reuters' review, which included research compiled by the Washington-based Jamestown Foundation and shared exclusively with the news agency, showed distillation is widely used by researchers linked to the People's Liberation Army and other military institutions. The papers suggest Chinese defence institutions see leading U.S. AI models as both a source of technical insight and a way to close the gap with American rivals. The dispute centres on unauthorised extraction, not distillation itself, a widely used industry practice. The issue has emerged as a major flashpoint ahead of U.S.-China talks on AI governance and safety. U.S. officials have accused some Chinese entities of using distillation to extract capabilities from American AI models, potentially undermining export controls and infringing intellectual property rights. China has rejected the accusations, saying Washington is pursuing AI "hegemonism" while arguing that U.S. firms have engaged in similar practices. Chinese developers have also disputed claims that their AI advances rely on foreign models. AI startup Moonshot last week denied allegations by the Trump administration that its Kimi K3 model was built using distillation, saying it was driven by proprietary innovations. Sunny Cheung, a Jamestown fellow who analysed over 60 of the papers, said Chinese military scientists are systematically capturing the reasoning steps of Western models to adapt them for surveillance, cyber warfare and tactical decision-making. "Teaching a model the right answer is one thing but teaching it the reasoning behind the answer is much harder," said Cheung. "These papers show Chinese military-linked researchers are trying to transfer that expensive, proprietary reasoning from Western models into smaller systems they can control and deploy locally." Reuters verified the academic literature and identified an additional two dozen military-linked case studies. One paper published last year by researchers in PLA Unit 96941, a military intelligence and cyber-warfare unit in Beijing, described using OpenAI's GPT-3.5 to process sensitive military source code. The researchers said third-party models were unsuitable for handling classified information. To overcome that limitation, they used GPT-3.5 to summarize software code and trained a domestic model on those summaries to run entirely within Chinese military networks. The White House, Pentagon, China's foreign ministry, the PLA and OpenAI did not respond to requests for comment. WIDE USE, FROM MONITORING TO MILITARY Chinese researchers have used distillation for purposes ranging from content monitoring to military deployment, Reuters' and Jamestown's review of the papers showed. At the North University of China, which has close links to the country's weapons industry, researchers used Anthropic's Claude 3 Haiku to generate synthetic training data for a text classification model for social media monitoring and content moderation. Anthropic said it does not provide commercial access to Claude in China or to Beijing-controlled firms and uses monitoring systems to detect policy violations. The company added that distilled models may lose the original systems' safety safeguards, potentially allowing sensitive capabilities to be transferred to models beyond its control. A 2024 paper from the PLA's National University of Defense Technology described using distillation to shrink an image-processing model for deployment on unmanned aerial vehicles, allowing drones to analyse live video and support navigation and targeting decisions in real time even when communications are cut. Similarly, researchers at China's Academy of Military Sciences used distillation to run a target-recognition model on tactical hardware during simulated maritime operations involving drones, ships and unmanned submarines, a study published earlier this year showed. BENEFITS AND LIMITS China has embraced distillation as it seeks to compete with the U.S. in frontier AI while facing constraints on advanced computing resources due to Washington's export controls on high-end chips. Central and local governments have promoted "model lightweighting" and edge computing, directing subsidies and research funding toward technologies that enable AI models to run on drones, satellites and other devices with limited processing power. Experts caution, however, that distillation has significant limitations. As Chinese AI models close the gap with their U.S. counterparts, military researchers are also examining distillation as a potential security risk. In January, researchers at the Army Engineering University published a paper on the threat of "data-free distillation," a method of reverse-engineering a model's capabilities without direct access to its core parameters. To counter that vulnerability, they proposed defense mechanisms designed to mask the hidden logical information exposed in a model's public outputs. Distilled models also inherit only selected capabilities and cannot fully replicate the broad intelligence of frontier systems. Trevor Koverko, co-founder of AI data company Sapien, said distilled models remain less capable than their teacher models. "It is best understood as transferring selected capabilities into a cheaper, locally controlled system, not achieving independence from frontier AI." (Reporting by Eduardo Baptista; Editing by Miyoung Kim and Shri Navaratnam)
Share
Copy Link
Chinese military researchers leveraged US AI models from OpenAI and Anthropic to train defence systems through model distillation, according to a Reuters review of over 80 academic papers. The findings expose how China's People's Liberation Army systematically extracts capabilities from proprietary models despite US export controls, intensifying the US-China flashpoint over AI governance.
Chinese military researchers have systematically used outputs from leading US AI models developed by OpenAI and Anthropic to train domestic defence systems, according to a Reuters investigation of more than 80 Chinese academic papers and patents
2
3
. The previously unreported findings reveal how military and security-linked institutions in China are leveraging cutting-edge US AI models as a shortcut to developing specialised systems, despite Washington's efforts to restrict Beijing's access to advanced chips and strategic technologies. The documents show widespread use of AI model distillation, a technique where outputs from a powerful AI system train smaller, specialised models that can be deployed locally without enormous computing requirements needed to build frontier AI systems from scratch1
2
.AI model distillation uses a large teacher model to train a smaller student model by generating examples such as answers and computer code as training material
1
4
. The student model does not inherit the teacher's weights, architecture or full capabilities. Instead, it learns selected behaviours enabling it to perform specific tasks more efficiently1
5
. The appeal of distillation is that it makes AI cheaper and easier to deploy. A frontier model may require large data centres and expensive chips to operate, while distilled models can run on less powerful hardware and be tailored for specific tasks1
4
. This makes them appealing to companies and governments looking to deploy AI more widely, from devices and factories to vehicles and private networks1
.Recent AI systems have increased interest in transferring not only final answers but also reasoning traces—the steps used to reach them
1
4
. Florian Tramèr, an assistant professor at ETH Zurich who researches machine-learning security, compared the process to human learning: "If I give you a book of complicated math problems with final solutions, you will have a much harder time learning how to solve problems than if I gave you detailed solutions that describe all steps to take"1
5
. As reasoning traces have become more valuable, access to AI outputs has become more sensitive because they may expose methods advanced systems use to tackle complex problems1
4
.Reuters' review, which included research compiled by the Washington-based Jamestown Foundation, showed distillation is widely used by researchers linked to the People's Liberation Army and other military institutions
2
3
. Sunny Cheung, a Jamestown fellow who analysed over 60 of the papers, said Chinese military scientists are systematically capturing the reasoning steps of Western models to adapt them for surveillance, cyber warfare and tactical decision-making2
3
. "Teaching a model the right answer is one thing but teaching it the reasoning behind the answer is much harder. These papers show Chinese military-linked researchers are trying to transfer that expensive, proprietary reasoning from Western models into smaller systems they can control and deploy locally," Cheung stated2
3
.One paper published last year by researchers in PLA Unit 96941, a military intelligence and cyber-warfare unit in Beijing, described using OpenAI's GPT-3.5 to process sensitive military source code
2
3
. The researchers said third-party models were unsuitable for handling classified information. To overcome that limitation, they used GPT-3.5 to summarize software code and trained a domestic model on those summaries to run entirely within Chinese military networks2
3
. A 2024 paper from the PLA's National University of Defense Technology described using distillation to shrink an image-processing model for deployment on unmanned aerial vehicles, allowing drones to analyse live video and support drone navigation and targeting decisions in real time even when communications are cut2
3
.Related Stories
The dispute centres on unauthorised extraction, not distillation itself, which is a widely used industry practice
2
3
. The issue has emerged as a major flashpoint ahead of US-China talks on AI governance and safety2
3
. US officials have accused some Chinese entities of using distillation to extract capabilities from American AI models, potentially undermining export controls and infringing intellectual property rights2
3
. Anthropic has accused Chinese entities including DeepSeek, Moonshot and MiniMax of conducting large-scale campaigns to obtain capabilities from Claude models, targeting capabilities including software engineering and advanced reasoning1
4
. OpenAI has also detected attempts by Chinese actors to use its models for distillation-related purposes1
4
.China has rejected the accusations, saying Washington is pursuing AI hegemonism while arguing that US firms have engaged in similar practices
2
3
. Chinese developers have disputed claims that their AI advances rely on foreign models. AI startup Moonshot last week denied allegations by the Trump administration that its Kimi K3 model was built using distillation, saying it was driven by proprietary innovations2
3
. Anthropic said it does not provide commercial access to Claude in China or to Beijing-controlled firms and uses monitoring systems to detect policy violations2
3
. The company added that distilled models may lose the original systems' safety safeguards, potentially allowing sensitive capabilities to be transferred to models beyond its control2
3
.The papers suggest Chinese defence institutions see leading US AI models as both a source of technical insight and a way to close the gap with American rivals
2
3
. At the North University of China, which has close links to the country's weapons industry, researchers used Anthropic's Claude 3 Haiku to generate synthetic training data for a text classification model for social media monitoring and content moderation2
3
. Researchers at China's Academy of Military Sciences used distillation to run a target recognition model on tactical hardware during simulated maritime operations involving drones, ships and unmanned submarines, a study published earlier this year showed2
3
. The White House, Pentagon, China's foreign ministry, the People's Liberation Army and OpenAI did not respond to requests for comment2
3
. Watch for escalating regulatory measures from Washington, potential restrictions on API access, and how this controversy shapes upcoming US-China technology battle negotiations on AI governance.Summarized by
Navi
13 Jul 2026•Policy and Regulation

23 Feb 2026•Technology

03 Mar 2025•Technology

1
Technology

2
Technology

3
Technology
