Japan Government Pushes AI Firms to Disclose Training Data Under New IP Protection Framework

2 Sources

Share

Japan's government adopted guiding principles requiring generative AI operators to disclose training data outlines and methods. The non-binding Principles Code aims to boost AI transparency while balancing innovation with protection of intellectual property rights amid growing concerns over unauthorized use of copyrighted materials.

Japan Government Introduces Transparency Framework for Generative AI

The Japan government adopted new guiding principles on Tuesday designed to address mounting concerns about intellectual property protection in the rapidly evolving generative AI landscape. The Principles Code calls on AI operators to disclose outlines of AI training data and methods used to develop their systems, marking a significant step toward greater AI transparency without imposing legally binding restrictions.

1

Source: MediaNama

Source: MediaNama

Businesses accepting all or part of the code must notify the government and publicly disclose learning processes, types of learning data, and data collection methods for their generative AI models on their websites. The framework applies not only to domestic businesses but also to overseas operators providing AI systems and services in Japan, casting a wide net across the global AI industry.

1

Comply or Explain Approach Balances Innovation and Rights Protection

Japan's Cabinet Office structured the revised "Principle-Code for Protection of Intellectual Property and Transparency for the Appropriate Use of Generative AI" around a comply or explain approach. Under this framework, generative AI businesses either follow the principles or publicly explain why they do not, creating accountability without rigid enforcement.

2

The proposal emerged from discussions by the Study Group on Intellectual Property Rights in the AI Era on August 18, following a public consultation that ran from December 26, 2025 to January 26, 2026. The Cabinet Office's meeting materials identify this as a revised draft incorporating feedback from stakeholders across the AI ecosystem.

2

Addressing Copyright Infringement Concerns Through Data Collection Standards

Growing concerns about unauthorized use of copyrighted materials have driven this policy initiative. Texts, images, and other materials are increasingly being used by generative AI for learning without permission, potentially resulting in copyright infringement and intellectual property rights violations.

1

The draft principles require businesses to establish processes ensuring their use of data to develop and train generative AI does not infringe others' intellectual property. Companies must respect access restrictions including paywalls, use crawler measures that follow machine-readable instructions such as robots.txt, and endeavor to avoid crawling pirate sites. They must also disclose their crawler measures for each user agent and provide notice when those measures change.

2

Enhanced Disclosure Requirements for Rights Holders and Users

The framework establishes mechanisms allowing rights holders to seek information about whether their works were used in AI development. While businesses are not required to release every individual item of training data publicly, they must provide specified information about models, training data, and collection methods. Rights holders pursuing legal remedies can ask whether a specific URL or identifier was used in training or validation, limited to what the business can readily access and confirm.

2

Companies accepting the code will make public whether learning data includes information that could lead to copyright infringement if requested by AI users and rights holders, provided certain conditions are met. This targeted disclosure balances transparency demands with protection of proprietary information including trade secrets.

1

2

Technical Safeguards and Ongoing Accountability Measures

Beyond data collection standards, the draft proposes that businesses retain training-related logs for a certain period and take technical measures to prevent infringing outputs where possible. Companies should implement measures such as digital watermarking and C2PA standards to verify content origin and provenance.

2

Businesses would establish contact points for rights holders, clarify requirements for approaching them, and maintain records of their responses. The principles also call for businesses to review their IP protection frameworks at least once a year and publish their substance, creating ongoing accountability.

2

Japan's Unique Legal Position on AI Training

Japan occupies a distinctive position in global AI policy debates. Article 30-4 of its Copyright Act permits certain uses of copyrighted works for purposes such as data analysis, including AI training, subject to specified conditions. This existing legal framework remains largely intact under the new principles, which layer transparency requirements onto Japan's permissive training regime.

2

Kimi Onoda, minister for intellectual property strategy, stated at a news conference on Tuesday that the government will respond appropriately, including considering taking further measures as the situation develops. This signals potential evolution of the framework based on industry adoption and emerging challenges.

1

Source: Japan Times

Source: Japan Times

The proposal arrives as the legal treatment of AI training remains under active international discussion, with Japan's approach potentially influencing how other nations balance innovation incentives against protection of intellectual property in the generative AI era.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved