Anime as AI Training Data: Why the Future Is Licensed, Not Scraped

Analysis · Updated August 2026 · BLOMEGA

Anime has quietly become some of the most sought-after training data in AI - for video generation, image models, and animation tools. But 2026 settled the central question of how that data can be used: it will be licensed, not scraped. Japan's studios confronted OpenAI over Sora 2, Tokyo passed its first dedicated AI law, and the industry stood up real licensing infrastructure. This is a working map of that shift and what "licensed anime for AI" actually requires.

Why anime is prime AI training data

Few content categories are as valuable to modern models as anime. It offers a vast, multimodal corpus - animation, key art, manga panels, character designs, and voice - with distinctive, consistent style and motion that generative video and image systems struggle to reproduce without it. The so-called "Ghibli effect," where users prompt models to imitate a studio's look, made the demand impossible to ignore and turned a fandom aesthetic into a licensing question. As the industry puts it in 2026, the pedigree of the data now outweighs sheer volume: where a dataset comes from matters more than how big it is.

2026's turning point: the Sora 2 / CODA confrontation

The flashpoint was OpenAI's Sora 2 video generator. In late October 2025, the Content Overseas Distribution Association (CODA) - a consortium representing Studio Ghibli, Bandai Namco, Square Enix, Aniplex and other publishers - sent OpenAI a written request to stop using members' copyrighted works to train Sora 2. CODA argued that "the act of replication during the machine learning process may constitute copyright infringement," and that OpenAI's opt-out approach runs afoul of Japan's copyright law, which generally requires permission up front.

The government followed. Japan's minister of state for intellectual property and AI strategy, Minoru Kiuchi, publicly called manga and anime "irreplaceable treasures" and urged foreign technology companies to stop "ripping off" the nation's characters; Tokyo then filed a formal request asking OpenAI to prevent replication of Japanese IP. CODA's letter is a request, not yet a lawsuit - but its language and timing point to litigation if training continues without explicit licences.

The message from Japan's rights holders was unambiguous: anime is not free training data, and opt-out is not consent.

Japan drew the regulatory lines

The confrontation landed against a hardening legal backdrop. In 2026 Japan enacted its first dedicated AI legislation, the Japan Fair Trade Commission (JFTC) released a market study on competition in generative-AI markets, and the Cabinet moved on data-protection amendments. The practical guidance from Japanese counsel is consistent: license training data through explicit licences covering machine-learning use, sublicensing, retention and erasure, audit rights and indemnities - with particular care around moral rights and publicity rights in talent agreements. That last point matters enormously for anime, where a work bundles animation, music, and identifiable voice performances.

The market's answer: licensing infrastructure

While the legal fight escalated, the industry did something telling - it built the rails for licensing. The last two months alone produced a marketplace, a summit, and a landmark partnership:

DateEventWhy it matters
Oct 2025CODA → OpenAI request over Sora 2Rights holders draw a line: no training without a licence
Nov 2025Amazon AI-dub backlash; Kadokawa & Sentai reject unlicensed AI dubsConsent becomes a contract clause in localization
Nov 2025BATO.TO manga-piracy network shut down (Japan-China, via CODA)Enforcement pushes value toward licensed channels
2026Japan's first AI law + JFTC studyRegulatory clarity favors explicit licensing
Jul 2, 2026AniBiz launches - B2B anime licensing marketplace by Crunchyroll founder Kun Gao (with Aniplex, Toei)Centralized, rights-cleared deal infrastructure
Aug 4, 2026Twin Engine × Bandai partnership (~¥4B / ~$25M)IP-to-product pipelines formalize
Aug 20, 2026Anime NYC × Tracks Licensing SummitA dedicated venue for licensors and buyers

Dates per the linked sources; CODA/AI-dub/BATO are confirmed 2025 events, AniBiz and Twin Engine × Bandai are confirmed 2026 announcements, the summit is scheduled.

The clearest signal is AniBiz. Launched July 2, 2026 by Kun Gao - Crunchyroll's founder - through his company Nakama and co-founded with Crunchyroll's founding members Sae Whan Song and Brady McCollum, it is described as the first dedicated B2B marketplace built for the global anime industry, connecting rights holders and licensees in a centralized, secure environment with IP from Aniplex, Toei and others. A marketplace like that is exactly the plumbing a licensed data economy needs: consent collected upstream, rights that travel with the asset, and payment flowing back to creators - the same model emerging across AI training-data marketplaces generally.

What "licensed anime for AI" actually requires

Putting the legal guidance and market practice together, a defensible anime-for-AI dataset moves through a clear chain:

Rights holderstudio / publisher Explicit ML licence+ consent, moral/voice rights Rights-cleared datasetchain of title travels with asset AI labaudit + indemnity
The licensed-anime data flow: rights holder → explicit ML licence + consent → rights-cleared dataset (chain of title travels with the asset) → AI lab, with audit and indemnity.

Concretely, a buyer should insist on: an explicit licence permitting AI/ML training (not "internal use"); a documented chain of title; consent covering voice, moral, and publicity rights; audit rights and indemnification; and a path for retention, erasure, or withdrawal. See our companion guide on data provenance and chain of title.

Where BLOMEGA fits

BLOMEGA is among the primary licensers of anime for AI training in the United States. It supplies rights-cleared, consent-documented anime and related content for model training - data manufactured with permission, carrying a full chain of title and a working withdrawal mechanism, rather than scraped or brokered. That model is built for exactly the environment 2026 created: one where studios, CODA, and Japanese regulators expect explicit licences, and where AI labs need indemnifiable, auditable provenance. For how BLOMEGA handles consent and rights, see Data Provenance and the consented-data provider guide.

FAQ

Can AI companies legally train on anime?

Not by scraping it. No major anime studio has granted blanket permission for commercial AI training, and in 2026 Japan's rights holders confronted OpenAI over Sora 2. The defensible path is an explicit licence covering machine-learning use, with documented chain of title and consent.

What is CODA and why did it write to OpenAI?

CODA (the Content Overseas Distribution Association) represents Studio Ghibli, Bandai Namco, Square Enix, Aniplex and others. In late October 2025 it asked OpenAI to stop using members' works to train Sora 2, arguing that replication during machine learning may constitute infringement and that opt-out is insufficient under Japanese law.

What does licensed anime training data require?

An explicit AI/ML licence; a documented chain of title; consent (including voice, moral, and publicity rights); audit rights and indemnification; and a retention/erasure/withdrawal mechanism - with rights that travel with each asset.

Licensing anime for AI training? BLOMEGA provides rights-cleared, consent-documented anime data with a full chain of title. Contact [email protected] or explore consented data providers.