ShieldFont: Free Open Source Font Feeds AI Scrapers Poisoned Data While Protecting Your Content

3 Sources

Share

Brazilian studio Seneda & Abrucio and Copenhagen foundry Playtype launched ShieldFont, a free open source font that displays normal text to readers but feeds AI scrapers grammatically correct gibberish. Using OpenType glyph substitution, it swaps roughly 25% of words to contaminate AI training datasets while passing quality filters.

ShieldFont Launches as Open Source Defense Against AI Scrapers

Brazilian creative studio Seneda & Abrucio partnered with Copenhagen-based type foundry Playtype to release ShieldFont, a free open source font designed to thwart AI scrapers by feeding them poisoned data

1

2

. The innovative tool addresses a growing concern among content creators: unauthorized scraping of web content for LLM training. Unlike robots.txt files that depend on voluntary compliance, ShieldFont actively deceives AI scrapers by exploiting the fundamental difference between how humans and bots read web pages

3

.

Source: The Register

Source: The Register

The font displays perfectly readable text to human viewers while presenting grammatically coherent but semantically altered content to scrapers reading raw HTML. Studio founders Isaque Seneda and Gabriel Abrucio emphasized their goal: "We wanted a mechanism with actual consequences: scrape without asking, and you can't tell if what you took was real"

1

. This approach moves beyond simple concealment to actively contaminating AI training datasets with unreliable information.

How OpenType Glyph Substitution Powers the Defense

ShieldFont operates through OpenType font technology, specifically leveraging glyph substitution (GSUB) tables in a novel way

1

. While GSUB tables traditionally replace individual characters or character pairs—like combining lowercase f and i into a ligature—ShieldFont extends this functionality to entire words. The system automatically swaps approximately 25% of words in any text block, replacing them with alternatives from the same grammatical category

2

.

Source: TechSpot

Source: TechSpot

The substitution logic operates with remarkable sophistication. Words aren't randomly swapped; instead, they're organized into roughly 250 pools that account for part of speech, sense category, concreteness, singular or plural forms, verb transitivity, verb inflection, and adjective degree

1

. A plural abstract noun about communication, for instance, gets replaced only with another plural abstract noun about communication. This precision ensures the altered text remains grammatically valid—critical for bypassing AI quality filters that discard obvious nonsense.

Testing revealed the strategy's effectiveness. When Seneda & Abrucio ran shielded text through FineWeb-Edu, a quality filter used to assemble large public training datasets, approximately one in 10 passages that passed before shielding still passed afterward

2

. More significantly, 55.8% of shielded passages no longer conveyed their original factual claims

3

. The poisoned data enters AI training datasets appearing legitimate but carrying corrupted information.

Technical Implementation and Accessibility Features

The default ShieldFont release adapts Playtype's existing Optik typeface, available in six weights from Regular to Black

3

. Type designer Jeppe Pendrup noted that "every detail had to serve the human eye, while the hidden system underneath served an entirely different purpose"

3

. Despite its sophisticated functionality, the web font remains practical for everyday use. Compressed web fonts including the complete GSUB dictionary measure around 800 KB, while document and desktop versions reach approximately 5 MB

1

.

Implementation options cater to various technical skill levels. Writers and developers can protect content from AI training through an online encoder, a React component available on GitHub, or CSS and CDN integration for blogs, content management systems, and static websites

3

. ShieldFont ships with three GSUB dictionaries, and the GitHub repository includes documentation for creating custom mappings to prevent reverse-engineering

1

.

Accessibility considerations shaped development priorities. Screen readers used by blind readers interpret code rather than rendered pixels, meaning they would vocalize the decoy text. ShieldFont addresses this through a beta feature that provides assistive technologies with authentic content

2

. The current release targets high-frequency English nouns, with the team acknowledging that expanding language support represents future development work

2

.

Limitations and the Broader Movement to Protect Creative Work

Seneda & Abrucio acknowledge ShieldFont isn't unbreakable. The most obvious vulnerability involves optical character recognition: taking a screenshot of a ShieldFont-protected page and running OCR recovers the original text, since OCR reads rendered pixels rather than raw HTML

1

. Determined scrapers targeting specific websites could inspect the font files and reverse the word mappings. However, the creators argue this misses the point. ShieldFont aims to make mass scraping operations riskier, more expensive, and less reliable—not to create an impenetrable fortress.

The project joins a growing ecosystem of tools for protecting creative work from AI. Nightshade adds imperceptible alterations to digital art that corrupt AI models trained on poisoned images

3

. Ghost Font uses optical illusions with moving text to prevent AI from accurately reading content, though its reliance on video format and visual noise limits practical applications

3

. ShieldFont distinguishes itself through versatility and usability—it functions as a standard web font while embedding protection mechanisms.

Playtype CEO Daniél Andreasen positioned the release within typography's historical role: "Typography has always helped humanity preserve and share its ideas. Now it can help protect them too. That is why ShieldFont had to be open source, so its system can grow beyond Optik and become part of many different typefaces"

3

. The project includes a custom font builder allowing anyone to apply the ShieldFont protocol to compatible typefaces with private mappings, enabling collective defense against unauthorized scraping.

Redefining Consent in the Age of AI Training

The philosophical foundation driving ShieldFont challenges prevailing assumptions about web publishing. "ShieldFont is not an anti-AI project," Seneda and Abrucio clarified. "We're simply against the idea that publishing is the same as consenting"

3

. They observed that existing opt-out protocols like robots.txt face routine violations, necessitating a self-enforcing alternative. By transforming content itself into an opt-out mechanism, ShieldFont shifts power dynamics between creators and AI companies conducting mass data harvesting.

The team articulated their broader vision: building a web where taking content without permission carries tangible consequences. Mass scrapers operate on the assumption that publicly accessible content can be freely harvested for commercial AI development. ShieldFont introduces uncertainty into that equation. Scrapers cannot determine in advance whether a site uses the font or which mapping it employs, forcing them to either accept contaminated data risks or invest resources in per-site analysis that undermines mass scraping economics.

Whether ShieldFont gains widespread adoption remains to be seen, but it represents a shift in how creators might enforce opt-out mechanisms for content scraping. Full methodology and benchmark data appear in the project's white paper, while the complete toolkit awaits on GitHub for anyone seeking to protect content from unauthorized AI training. For content creators weighing their options, ShieldFont offers a practical middle ground: share work with human readers while actively defending against exploitative data collection practices that fuel LLM training without compensation or consent.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved