Skip to content

Cara was built to protect artists from AI. Then scrapers came anyway.

After Cara was scraped three times in ten days, the artist-led platform has become an uncomfortable test of what "opting-out" of AI training actually means online.

Cara's profile portfolio page on desktop. Image credit: https://blog.cara.app/blog/caras-app-features

Table of Contents

When Cara first launched, the premise of the platform was to enable artists to share their work online without automatically consenting to its use in generative AI training. However, in August, this was tested three times in ten days.

The artist-focused social platform, founded by photographer, Jingna Zhange, was hit by a series of large-scale scraping attacks beginning on August 13. The first reportedly collected around 12 million publicly available artworks, effectively Cara's entire public image library. A second gathered approximately 8.5 million links alongside usernames, titles, and tags, while the third scrape on August 22 included 123,000 images as well as text posts and user bios.

Photo by Firosnv. Photography on Unsplash

For a platform that has attracted around 1.5 million users, partly because of its stance against AI training, these attacks expose the difficult reality. Declaring that your work is off-limits and technically preventing someone from taking it are two very different things.

The problem with consent online

Cara first emerged when artists became increasingly concerned about how their work was being collected to train image generators. Early generative AI systems were developed using enormous datasets of images gathered across the public internet, including copyrighted material.

That left artists in a difficult position. A working illustrator, photographer, or designer often needs an online portfolio to find clients and build an audience. Although, making that portfolio publicly accessible can also make it available to automated scrapers.

Cara attempted to create an alternative. Its terms explicitly prohibit unauthorised AI training, the platform filters AI-generated images and it has offered tools such as Glaze, developed by researchers at the University of Chicago to make the images harder for AI systems to mimic. None of these measures can make a public website impossible to scrape.

That limitation matters, but it does not necessarily make Cara's approach pointless. The platform establishes something that can otherwise become murky online - the users have expressly said no.

The controversy therefore questions digital consent. Does making an artwork available for people to view also make it fair game for automated collection, redistribution, and eventual AI training?

For many artists, those are clearly different permissions. Some proponents of AI development disagree, arguing that publicly accessible data needs to remain available for research and technical development.

The individual behind the first Cara scrape initially defended the practice on similar grounds, comparing data collection for AI development to building infrastructure for the broader public good. But after seeing the responses from artists, he reversed course, deleted the dataset, and apologised for deliberately targeting the community.

He has since begun working alongside Zhang on an open-source project called Lantern, designed to alert artists when their work appears in newly published AI datasets.

Copyright law offers an incomplete answer

The second Cara scrape shows why the issue cannot easily be reduced to theft versus ownership. After millions of Cara links and associated metadata appeared on AI platform Hugging Face, artists submitted copyright complaints. Hugging Face acknowledged that the artists owned their work but declined to remove the URLs, arguing that copies of the artworks themselves were not hosted on its servers.

That distinction reveals one of the central problems surrounding AI datasets. Copyright can establish who owns an image, but questions around scraping, linking, dataset creation, and AI training remain far less settled.

Photo by Umberto on Unsplash

Zhang is already involved in separate class-action lawsuits concerning the alleged use of copyrighted artwork in AI training, while Cara has turned to crowdfunding to help cover legal costs generated by the recent attacks.

There is also a practical imbalance at play. Scraping millions of publicly accessible files can be relatively cheap for an individual, while defending against that activity can create substantial server and legal costs for a small platform.

Cara was never an impenetrable safe haven

There is another important distinction to be made. The appearance of Cara's images in scraped datasets does not mean that all 12 million artworks are suddenly being used to train major commercial AI models.

The first scraper himself has since argued that the direct training risk was overstated. Major AI systems rely on datasets at an enormous scale, and there is no guarantee that a dataset uploaded to an open repository will ever become training material for a commercial model.

For Cara's artists, however, that does not erase the issue. The objection is also about control. A community explicitly formed around withholding consent was deliberately targeted, and its work was redistributed anyway. Cara cannot make publicly art impossible to copy. Few platforms can. What is can do, however, is make the artist's intentions unambiguous.

Lantern reflects that reality. Rather than promising to stop scraping altogether, the tool would allow artists to create digital fingerprints of their work and check new datasets for matches. It is detection after fact rather than perfect prevention.

That may be where the Cara story becomes most significant. The internet was designed around making information easy to access and reproduce. Generative Ai has increased the value of collecting that information at scale, while laws and technical safeguards have struggled to catch up.

Artists are left with an uncomfortable choice: keep their work private and lose much of the visibility required for a creative career, or publish it knowing that saying "do not use this for AI" may not actually stop anyone.

Cara's three scrapes have not proved that opting out is meaningless. They have shown how little infrastructure currently exists to make that choice enforceable.

Latest