He Scraped All of Their Art for AI. Now He’s Collaborating on a Tool to Help Them

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane August 28, 2026 4 min read
He Scraped All of Their Art for AI. Now He’s Collaborating on a Tool to Help Them

Photographer Jingna Zhang and a volunteer crew have spent months maintaining Cara, an image-sharing app that now hosts 1.5 million artists. The platform exists because creators wanted a space free from the unauthorised training of their work by artificial intelligence models.

Despite filters and protective tools like Glaze, which attempts to obscure image styles, stopping scrapers remains difficult. In August, Cara faced three major data harvests. These attacks drove up server costs and alarmed users who had moved there from Instagram to avoid Meta’s data collection.

The first incident emerged when a user posted a 12-terabyte archive of 12 million images from Cara to the subreddit r/DefendingAIArt. The poster, known as MandarinDawnPoppy994, described the process as a fun project that cost him less than $10.

Zhang learned of the breach when users tagged the team. The scraper was openly boasting and seeking collaborators on Reddit, which sparked a heated debate about ethics. Zhang told WIRED she felt the act was targeted and hurtful. She noted that current laws do not adequately protect against such data harvesting, allowing scrapers to claim technical legality. Zhang is already involved in two class action lawsuits against Stability AI, Midjourney, and Google.

Surprisingly, the original scraper later regretted his actions and agreed to work with Zhang on a new open-source tool. Meanwhile, other attackers exploited Cara’s vulnerabilities. Some appeared emboldened by the first incident to launch similar copycat attacks.

A second scraper downloaded 8.5 million links, along with usernames, titles, and tags, uploading them to Hugging Face. The platform received numerous takedown requests. In a statement, Hugging Face said it would ask the user, CaptiveDreamer, to remove personal metadata. However, the company could not remove the links themselves because the artworks are hosted elsewhere. They stated that further reports on the same basis would not change this outcome.

On August 22, a third scraper took 123,000 images, text posts, and user bios containing personal information from a site called Academic Torrents. Zhang launched a GoFundMe campaign to cover legal fees, setting a goal of $120,000. The funds would support cyber and copyright defence strategies. As of Thursday, she had raised more than $100,000 and is seeking additional legal assistance.

Zhang is frustrated by the confusion regarding what the team can realistically do to protect artists. Some users have already deleted their portfolios and left the site. She explained that current safeguards, including temporary login gates, do not solve the internet-wide problem.

She also worries that users blaming her for the attacks may not realise that Cara cannot guarantee complete security. Any site can be scraped. Deleting work does not make people safer elsewhere, as larger platforms face even more risk.

Zhang, who is not a tech founder by design, now has a new ally. The person who started the scraping frenzy apologised and deleted his dataset after seeing how hurt the community was.

Heft, a student in North America with a background in software and digital preservation, requested anonymity due to death threats he received. In a conversation over Discord, he told WIRED that scraping Cara was a technical project with no intention of making the data public.

He admitted to making a foolish decision to ragebait with the dataset on Reddit and being carried away by trolling. He knew it would provoke artists but did not anticipate the sheer anguish. He saw people sharing panic attacks and deleting entire portfolios. Direct conversations with artists showed him how personal their work was and how much they valued ownership.

In retrospect, Heft said deliberately targeting Cara and presenting it that way was cruel and thoughtless. He missed the consequences beyond causing anger.

He noted that while he was curious if the data could train an AI, he never believed it would. Twelve million images is not enough compared to large-scale web scrapes used by commercial labs. Random users might dump data on Hugging Face, but he considers it highly unlikely that OpenAI or Anthropic scans every new dataset.

Once shaken from that mindset, Heft joined Cara’s Discord server as a troubleshooter. He identified why proposed fixes were unlikely to succeed. Zhang said he helps explain structural weak spots that allow scraping. When a tool or design is discussed, he can demonstrate how to break through in minutes.

Because Heft believes no site can be truly unscrapable, he and Zhang are collaborating on a countermeasure. The tool, Lantern, lets artists create a one-way fingerprint for their images without storing them on the platform. It regularly scans new publicly available AI image datasets. If an artist’s work appears, they receive a notification and a link to request removal or submit a takedown notice.

Lantern is new and is an imperfect workaround inspired by a regulatory vacuum. Zhang and Heft aim to give artists more control and awareness of how content is collected, regardless of a website’s terms of service. Because it is open-source, anyone can contribute to the code.

Zhang hopes the fallout from the scraping spree draws attention from policymakers looking beyond copyright, which is just one piece of a complicated puzzle.

Getting attacked by AI is going to become commonplace. Most of the time, the attacker will not say sorry, let alone try to set things right.

What it means

Artists now have a way to flag their work if it appears in scraped datasets, but this does not stop the initial theft. The collaboration between a former attacker and a defender highlights how technical fixes alone cannot solve a legal and ethical gap.

Scroll to Top