Cara and former scraper develop open-source tool to alert artists when work appears in AI datasets
After several large-scale scrapes of the artist platform Cara, its founder is working with one of the people involved on an open-source detection tool called Lantern.
Cara, an image-sharing and portfolio platform built for artists who do not want their work used to train AI systems without permission, has faced a series of large data scrapes in August. The platform, founded by photographer Jingna Zhang and maintained by a small team, has attracted about 1.5 million artists.The first major incident began on August 13, when a scraper said he had collected a 12-terabyte archive containing about 12 million publicly available works from Cara. He later said the project cost less than $10. A second scraper collected about 8.5 million links along with metadata including usernames, titles and tags and uploaded the material to Hugging Face. The company said it would seek removal of personal metadata but would not remove links that simply pointed back to material already published on Cara.On August 22, a third scrape obtained about 123,000 images together with text posts and user biographies containing personal information and distributed the data through Academic Torrents. Zhang subsequently launched a fundraising campaign seeking $120,000 for legal work; by the time of the report, more than $100,000 had been raised.The person behind the first scrape, identified by the screen name Heft, later deleted his dataset after being confronted by a Cara user and said he regretted the impact on artists. He has since joined Cara's community as a technical troubleshooter and is collaborating with Zhang on an open-source project called Lantern.Lantern is designed to create a one-way fingerprint of an artist's image without storing the image itself. It then scans newly available public AI image datasets and alerts the artist if matching material appears, allowing the creator to seek removal or submit a takedown request. Zhang and Heft acknowledge that no public website can be made completely resistant to scraping, making detection and response a more practical goal than promising total prevention.
G.Hubert--PP