Since early 2023, photographer Jingna Zhang and a small crew of volunteers have worked tirelessly to maintain an image-sharing social media and portfolio app called Cara. So far, it has attracted about 1.5 million artists. What drew them to the platform? A shared opposition to the unauthorized use of their work to train AI models and a desire to publicize their art while avoiding exploitation by Big Tech.
But while Cara filters out AI images and offers protective features—including Glaze, a tool meant to mask the style of the images picked up by scrapers in order to disrupt AI mimicry—preventing scrapes themselves is nearly impossible. And just this month, beginning on August 13, Cara was subjected to three major scrapes, which spiked its server fees and alarmed creators who had migrated there from platforms like Instagram, where all content is explicitly available to Meta as training data.
The first of these incidents came to light when the individual responsible posted a 12-terabyte archive of 12 million works from Cara—more or less its entire library of publicly available images—on the subreddit r/DefendingAIArt. “It was a fun project,” wrote the redditor, MandarinDawnPoppy994, in his since-deleted post, saying the process cost him less than $10.
“We actually found out about it through our users tagging us,” Zhang tells WIRED, since the scraper was “gloating and looking for other people to join him to do something with the dataset on Reddit,” sparking a fierce debate across AI-related forums about the ethics of what he had done.
Source link







