With the exploding popularity of generative artificial intelligence, many open-source models are now available online for anyone to adapt for their task, such as generating product renderings in a certain artistic style.
But these models also find their way into the hands of nefarious actors who may optimize them to produce illegal content, like hate speech or child sexual abuse material (CSAM). This is a growing problem — the National Center for Missing and Exploited Children received more than 1.5 million reports of AI-generated CSAM in 2025, an increase from 67,000 in 2024.
Engineers usually test AI for harmful capabilities by prompting the model and inspecting its outputs, but this is impossible for CSAM, since it is illegal in the U.S to generate such content, regardless of intent.
To avoid this dilemma and improve AI safety, Associate Professor Ashia Wilson and her graduate student, Vinith Suriyakumar, teamed up with researchers from MIT’s Healthy ML Lab, led by Marzyeh Ghassemi, and child safety non-profit Thorn to develop a new auditing approach that determines whether a model can produce CSAM without prompting it. Thorn is a child safety nonprofit whose mission is to transform how children are protected from sexual abuse and exploitation in the digital age.
Their technique examines how the inner workings of a model have been adapted, but it never generates an output. By examining hidden representations, it can reliably infer whether a model has been specialized to produce harmful imagery.
When tested, the auditing procedure identified model variations that had been specialized to generate CSAM with 100 percent accuracy.
Source link







