For her safety, Doe has opted to receive alerts from the US Department of Justice Victim Notification System any time she may be a victim in a new criminal investigation. Although she has received countless alerts, she was shocked when the CCCP notified her that it had identified AI-generated CSAM on xAI that depicted her. This re-traumatized Doe, whose complaint alleged that messages were found on online forums “between offenders chatting about creating AI generated CSAM of Plaintiff and other similarly situated known, legacy, victims of CSAM.”
Now, Doe fears that xAI has not only made it easier to make more violative images of the most distressing time in her life, but also that xAI allegedly has stored the images that Grok generates and uses those outputs to further train Grok. Because of this, she believes that Grok has been trained on both the initial set of images that have haunted her for more than 20 years and the more recent AI-generated ones.
This is the first case to accuse xAI of training on CSAM, and the complaint does not go into great detail on that claim. Previously, Ars reported on a controversial dataset that was later scrubbed after researchers found CSAM in the training data, but there’s no indication xAI trained on that data. In a press release from lawyers representing Doe, it explained that Doe’s images were included in a CSAM Hash List maintained by NCMEC, and “that same material” allegedly “was part of the dataset xAI used to build Grok’s image and video generating capabilities.” The complaint similarly only alleged that “CSAM depicting Plaintiff with its longstanding well-known hash values has been used as a part of the dataset used by xAI.”



Most likely from Musk’s personal collection.
They’re using photo checksums to identify known CSAM? As in, you can change a single pixel to fool it? Is that right?
No. You are thinking about cryptographic hashes, there are locality based hashes that give you a kind of similarity metric
That would be very stupid. I’m sure they have smarter algorithms that can handle a picture being resized of cropped.
Smarter, so long as the end user doesn’t do something crazy like adjust hue and a tiny bit of compaction?
I don’t know what they use but i can come up with multiple smarter ways on the spot.
Its probably a multitude of hashes and finding multiple ones in a picture is a red flag.
one i would do is to compare the difference between pixels. You can change the colors all you want. Dark hair versus teeth will always read as opposites per example.
Actually i don’t need to argue that it works… ai generation is always an entirely new picture with new pixels in a set resolution … thats why ai cant do colorisation right.
The fact that the existing hashes where detected in such recreation means the tech is very impressive. I doubt false positives are common with these either.