For her safety, Doe has opted to receive alerts from the US Department of Justice Victim Notification System any time she may be a victim in a new criminal investigation. Although she has received countless alerts, she was shocked when the CCCP notified her that it had identified AI-generated CSAM on xAI that depicted her. This re-traumatized Doe, whose complaint alleged that messages were found on online forums “between offenders chatting about creating AI generated CSAM of Plaintiff and other similarly situated known, legacy, victims of CSAM.”
Now, Doe fears that xAI has not only made it easier to make more violative images of the most distressing time in her life, but also that xAI allegedly has stored the images that Grok generates and uses those outputs to further train Grok. Because of this, she believes that Grok has been trained on both the initial set of images that have haunted her for more than 20 years and the more recent AI-generated ones.
This is the first case to accuse xAI of training on CSAM, and the complaint does not go into great detail on that claim. Previously, Ars reported on a controversial dataset that was later scrubbed after researchers found CSAM in the training data, but there’s no indication xAI trained on that data. In a press release from lawyers representing Doe, it explained that Doe’s images were included in a CSAM Hash List maintained by NCMEC, and “that same material” allegedly “was part of the dataset xAI used to build Grok’s image and video generating capabilities.” The complaint similarly only alleged that “CSAM depicting Plaintiff with its longstanding well-known hash values has been used as a part of the dataset used by xAI.”



Yeah, it’s absolutely valid to question the degree to which this is an outcome that we should even want. If someone hides a CSAM image in the Linux kernel should Linux become illegal?
I think there is a valid distinction to be drawn in this case, because removing individual components of an LLM isn’t really something we know how to do. So there’s a fair argument that a model which is built using illegal content should be illegal, and if that means they have to completely retrain from scratch, so be it.
But yes, I’d want to be very careful about the lines around a law or ruling like that and exactly what it’s extent is. Child sex crimes and child safety are topics that tend to short-circuit all reasonable objections, and are frequently exploited as a means of getting bad laws onto the books. Bill C-22 up here in Canada is a great recent example.
Globally, there’s powers at work to get all kinds of “age verification” laws on the books, which we all know is just a thin excuse for de-anonymization and identity gathering. These same powers were fucking with Steam and Itch.io’s payment methods, because “what about the children”, which then ties to similar events with PornHub a few years earlier. Congress passed three different “Internet child safety” laws in the 2000s and 2010s, and all three were struck down by the Supreme Court for being unconstitutional. Decades earlier, DARE abused “child safety” justifications for personal gains.
This excuse is used all the time, and the emotional weight it carries makes it shocking effective to the proles that buy it at face value.