For her safety, Doe has opted to receive alerts from the US Department of Justice Victim Notification System any time she may be a victim in a new criminal investigation. Although she has received countless alerts, she was shocked when the CCCP notified her that it had identified AI-generated CSAM on xAI that depicted her. This re-traumatized Doe, whose complaint alleged that messages were found on online forums “between offenders chatting about creating AI generated CSAM of Plaintiff and other similarly situated known, legacy, victims of CSAM.”

Now, Doe fears that xAI has not only made it easier to make more violative images of the most distressing time in her life, but also that xAI allegedly has stored the images that Grok generates and uses those outputs to further train Grok. Because of this, she believes that Grok has been trained on both the initial set of images that have haunted her for more than 20 years and the more recent AI-generated ones.

This is the first case to accuse xAI of training on CSAM, and the complaint does not go into great detail on that claim. Previously, Ars reported on a controversial dataset that was later scrubbed after researchers found CSAM in the training data, but there’s no indication xAI trained on that data. In a press release from lawyers representing Doe, it explained that Doe’s images were included in a CSAM Hash List maintained by NCMEC, and “that same material” allegedly “was part of the dataset xAI used to build Grok’s image and video generating capabilities.” The complaint similarly only alleged that “CSAM depicting Plaintiff with its longstanding well-known hash values has been used as a part of the dataset used by xAI.”

  • Treczoks@lemmy.world
    link
    fedilink
    English
    arrow-up
    13
    ·
    13 hours ago

    If Doe is actually recognizable, then chances are high that actual pictures of her or him were used for training the AI. Wouldn’t surprise me. They scrape the internet for everything they can get their digital hands on, regardless of copyright or criminal law. They are bound to find illegal stuff on that track.

    The next thing is that there are probably confidential data in the training sets, either exposed by neglient users, or by hackers that breached sites and blackmailed them.

    And of course all the copyrighted material they used without permission.

    If AI companies would really get sued on those three illegal sources, they could probably close their doors.

    • matlag@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      4
      ·
      6 hours ago

      Next is a “funny” issue: scraping the web without checking sources will inevitably lead to AI trained on illegal content.

      The only way for them to prevent that is to train their AI to recognize illegal content. And the only way to do that is to feed it that content with label.

      AI corps would have to pay people to watch children sexual abuse and label the videos. (Not that I have the slightest doubt they would proceed if that was allowing them to keep going rampage on web scraping)

      • SleeplessCityLights@programming.dev
        link
        fedilink
        English
        arrow-up
        1
        ·
        2 hours ago

        People have that job already. How do you think that websites where you can upload amateur porn, filter and block CSAM? Someone has to view the flagged material.