For her safety, Doe has opted to receive alerts from the US Department of Justice Victim Notification System any time she may be a victim in a new criminal investigation. Although she has received countless alerts, she was shocked when the CCCP notified her that it had identified AI-generated CSAM on xAI that depicted her. This re-traumatized Doe, whose complaint alleged that messages were found on online forums “between offenders chatting about creating AI generated CSAM of Plaintiff and other similarly situated known, legacy, victims of CSAM.”
Now, Doe fears that xAI has not only made it easier to make more violative images of the most distressing time in her life, but also that xAI allegedly has stored the images that Grok generates and uses those outputs to further train Grok. Because of this, she believes that Grok has been trained on both the initial set of images that have haunted her for more than 20 years and the more recent AI-generated ones.
This is the first case to accuse xAI of training on CSAM, and the complaint does not go into great detail on that claim. Previously, Ars reported on a controversial dataset that was later scrubbed after researchers found CSAM in the training data, but there’s no indication xAI trained on that data. In a press release from lawyers representing Doe, it explained that Doe’s images were included in a CSAM Hash List maintained by NCMEC, and “that same material” allegedly “was part of the dataset xAI used to build Grok’s image and video generating capabilities.” The complaint similarly only alleged that “CSAM depicting Plaintiff with its longstanding well-known hash values has been used as a part of the dataset used by xAI.”



This should make the entire model illegal. Train it on illegal stuff: delete the whole model.
The US legal system still runs on powdered wigs and good ol boys and closed door midnight meetings. It didn’t catch up to the technology booms during our childhood, and it’s sure as shit not catching up to anything happening now, hence the unbridled tech bro dystopia currently enshitifying everything you can possibly imagine.
In all seriousness, there are some very interesting legal questions that will be raised if this case makes it that far.
The problem is that there’s no existing law that would effect this on its own. To my knowledge, no country in the world has a law on the books specifically dealing with AI models trained on CSAM. So the question, under existing laws, would turn on whether the data stored within the model itself would constitute CSAM.
The problem, in no small part, is that we have serious gaps in our public consensus knowledge about how LLMs actually work.
There’s a case that, AFAIK, is still being argued in Germany pushing the theory that LLMs actually do, in effect, store a copy of all their training data, just in a compressed form. This certainly seems to hold some water given both the tests they relied on, and the situation with this Jane Doe where the model produced images so alike to real images of her that they tripped hash detections.
The German case argues that this is analogous to the difference between an MP3 and a WAV, or a JPEG and a PNG. That sharing a lossy copy of a work is no less infringing just because it’s imperfect.
If the underlying claim - that LLMs function as a form of lossy compression - can be substantiated then there would be a real argument that the model itself would constitute CSAM. Since there would be no realistic method that I’m aware of for removing the offending material from the model - and presumably SpaceX would have to somehow prove that they’ve done so - that would make the entire model contraband. They’d have to retrain on a clean dataset.
Of course I said “if the case makes it that far” at the top because I don’t think it will. SpaceX will do anything and everything to avoid handing over meaningful discovery in this case, including, I suspect, outright destruction of evidence. If there is anything that actually proves that they used CSAM in the training data then they are so far beyond fucked that there’s simply no downside to further illegality in pursuit of concealing their crimes. They have the world’s wealthiest asshole in a position to throw literal billions at making this go away. I genuinely wouldn’t be surprised if people turn up dead off the back of this if that’s what it takes.
CSAM laws are already overreaching in dimensions that are rife with abuse. It’s rare to have a law where owning a picture of something is highly illegal. It’s the only law I know of where somebody can just send you a picture on your phone, and suddenly, you’re breaking the law. You could be arrested and thrown in prison for even admitting that somebody else sent you the picture.
Police use CSAM as an excuse all the time to search without a warrant. Remember the duress passcode case? Police claimed he had CSAM on his phone, when everybody knew they were targeting his activist work.
And what should we really care about? The CSA. Go after the CSA. Focus on the abuse. The Epstein files are right over there.
Yeah, it’s absolutely valid to question the degree to which this is an outcome that we should even want. If someone hides a CSAM image in the Linux kernel should Linux become illegal?
I think there is a valid distinction to be drawn in this case, because removing individual components of an LLM isn’t really something we know how to do. So there’s a fair argument that a model which is built using illegal content should be illegal, and if that means they have to completely retrain from scratch, so be it.
But yes, I’d want to be very careful about the lines around a law or ruling like that and exactly what it’s extent is. Child sex crimes and child safety are topics that tend to short-circuit all reasonable objections, and are frequently exploited as a means of getting bad laws onto the books. Bill C-22 up here in Canada is a great recent example.
Globally, there’s powers at work to get all kinds of “age verification” laws on the books, which we all know is just a thin excuse for de-anonymization and identity gathering. These same powers were fucking with Steam and Itch.io’s payment methods, because “what about the children”, which then ties to similar events with PornHub a few years earlier. Congress passed three different “Internet child safety” laws in the 2000s and 2010s, and all three were struck down by the Supreme Court for being unconstitutional. Decades earlier, DARE abused “child safety” justifications for personal gains.
This excuse is used all the time, and the emotional weight it carries makes it shocking effective to the proles that buy it at face value.
It’s illegal to posses. Anyone training the models on that is criminally liable for possession as is the company. And it’s a conspiracy to posses that was directed by someone, which is now RICO. Will they prosecuted? No.
Sure. Never argued against that. I was discussing the assertion that the model itself would be illegal as a result. Different thing.
If it was trained on it, that means they are in possession of it, which that right there is straight to jail. I have a feeling they’re scrubbing everything they can right now as we chat
Was going to say basically say the same exact same thing. There is no law saying how AI is handled when using stollen materials or other “illegal” content.
But somehow we have all been brought to believe that somehow “new” technology isn’t subject to existing laws.
If x or any other company downloaded CSAM everyone in the company should be arrested.
I assume you cannot possibly mean that as written, right?
I’m absolutely for arresting anyone who was involved with this, or had knowledge that it was happening. But we’re obviously not talking about going after Jane the intern here, right?
I won’t say everyone. I’ve actually been at a company who was investigated (not for CSAM, but other things that happened). I had no idea it even happened, and luckily was not involved with any of it. So for me no, I wouldn’t have wanted that. That being said 2 things, say I had been in the position. If our scraper was downloading it and it was my scraper, damn right I would have flagged it to legal, HR, and everyone I could have, along with writing some way to prevent it, and written everything down in a complete log (off company computer). If it wasn’t stopped immediately I would either quit, whistleblowed, or happily talked with anyone raiding and making sure any of the decision makers were hauled off. I don’t blame someone for being lowest level at a shit company, been there. (Although I will say, xAI, come on, no one is “stuck” there, but that doesn’t mean that Dave the brand new intern out of college should be hauled off). Who I blame are the suits who were probably told it was happening and chose to ignore it, and any engineer I do blame if they knew about it and chose not to do one of the above.
As an engineer I’ve had my fair share of let’s say… challenges that I’ve had to morally grapple with. Things I’ve been asked to do that may not be moral. However, there’s a pretty wide chasm between “Implement this dark pattern so people won’t unsubscribe” and “host this and don’t tell anyone”
If you started a company that made CSAM do you think you and every one involved should be let off, because you claim it was a technological oopsie?
People have been pointing out that xAI is a CSAM machine since basically the first day it came out. And when they basically said they don’t care. All the employees that stayed are all accomplices to the crimes… let alone the people who started working there after.
The only way to stop these companies is to start holding people accountable. And they will never hold the people at the top to account.
So make it hurt for everyone involved and then people will think twice before they sign up to do evil deeds for evil people.
But you didn’t say “The people at the top”, you said “Everyone.” I think you need to take a moment to figure out what you’re actually arguing for here.
Again, I would happily see anyone who had knowledge of this arrested. They either supported it, or knew of it and said nothing. And yeah, we can throw in anyone who maintained wilful ignorance too. If that includes Elon himself, so much the better. We already know the dude is a fucking pedophile, maybe this is how they’ll finally nail him.
But if you’re arguing for arresting the cafeteria lunch guy over this, that is an insane position to hold.
So which is it?
I have been thinking about it and I agree it’s an extreme position.
But these deeds are not being done by one person. And at some point whether you are an active participant or not you are continuing to work there so you have some level of complicity.
If you came into work today to find your building full of feds packing up evidence, because it turned out that, entirely unbeknownst to you, five members of the C-suite had been running a child sex ring, would you consider yourself “complicit”?
deleted by creator
God help us all
Models do not store the original, they store the model, which is a huge graph of probabilities.
I’m aware. But if you read my explanation a little more carefully, you’ll see that the argument being made is that this de facto constitutes a form of lossy compression.
The easy comparison is that a JPEG does not store an “original” image, but it contains information that can be used to almost perfectly recreate that image, with the help of a little math.
If the same argument can be said to hold true of LLMs - and yes, that is very much a load-bearing “if” - then they would constitute a form of lossy compression.
How far does one take that? A picture of a horse could be transformed into Abraham Lincoln with the right algorithm. Is that lossy compression?
That’s what courts exist to decide.
Not in the US though.
I think there’s no real change that needs to be made to the law. The images are stenographically embedded into these models. There exists some generic input data that results in the reproduction of the embedded data.
This is not really any different than a password protected zip file.
Well, like I said, that comes down to whether, legally, that stenographic embedding* would constitute “reproduction” or not. That’s what the German case hinges on. Obviously not relevant to US law, but a) a similar case could be made in the US, and b) X operates in the EU.
To play devil’s advocate, SpaceX would most certainly argue that what they’re doing is equivalent to storing a hash, like how Microsoft’s PhotoDNA system works. PhotoDNA can detect CSAM without storing CSAM because it only stores the image hashes, not the images themselves. So there’s pretty clear legal precedent for them to point to.
(NB: There has been work done by security researchers on reverse engineering images from hashes, so even that isn’t absolute.)
It’s a legally complex area where we’re likely to see case law evolving rapidly.
(ETA:) *Note that this is not actually the correct term for how the data is being stored. See discussion below. I just didn’t want to derail things by getting into it here.
You can’t reproduce an approximation from any type of hash, so that argument is dead in the water.
Do you understand what I mean by stenographically embedded?
https://www.pseudodna.eu/. Scroll down to “Hash reversal” where they demonstrate the technique.
I took it to be an imperfect attempt to describe more broadly the way that data is mathematically encoded into LLMs.
Technically, stenography would require that the original be retrievable, since stenographic embedding is the process of concealing one thing inside another. Stenos, from the Greek “covered”. Personally, I’d argue that to conceal, you have to be able to reveal. If I throw a photograph into a fire I haven’t hidden the image in the fire. Modern stenographic image embedding techniques use methods of encoding data into another dataset without visibly altering the second set, with the intent being that that data can later be retrieved by someone who knows that it’s there (eg, least significant bit, where you change only the “1” bit of each pixel. This imperceptibly shifts the colour values of the image to a human viewer, but allows you to read out that stored data at a later time).
Now, since your argument rests on the exact opposite, that the data is not retrievable, I simply accepted the term as a “close enough” approximation for what I believe we’re both talking about - the extremely complex multidimensional data arrangement that forms the core of an LLM - and carried on from there because I find that sometimes it’s better to just roll with a person’s choice of language rather than quibble over it.
But since you clearly feel that your meaning was either improperly expressed, or improperly understood, you’re welcome to elaborate.
Wow they fucked that up so badly.
My understanding of hashing is a math function that reduces information.
The models and apparently that piece of software have retrievable information, so using that software to generate these images should count as reproduction under existing law.
Close, but not quite. A hash is better understood as a “fingerprint” for a given piece of data. Any given input will produce the same length of output, so it doesn’t really “reduce” information - a hash of “Hello world” would actually be significantly longer - but rather it identifies a matching piece of data.
This has a number of uses. For example,if you have a password system on a website, you don’t store a user’s actual password. That would be terrible for security. Instead you store the hash (yes, salted, for that one pedant who was about to interject). Then when the user enters their password to log in, you hash the entered password and compare it to the stored hash. If they match, you know they’re the same string. You can also create a hash of a file and provide that along with the file itself. If the receiver also hashes the file, they can then compare their hash to yours; if they don’t match, the file is corrupt or has been tampered with. You’ll often see this referred to as a checksum.
Technically, the reproducible data from a hash should be zero, so they don’t so much contain information as verify it. But as we’ve seen from that research, there are ways around that (not the only ones, to be sure). But the idea in theory is that you produce a hash through a one-directional algorithm; very easy to compute A from B, but very, very hard to compute B from A.
It should be noted that this is very different from stenography, which is a variety of systems for storing data, not fingerprinting it. Hashing is also distinct from encryption as a means of securing data, because encryption is intended to be reversible, hashing is not.
I’ve seen and used fixed length output on hashes, I assumed that is more common than variable length output.
I would not trust variable length output, and it looks like they screwed up their math enough that the function is reversible.
Either people end up dead and yet it’s still somehow a nothing burger or we don’t even make it that far
Yeah, the stakes are just too high for SpaceX to let this get to trial. Any amount of illegality becomes worth it when you consider the alternative.
The only other possibility I can see is that they pull a Bungie; “perform an internal investigation”, find an intern to blame for everything, and throw a huge settlement at Jane Doe. But even that would require an absolutely insane cover up to pull off.
Yep, shut that shit down. Also, pay her a shit ton of money.