• AmyAye@nord.pub
      link
      fedilink
      English
      arrow-up
      4
      ·
      2 days ago

      Yeah, like how they are destroying rare books.

      I mean, you don’t get that back just by asking the content theft machine.

    • jafra@slrpnk.net
      link
      fedilink
      arrow-up
      18
      ·
      edit-2
      2 days ago

      Yeah. I first thought it’s about jpeg, then i read ai and i wasnt sure if its worrying stupidity reporting or bad satire. Edit: i thought '92 i mean

    • Maeve @lemmygrad.ml
      link
      fedilink
      arrow-up
      5
      ·
      edit-2
      2 days ago

      It’s what Apple and MS basically do by uploading all your images to iCloud and showing the name in your images folder.

      Edit to ask: is this what the human brain does when retrieving a memory?

  • whereitsat@lemmy.zip
    link
    fedilink
    arrow-up
    13
    ·
    1 day ago

    brilliant satire that critiques all of magazine journalism.

    i love going to [insert publication here] and reading another article about ‘so and so is ready for their next chapter in life.’ the so and so always an uninteresting, overly wealthy fuckwit that hasn’t accomplished anything other than going to college and having a wealthy parent.

  • DaddleDew@lemmy.world
    link
    fedilink
    arrow-up
    157
    ·
    edit-2
    2 days ago

    Amateur. I can compress entire seasons of a TV series to a few bytes. All I have to do is type its title in Netflix and then BAM, gigabytes of video come out.

  • OldGrayDog@fedinsfw.app
    link
    fedilink
    English
    arrow-up
    49
    ·
    2 days ago

    I’ve read that the Trump administration is hiring him to archive all of the Epstein files using that format.

    • lad@programming.dev
      link
      fedilink
      English
      arrow-up
      12
      ·
      edit-2
      2 days ago

      That’s nice, albeit I want to point out for anyone wondering that this is only conjectured and not guaranteed:

      One of the properties that π is conjectured to have is that it is normal, which is to say that its digits are all distributed evenly, with the implication that it is a disjunctive sequence, meaning that all possible finite sequences of digits will be present somewhere in it.

      There is no guarantee for any specific sequence to appear in π, but for short chunks chances are better (it’s not really a probability, but it’s simpler to say and I can’t explain in details anyway). That’s because (from wiki):

      It is widely believed that the (computable) numbers √2, π, and e are normal, but a proof remains elusive.

  • tal@lemmy.today
    link
    fedilink
    English
    arrow-up
    4
    ·
    1 day ago

    It’s humorous, but last I checked, the best general-purpose compressors with the highest levels of compression—even lossless, which is probably not what most people think of when they think of neural nets—are neural net based.

    Neural net-based compressors are computationally expensive, which is why we don’t normally use them for most day-to-day tasks, but they really can produce really small outputs.

    I’m going to take the text of the US Constitution and stick it in a text file.

    $ wget https://www.gutenberg.org/cache/epub/5/pg5.txt
    $ stat -c %s pg5.txt 
    48326
    

    Okay, so 48326 bytes.

    Let’s do lzo. You’d expect a limited amount of compression — LZO is “fast” compression, usually only used where compression speed is really important, like where you want to be compressing stuff that’s going to be decompressed once and your bottleneck is throughput to disk:

    $ lzop <pg5.txt >pg5.txt.lzo 
    $ stat -c %s pg5.txt.lzo
    24843
    

    Okay, how about gzip? That’s Deflate, an older, but pretty-widely-used general-purpose compression algorithm.

    $ gzip <pg5.txt >pg5.txt.gz
    $ stat -c %s pg5.txt.gz
    16660
    

    Okay, what about LZMA? That’s a newer, more-CPU-intensive thing that’s probably a good general-purpose choice that’ll generally give better compression ratios. It’s the kind of thing that I’d probably use in a lot of cases. (Personally, these days, I tend to use pixz, which provides both indexed access for tarballs and parallel compression and decompression, which is important for modern processors.)

    $ xz <pg5.txt >pg5.txt.xz
    $ stat -c %s pg5.txt.xz 
    15488
    

    Okay, now PAQ, a neural-net-based compressor:

    $ zpaq a pg5.txt.zpaq a pg5.txt -method 5
    $ stat -c %s pg5.txt.zpaq 
    13063
    
    • BradleyUffner@lemmy.world
      link
      fedilink
      English
      arrow-up
      3
      ·
      edit-2
      1 day ago

      “Neutral net based” compression isn’t even in the same universe as “compressed to a prompt” via LLM

      • tal@lemmy.today
        link
        fedilink
        English
        arrow-up
        1
        ·
        edit-2
        1 day ago

        It actually is. I mean, it’s building a dictionary off of a variety of content ahead-of-time, rather than training it on the specific item in question, but that’s not uncommon for non-general-purpose compressors.

        I mean, doing so to a (probably short) prompt is (a) lossy (and I gave a lossless example) and (b) lossy to an extreme degree, to where it’s probably not incredibly useful option for the kinds of systems that exist today.

        But…existing diffusion models aren’t actually intended for this, either. I’d bet that you could train a model to do image compression along these lines, with a large dictionary, that could do usable compression along the lines of what is (jokingly) described in the article. Probably have a larger compressed form than what they’re thinking of.

        EDIT: At one point in time, about over a quarter-century ago now, I went out and banged on a neural net post-processor for JPEG artifacts. The idea here is that JPEG very probably isn’t optimally representing the final image, as a human, using their knowledge of what the world looks like, can manually (if time-consumingly) clean these up. I didn’t meet with a lot of success; I only wanted to put a small amount of time into it, and I was working with much weaker hardware than people are running neural nets on today. But that generated a pre-existing dictionary, a pre-trained neural net, off a training corpus of uncompressed images. It didn’t try to reconstruct the image from scratch, the way something like this would, just clean up artifacts, but it has that same pre-generated neural net approach.

  • TrickDacy@lemmy.world
    link
    fedilink
    arrow-up
    22
    ·
    2 days ago

    a typical jpeg of 20 Mb

    Uh what. Literally should be the highest fucking possible quality from a $10K camera if it’s that big. My raw images aren’t even that big usually.

    • PieMePlenty@lemmy.world
      link
      fedilink
      arrow-up
      14
      ·
      edit-2
      2 days ago

      Its not typical, but you can get 20 Mb+ jpegs out of an entry level 18MP dlsr. Especially if there’s lots of color and at like 5500x3300 resolutions and created with 100% quality preset.
      I checked my immich and I have some (and larger), but yeah, not exactly typical.

      • TrickDacy@lemmy.world
        link
        fedilink
        arrow-up
        2
        ·
        2 days ago

        I am not sure I’ve ever used the 100 quality setting on jpegs. Many years ago I experimented with that a lot and decided that anything over 90 was not different to my eye but the file size was much bigger, relatively speaking. So yeah I suppose if you did use 100, a 20 MB jpeg image is not hard to reach.

    • bstix@feddit.dk
      link
      fedilink
      arrow-up
      11
      ·
      2 days ago

      Raw images is where you go wrong. You need at least 10 mb of metadata tags to achieve professional levels of file sizes. How can you even look at a picture without having a full description of all your childhood memories that lead you to take this beautiful picture of yesterdays mac’'n’cheese dinner. This is why we need more data centers.

  • ShellMonkey@piefed.socdojo.com
    link
    fedilink
    English
    arrow-up
    35
    ·
    2 days ago

    On a similar note, I saw a story a bit back of someone saving input tokens by feeding the bot an image of a wall of text rather than the text itself and having it read the image via OCR.

    Satire and reality are too hard too distinguish these days.

  • Shanmugha@lemmy.world
    link
    fedilink
    arrow-up
    13
    ·
    2 days ago

    Here is another one, no compression:

    • read image
    • send a prompt “regenerate image, here is pixel-by-pixel description”
    • enjoy the result

    (sarcasm)