• hirihit640@sh.itjust.works
    link
    fedilink
    English
    arrow-up
    1
    ·
    34 minutes ago

    Anubis should add an option to “trust” certain tokens or public keys, so good scrapers like Internet Archive can get past.

  • ISO@lemmy.zip
    link
    fedilink
    arrow-up
    1
    ·
    8 hours ago

    --user-agent "Mozilla/5.0 (X11; Linux x86_64; rv:154.0) Gecko/20100101 Firefox/154.0 Scrapping-For-Research-Hope-Its-Okay"

    Here, It’s disclosed, no?

  • hendrik@palaver.p3x.de
    link
    fedilink
    English
    arrow-up
    4
    ·
    edit-2
    15 hours ago

    I feel at this point it doesn’t even matter. I’ve recently renovated parts of my IT infrastructure… And I mean they’re not only harming news outlets and Wikipedia… Everyone is under attack. They’re scraping my sites, searching for all kinds of stuff until the database process catches on fire. They put weird stuff into every web form, send plenty spam mails… Sign up random people for your newsletter, unsubscribe people from your newsletter… The OpenClaws are able to solve the CAPTCHAS and/or sign up… It’s crazy. Law aside, I don’t think I have any option but to block them myself. This will also block legitimate crawlers because I have to put in place some drastic measures with some services. But at least I can continue to run my website.

    I see other people do the same. And I also run into rate limiting all the time. Back in the day, my Home Assistant could just query trashday, whether the train is on time, which flights are overhead… I could just run yt-dlp to save some YouTube videos. That’s all becoming more and more difficult because of all the countermeasures. There’s still some chance to force people to sign up and offer some protected API…

    But whether someone passes a law or doesn’t… I think the internet admins are still going to (have to) kill the bots.

    • hirihit640@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      2
      ·
      37 minutes ago

      Depressing. And this is all leading to human verification. The government has always wanted it, and now the people want it too. Say goodbye to privacy.

      • hendrik@palaver.p3x.de
        link
        fedilink
        English
        arrow-up
        1
        ·
        edit-2
        18 minutes ago

        Seems to me like the big tech companies are way ahead anyway. Google probably has enough data collected to know whether someone (or something) is eligible to watch a video. At least they distinguish based on network, maybe behaviour and in doubt they’ll force me to log in and present my session cookie. That’s all bad for privacy as is.

        I wouldn’t be surprised if they follow up revamping the captchas. Maybe in the future it won’t make you click on pictures with an omnibus in it, but directly talk to the Chrome browser on your device and make that (or some hardware chip in the computer/phone like a TPM) verify you’re a human. Of course that one might as well contain a unique ID and an advertiser ID… I think I read on the news how they’re considering to redo the captchas. Not sure if it’s going to end up like that. But my best bet is it’s gonna be bad in some way or another. Perfect opportunity to tie something down for the users, remove use-cases that aren’t consume only, and track them some more.