Stay discoverable in search while disallowing AI training

(blog.cloudflare.com)

37 points | by djfergus 2 hours ago

13 comments

  • 1vuio0pswjnm7 18 minutes ago
    "Cloudflare classifies bots by behavior, and a single bot can exhibit more than one behavior."

    Is that really true

    CF classifies anyone not using a popular browser with Javascript enabled as a "bot"

    CF fingerprints www users

    It seems CF is like Mozilla, deeply supportive of ad tech

    Online advertising is their first priority, www users, e.g., ones who dislike ads and tracking, are not important

    • devmor 1 minute ago
      As far as I can tell, after months of fighting being DDoSed by Anthropic and OpenAI across 50+ sites - Cloudflare also allows what it considers "good bots" through all of your bot blocking rules, with no option to turn this off unless you pay them money.
  • nirmeetimthebes 23 minutes ago
    "Accountable" is just a fancy word for "pinky promise, but with a label." Nothing stops the data from ending up in a training run once it's already been fetched.
  • skybrian 1 hour ago
    I didn't know websites could opt out of providing data to Google's AI training. Looks Google added support for this via 'Google-Extended' in robots.txt back in 2023:

    https://blog.google/innovation-and-ai/products/an-update-on-...

  • kinduff 1 hour ago
    I maintain a cloud IP ranges database, and I'm going to test this out.

    I have my doubts, though. A formal title like "Accountable" (capitalized) sounds deliberate, but I can't help imagining the renewal email:

    "Hey, want to renew your Accountable™ license? Just pinky promise again that you use your IPs for what you say you do."

  • mskalski 1 hour ago
    I wonder if protocols like Web Bot Auth [1] will see wider adoption. At least as a supported mechanism for those bots which identify themselves. The rest probably still have to be treated with Anubis. In my free time I've recently been experimenting with a Web Bot Auth implementation as an Envoy dynamic module [2] to have a way to define some additional policies for the traffic from bots.

    [1] https://datatracker.ietf.org/doc/draft-ietf-webbotauth-https... [2] https://github.com/michalskalski/envoy-web-bot-auth

    • fraywing 1 hour ago
      I'm playing around with it for my MCP hiring protocol ojcp[1] and it seems to work very well for signing attestations at the header level.

      [1] https://github.com/ojcp-org/ojcp

      • mskalski 54 minutes ago
        Thanks for sharing, it is interesting use case
  • dzhiurgis 11 minutes ago
    If this admin is serious about AI growth they’d make anti-scrapping illegal.
  • AnonC 1 hour ago
    > Accountable mixed-use crawlers remain allowed for search. Every other training crawler is blocked, including the training-only crawlers run by Amazon, Anthropic, Meta, and OpenAI — blocking those does not affect search.

    > We also categorize the relevant crawlers from Amazon, Anthropic, Meta, and OpenAI as Accountable. These organizations separate their Search and Training crawlers, so Cloudflare can block the Training crawler without affecting search.

    I find it difficult to trust that either Meta or OpenAI would use their separate search and training crawlers only for the respective purposes. Their pinky promises have no value, IMO. Both companies are premised on deceptive behaviors.

  • zergrush 1 hour ago
    what weirds me out is the analytics

    theres no way my index.html page with nothing is getting 10000 hits a day

    wtf?

    • itake 1 hour ago
      7 per minute.

      I use my high school’s website to test Internet connectivity bc the domain is short and they don’t do a TLS redirect (making it easy to detect WiFi portals).

  • gleezard 1 hour ago
    A bit too late honestly (?). With so many people who have shifted over to reading AI summaries as a primary search response, those with AI-enabled sites will win by attrition.

    There is no going back from this. And the internet is a relatively new phenomenon. Recklessly, blindly applying ads to pages in hopes of generating revenue is a very silly thing to do. Technology with ad blockers and now AI summaries has taken that away. New business models, perhaps actually decent ones are required.

    Death of ads everywhere? Good fucking riddance.

    Posted from LibreWolf.

    • tracerbulletx 1 hour ago
      All this attitude does is tear down the only viable income source for independent publishers and demonizes them for trying to make money, while everyone let's huge corporations off the hook for it because "well that's just what they do"
      • octoberfranklin 13 minutes ago
        Independent publishers can't afford to run an ad network.

        You aren't independent. You work for the BigTech company that serves ads on your site.

      • gleezard 55 minutes ago
        Advertising in the way it’s done is demonic in and of itself. I don’t care - find a better business model.
  • gchamonlive 2 hours ago
    Cloudflare, enabling the problem and the solution since, how long has it been?
    • arm32 1 hour ago
      (checks watch)

      17 years

      Fun lava lamp story, though

  • India_InfraNote 25 minutes ago
    [flagged]
  • aaron695 16 minutes ago
    [dead]
  • arm32 1 hour ago
    What does this new setting actually do? Does it block their IP ranges too, I hope? As if Meta, for instance, is actually going to respect Accountable, via themselves or their partners, quite frankly is eyebrow raising at best.