All Episodes
23 minEpisode 231

231: A Paywall Stopped ChatGPT From Reading The Article. It Confidently Summarised It For Me Anyway.

SpotifyApple
HOSTED BY
Slobodan "Sani" Manic

Slobodan "Sani" Manic

Website Optimisation Consultant, No Hacks Founder & Keynote Speaker

CXL-certified conversion specialist and WordPress Core Contributor helping companies optimise websites for both humans and AI agents.

SHOW NOTES

A magazine put a paywall in front of its journalism to stop AI reading it. I tested what that paywall actually checks. It reads one line of your request, the name you type in yourself, and compares it against a list. Say who you are and you get a bill. Make something up and you walk straight in. Then I asked two assistants to read the article anyway. Both got charged, both answered, and neither had read a word of it. It keeps going after the sign-off, about where this podcast is headed.


Chapters

  • 00:00 - A story, and a bill instead of a page
  • 02:19 - What a 402 is, and who sells it
  • 04:02 - What the wall checks, and the name I made up
  • 06:46 - robots.txt says one thing, the server does another
  • 08:08 - Googlebot walks in free while Penske sues Google
  • 10:08 - Is anybody actually paying
  • 12:13 - I asked Claude and ChatGPT to read it
  • 15:22 - Being read or being cited
  • 17:42 - Ask your chatbot where it got the facts
  • 19:17 - After the episode: where this is headed


Key Numbers

  • 402 Payment Required has been in the web standards since 1992 and almost nothing ever used it
  • robots.txt has been a convention since 1994, and nothing makes a robot obey it
  • 6,000 or 8,000 or 9,500 websites, depending which of TollBit's own pages you read
  • 25 robots named in Variety's robots.txt, 24 of them blocked
  • 3 crawlers get charged that the file never names at all
  • 0 AI companies named on TollBit's page aimed at AI companies
  • 1 paying customer ever named publicly, a news reader app


Key Takeaways

  1. The paywall checks a name, not a robot. Every crawler that identified itself honestly got a bill. A name I invented, belonging to no company on earth, got the whole page one second later.
  2. Your server sets your policy and your robots.txt only describes it. Variety's two disagree right now, and three crawlers the file permits are charged anyway.
  3. Blocking the machine does not block the answer. It decides who gets credited when the machine repeats you. Two assistants credited a search engine and a podcast directory.


What to Do

  • Request a page from your own website with a crawler name in the user agent, then with a name you invent, and compare
  • Read your robots.txt beside what your server actually returns, line by line
  • Check what Google-Extended covers before you rely on it, because it has never covered AI Overviews
  • Ask your assistant whether it read the page or searched around it
  • Every link and every check I ran is in the newsletter, nohacks.co/subscribe


Sources & Links

No Hacks runs no sponsorships and is funded by advisory and audit work. 

If your website needs to work for machines as well as people, start with a fixed-scope Machine-First Architecture audit: https://nohacks.co/audit

ENJOYING THIS EPISODE?

Practical strategies for making your website work for AI agents and the humans using AI to find you. Once a week you get the new articles, the latest podcast episode, and a few links worth keeping.