231: A Paywall Stopped ChatGPT From Reading The Article. It Confidently Summarised It For Me Anyway.

Slobodan "Sani" Manic
Website Optimisation Consultant, No Hacks Founder & Keynote Speaker
CXL-certified conversion specialist and WordPress Core Contributor helping companies optimise websites for both humans and AI agents.
SHOW NOTES
A magazine put a paywall in front of its journalism to stop AI reading it. I tested what that paywall actually checks. It reads one line of your request, the name you type in yourself, and compares it against a list. Say who you are and you get a bill. Make something up and you walk straight in. Then I asked two assistants to read the article anyway. Both got charged, both answered, and neither had read a word of it. It keeps going after the sign-off, about where this podcast is headed.
Chapters
- 00:00 - A story, and a bill instead of a page
- 02:19 - What a 402 is, and who sells it
- 04:02 - What the wall checks, and the name I made up
- 06:46 - robots.txt says one thing, the server does another
- 08:08 - Googlebot walks in free while Penske sues Google
- 10:08 - Is anybody actually paying
- 12:13 - I asked Claude and ChatGPT to read it
- 15:22 - Being read or being cited
- 17:42 - Ask your chatbot where it got the facts
- 19:17 - After the episode: where this is headed
Key Numbers
- 402 Payment Required has been in the web standards since 1992 and almost nothing ever used it
- robots.txt has been a convention since 1994, and nothing makes a robot obey it
- 6,000 or 8,000 or 9,500 websites, depending which of TollBit's own pages you read
- 25 robots named in Variety's robots.txt, 24 of them blocked
- 3 crawlers get charged that the file never names at all
- 0 AI companies named on TollBit's page aimed at AI companies
- 1 paying customer ever named publicly, a news reader app
Key Takeaways
- The paywall checks a name, not a robot. Every crawler that identified itself honestly got a bill. A name I invented, belonging to no company on earth, got the whole page one second later.
- Your server sets your policy and your robots.txt only describes it. Variety's two disagree right now, and three crawlers the file permits are charged anyway.
- Blocking the machine does not block the answer. It decides who gets credited when the machine repeats you. Two assistants credited a search engine and a podcast directory.
What to Do
- Request a page from your own website with a crawler name in the user agent, then with a name you invent, and compare
- Read your robots.txt beside what your server actually returns, line by line
- Check what Google-Extended covers before you rely on it, because it has never covered AI Overviews
- Ask your assistant whether it read the page or searched around it
- Every link and every check I ran is in the newsletter, nohacks.co/subscribe
Sources & Links
No Hacks runs no sponsorships and is funded by advisory and audit work.
If your website needs to work for machines as well as people, start with a fixed-scope Machine-First Architecture audit: https://nohacks.co/audit
ENJOYING THIS EPISODE?
Practical strategies for making your website work for AI agents and the humans using AI to find you. Once a week you get the new articles, the latest podcast episode, and a few links worth keeping.
