Half the GEO Problem Is a robots.txt Line Somebody Added in 2024
Before spending anything on content optimization, check whether the engines can reach you at all. A meaningful share of AI visibility problems are access problems, and they are the cheapest ones to fix, often a one-line change that has been sitting there since someone reacted to a news cycle about AI training two years ago.
The three places access gets blocked
robots.txt. In 2024 and 2025 a lot of sites added blanket disallows for AI user agents, frequently copied from a blog post, frequently without distinguishing between the bots involved. The distinction matters: several vendors run separate agents for training-corpus collection and for live retrieval when answering a user's question. Blocking the first is a legitimate business decision. Blocking the second removes you from answers in real time. Teams that intended the first and implemented both are common.
The WAF or CDN. Bot-management rules at the edge will silently return 403s to crawlers that robots.txt permits, and nothing in your analytics will tell you. Managed rulesets are updated by the vendor, so a category that was allowed last quarter can be challenged this quarter without anyone at your company touching a control. If you use a bot-management product, check the actual disposition per user agent rather than the policy you think you set.
JavaScript rendering. Not a block exactly, but the same outcome. If your content only exists after client-side hydration, fetchers that do not execute scripts see an empty shell. This disproportionately hits modern marketing sites built on frameworks where the team assumed server-side rendering was on and never verified it in production.
The decision you should make explicitly
There is a real trade here and it deserves a deliberate answer rather than a default. Allowing retrieval crawlers means your content can appear in answers, driving referral traffic and brand presence. It also means your content is summarized in a surface where the user may never click through. Blocking them protects the content and removes you from the answer.
For most B2B SaaS companies the arithmetic favors allowing retrieval: your content is marketing, not the product, and being absent from the answer does not stop the answer from being generated; it stops it from mentioning you. For publishers whose content is the product, the calculation is genuinely different, and several have concluded that licensing or blocking beats free summarization.
The failure is not picking either side. It is having a policy that blocks retrieval while the marketing team buys a platform to measure why you are not in the answers.
A ten-minute audit
- Read your live
robots.txtand list every AI-related user agent you disallow. For each, confirm whether that agent is used for training-corpus collection, live retrieval, or both. Keep the training blocks if you want them; reconsider the retrieval ones. - Request a key page with
curl, spoofing each of those user agents, and check you get a 200 with real HTML, not a challenge page or a shell. - Diff the
curloutput against what the page shows in a browser. If the served HTML is missing your headings, pricing or product copy, your rendering is the problem, not your content. - Check your edge logs for 403s and challenge responses by user agent over the last 30 days. This is where WAF rules you did not know about surface.
None of this improves how persuasive your content is. It just establishes whether the content is reachable, which is the precondition for everything else, and the step most audits skip because it belongs to infrastructure rather than marketing.