How often do AI-cited pages block access?
Rarely. Of 154,986 AI-cited pages checked at or after their latest recorded citation, 2,435, or 1.57%, returned an authentication-required, forbidden, or rate-limit response.
By Dimitry Apollonsky · July 17, 2026 · 9 min read
Contents
- Only 1.57% of checked AI-cited pages limited access
- 97.52% returned a successful, redirect, or not-modified response
- Forbidden responses account for 91.83% of access-limited pages
- Google AI Mode's cited pages had a higher rate
- Government, directory, news, and social pages had the highest large-group rates
- Reddit and Quora had the highest rates in the domain table
- Only 1.38% of pages with an access-limited result later returned a normal response
- What the count leaves out: pages we could not check in time
- What marketers should do
- Get the data
- Sources
- Related research
In one observed cut of Parse data, we analyzed the latest HTTP result for 154,986 unique pages across 67,005 domains, cited 322,681 times in 382,708 AI answers to 16,994 organic prompts on ChatGPT Search and Google AI Mode from May 24 through July 17, 2026.
Only 1.57% of checked AI-cited pages limited access
An access-limited page is a cited URL whose latest eligible check returned HTTP 401, 403, or 429. We kept one latest result per page and required that check to occur at or after the page's latest recorded citation. Of 154,986 eligible pages, 2,435 limited access. The exact rate was 1.5711%.
Ahrefs reports that 67% of
ChatGPT's top citations are off-limits to marketers. Its study defines off-limits as not pitchable or influenceable through traditional outreach. It does not mean 67% of cited servers refused a request. Audit outreach opportunity and direct page access as separate questions.
Takeaway
97.52% returned a successful, redirect, or not-modified response
The latest eligible result was a successful response for 144,774 pages, or 93.4110%, and a redirect or not-modified response for 6,367, or 4.1081%. Together they account for 151,141 pages, or 97.5191%. Access-limited results account for 1.5711%, page-not-found or gone results for 0.6626%, server errors for 0.2045%, other 4xx responses for 0.0419%, and other responses for 0.0006%.
A page can be hard to influence, omitted from later answers, or disallowed for a named crawler while still returning a normal response to a direct check. Semrush measures declared crawler rules in robots.txt. This study measures the direct page response after a citation.
Takeaway
Forbidden responses account for 91.83% of access-limited pages
HTTP 403 accounts for 2,236 of the 2,435 access-limited pages, or 91.8275%. HTTP 429 accounts for 185, or 7.5975%, and HTTP 401 accounts for 14, or 0.5749%.
Most access-limited results were explicit refusals rather than rate limits or authentication prompts. The status code identifies the next diagnostic step, but it does not reveal the publisher's intent or the rule that produced the response.
Google AI Mode's cited pages had a higher rate
Google AI Mode had 2,249 access-limited pages among 136,583 checked pages, or 1.6466%.
ChatGPT Search had 226 among 20,095, or 1.1247%.
Keep the two engines separate during a citation audit. The difference is descriptive because the same page can be cited by both engines and the engines cite different sources.
Government, directory, news, and social pages had the highest large-group rates
Among source groups with at least 1,000 checked pages, government pages were at 4.0189%, directories 3.0335%, news 2.3683%, social 2.3382%, health 2.0268%, and general web pages 2.0144%. SaaS pages were at 0.7345% and developer pages at 0.1805%.
Source type changes the baseline, but no displayed group reached 5%. Use the group rate to prioritize checks, not to assume that every page in the group limits access.
Reddit and Quora had the highest rates in the domain table
The table requires at least 50 checked pages and at least three access-limited pages. Reddit had 55 access-limited pages among 63 checked pages, or 87.3016%. Quora had 49 among 117, or 41.8803%. Byrdie was at 20.5882%, Food & Wine at 15.8730%, Serious Eats at 10.9589%, and
NerdWallet at 8.8235%.
Domain behavior is more actionable than the aggregate rate. Recheck the exact cited URL and response before treating a whole source as unavailable.
| Social | 63 | 55 | 87.302 | |
| Community | 117 | 49 | 41.88 | |
| General web | 68 | 14 | 20.588 | |
| General web | 63 | 10 | 15.873 | |
| General web | 73 | 8 | 10.959 | |
| General web | 136 | 12 | 8.824 | |
| General web | 91 | 6 | 6.593 | |
| General web | 66 | 3 | 4.546 | |
| Social | 498 | 12 | 2.41 | |
| General web | 904 | 8 | 0.885 |
Takeaway
Only 1.38% of pages with an access-limited result later returned a normal response
The eligible window contained 2,469 pages that returned 401, 403, or 429 at least once. On the latest check, 2,435 were still access-limited and 34 returned a successful, redirect, or not-modified response. The later-normal-response rate was 1.38%.
A recheck can clear a temporary response, but most access-limited pages in this cut remained limited on the latest result. Confirm the latest response before escalating to the publisher or changing content strategy.
Takeaway
What the count leaves out: pages we could not check in time
The two-engine corpus contained 1,823,717 cited pages. We excluded 1,557,618 without a check in the window, 14,854 without a valid HTTP result, and 96,259 checked before their latest recorded citation. That left 154,986 pages.
The 1.57% result applies only to eligible checked pages. A strict 403-only cut was 2,236 of 154,986, or 1.4427%. An expanded definition that also includes 402, 406, and 451 was 2,453 of 154,986, or 1.5827%. The result should not be applied to cited pages Parse did not check, to robots.txt policy, or to the access behavior of an AI engine.
| No valid HTTP result | 14,854 | |
| No check in the window | 1,557,618 | |
| Included pages | 154,986 | |
| Expanded access-limited cut | 2,453 | 1.583 |
| Cited pages in the corpus | 1,823,717 | |
| Check before the latest citation | 96,259 | |
| 403 only | 2,236 | 1.443 |
| 401, 403, or 429 | 2,435 | 1.571 |
What marketers should do
The observed cut isolates 2,435 pages with a latest authentication-required, forbidden, or rate-limit response. The domain and source-group results show where those pages concentrate.
Check three things separately. Read robots.txt for declared crawler rules, request the exact cited URL to see the current response, and rerun the prompt to see whether the engine still cites it. Fix only the one that failed, then run the same check again.
Get the data
Sources
- Ahrefs: 67% of ChatGPT's top 1,000 citations are off-limits to marketers · accessed 2026-07-17
- Semrush: Blocked from AI Search in Site Audit · accessed 2026-07-17
- Consent in Crisis: The Rapid Decline of the AI Data Commons · accessed 2026-07-17
- Cloudflare: From Googlebot to GPTBot · accessed 2026-07-17
- Do generative AI assistants respect robots.txt? Tracing web access beyond visible answers · accessed 2026-07-17