Check what is allowed
We show which AI visitors your site welcomes, which it turns away, and where two instructions disagree.
See what AI gets when it visits your site — and the exact reason it cannot read something.
We check your page, what search visitors can read, and whether the words and answers are easy to use.
CrawlProof does not guess what an AI will say or promise a citation. It reads your public page, compares the responses, and shows the proof in plain language.
No signup. No model tokens. Paste a public page and see what AI is allowed to read — and whether it actually receives useful content.
The report below is a finished scan of a real public page. It shows what the site tells AI visitors, what they actually receive, and the proof behind each warning.
source: www.notion.so / scan captured 09 Aug 2026
In our measured sample, sites that blocked AI often blocked search and training together. CrawlProof shows the cost plainly: you blocked AI training and also removed yourself from AI search — then gives you the configuration that keeps search access while staying out of training.
OPENAI / entirely blocked incl OAI-SearchBot: 8/75
ANTHROPIC / entirely blocked: 13/75
TRAINING ONLY / OpenAI: 10/75 · Anthropic: 9/75
75 sites · root-path robots sample · CrawlProof measurement · not a web-wide prevalence estimate.
Plain answers about what your site says, what AI visitors receive, and what a buyer can understand quickly.
We show which AI visitors your site welcomes, which it turns away, and where two instructions disagree.
We compare the page a normal visitor sees with the page an AI-style request receives.
We look for clear answers, useful headings, real text, structured data, and sections that make sense on their own.
Each result points to the page, response, or instruction behind it. No guessed score and no promise of a ranking.
Small checks, assembled into a report an owner or agency can explain to a client.
We compare the page a browser sees with the page each crawler receives. A 200 response can still be an empty app shell.
We join robots rules, noindex headers, page tags, canonicals, sitemaps, and redirects instead of listing them separately.
We also check page tags and server headers that can say “do not list this page” even when the main rules look fine.
A page can look fine in a browser but send almost nothing to a text-only visitor. We show that difference instead of calling it healthy.
We flag headings that sound like slogans and show page-derived prompt ideas. They come from your words; they are not predictions of what AI will say.
We look for short passages that start with “it”, “this”, or “they” and may lose their meaning when copied without the rest of the page.
We check useful page types, names, dates, sources, numbers, definitions, and comparisons without claiming they guarantee visibility.
Agency scans check a list of client or prospect sites, sort the real issues first, and give you a report you can hand to a client.
Paid users can save a site, run it again no more than once a day, compare the result with the last run, and review alert records when access, findings, or structured data change. No email is sent.
We keep a sourced, dated registry and label its entries honestly: a fixed subset is live-probed, while the rest are rules-only checks. More registry rows do not create an unbounded request bill.
“20 minutes per client per week manually Googling their brand in ChatGPT and Perplexity, screenshotting results, and pasting them into a Google Doc. That is not a workflow. That is desperation.” — Freelance SEO managing nine client sites, Indie Hackers user report; source
These are the settled comparison figures from our research. Third-party captures remain labelled as such.
| Tool | Price observed | What it does | What CrawlProof does instead |
|---|---|---|---|
| Profound | $99–$399/mo | LLM visibility and citation tracking | One-time technical evidence, no model inference |
| Peec AI | ~$95+/mo, third-party capture | AI search monitoring | AI visitor access and page-content checks |
| AthenaHQ | ~$295/mo, third-party capture | Answer-engine analytics | Robots and HTTP evidence without a subscription |
| Semrush AI Visibility Toolkit | $99/mo per domain | AI visibility and an AI-readiness site audit | Focused deterministic audit, $69 once |
| Ahrefs Brand Radar | from $199/mo | Brand visibility in AI answers | No citation tracking, no recurring model cost |
| ZeroRank AI | $89 LTD observed | AI-search citation monitoring | Technical half of GEO, not citation monitoring |
| CrawlProof Solo | $69 once | AI access checks, content differences, hidden blocks, fix files | What they do that we don't: citation tracking. |
Paid access is capacity and deliverables, not a metered scan. The ten included materials are practical files you can use after the audit.
A practical guide to keeping search access while controlling training. Delivered through the activation page after purchase.
Working code that logs real crawler hits on your site, checks published IP ranges where possible, and includes a CLI and fix-file generator.
A plain-language handoff that explains what the evidence means, what it does not prove, and what the client should fix first.
A dated reference for the sourced crawler registry, its purposes, and the boundary between a live probe and a robots-only assessment.
Short, evidence-led outreach and intake scripts for turning a real blocked or unreadable page into a useful conversation.
Parse your own server logs to compare observed crawler hits with CrawlProof's supplied-request evidence.
Runnable examples for putting robots and response checks into a deployment workflow without adding model calls.
Small repeatable tests for allow, disallow, wildcard, and conflicting-rule cases.
A checklist for distinguishing robots policy from CDN, WAF, challenge, and response-header behavior.
A simple way to review what changed between two robots files before a crawler-access regression reaches production.
CSV fields and issue templates that carry the exact URL, evidence, owner and acceptance test into an agency workflow.
A local command that compares status, robots, findings and answer-unit changes without assigning a visibility score.
Original templates for rewriting weak passages and assigning every important claim a dated owner and source.
Reverse DNS, forward confirmation, ASN and published-range fields for evidence stronger than a User-Agent string.
A local SHA-256 manifest for saved reports, registry versions, assumptions and answer-unit calibration.
No per-scan credits or CrawlProof-funded model meter. Published page, site, monitor, safety and platform limits apply. No tier remapping or current feature moved behind a later paywall; third-party crawler behavior and future external services remain outside our control.
I built CrawlProof around the part of GEO that can be honest as a lifetime product: fetching public pages, comparing crawler purpose, and reporting deterministic evidence. It measures access and extraction conditions — not whether an answer engine retrieves, cites, recommends, or selects your site.
Solo is $69 once, with no per-scan credits or recurring model meter. Published caps apply.
CrawlProof measures access evidence, extraction evidence and answer-unit coverage. Provenance and retrieval-surface experiments remain API evidence, not buyer-facing scores. It does not measure engine retrieval, citations, recommendations, source selection or verified vendor-network identity.
No inflated promise, no fake customer wall, no mystery score.
Yes. Several free tools check more AI visitors than we do. We show what the result means, point to the exact rule or page tag, give you the corrected file, and let paid users compare and scan a client portfolio. You keep the report instead of getting an unexplained score.
No. CrawlProof cannot promise citations. Schema's causal effect is unsettled, and answer engines change. We show whether a supplied request was allowed, blocked, challenged, or served an almost-empty page. That is evidence for remediation, not a citation guarantee.
We can report whether it exists, but adoption by major providers is unconfirmed. It is not the headline feature and we do not claim it is a proven ranking lever.
No. We test a request from CrawlProof, not a verified crawler from the vendor's network. Treat it as strong proof of what our request received, then use the report to find the rule or server setting responsible.
No per-scan credits or CrawlProof-funded model meter, no tier remapping, and no current feature moved behind a later paywall. Published capacity and safety limits apply. It covers the current deterministic CrawlProof product for its lifetime; it cannot control third-party crawler behavior or promise every future feature forever.
60 days, stated plainly. If the report is not useful for your site, request a refund within 60 days through the merchant-of-record checkout.
Run the free scan first. If the evidence is useful, choose Solo or Agency access for repeat audits.
The founder allocation is still live: 200 licences remain. Solo is $69 once, with no metering.
Keep the one-time access