AI crawler defaults vary by vendorcheck vendor docsRun a free scanGet lifetime access
A clearer way to audit AI access

See what AI visitors actually get from your site.

See what AI gets when it visits your site — and the exact reason it cannot read something.

We check your page, what search visitors can read, and whether the words and answers are easy to use.

CrawlProof does not guess what an AI will say or promise a citation. It reads your public page, compares the responses, and shows the proof in plain language.

FREE CHECK See what AI visitors can read, what your page says, and what may be hard to use. Run the free scan →

Scan one URL now

No signup. No model tokens. Paste a public page and see what AI is allowed to read — and whether it actually receives useful content.

01GET /robots.txt200
02check your access rulesREADY
03ask the AI visitorsWAIT
04compare what they receiveWAIT
05find hidden page blocksQUEUE
output appears here after submit / no model call involved
Ready. Your URL stays in your browser until you submit.

Privacy note: CrawlProof records aggregate funnel stage counts only (for example, scan started or activation succeeded). No cookies, analytics profile, license key, URL or client identity is sent for this measurement, and the free scanner does not require tracking consent.

n/apage response
n/acrawler checks
n/aH1 headings
n/apage needs scripts

This is the result, not a design concept.

The report below is a finished scan of a real public page. It shows what the site tells AI visitors, what they actually receive, and the proof behind each warning.

source: www.notion.so / scan captured 09 Aug 2026

READ 01what AI is allowed to readREAD 02what AI actually receivedREAD 03where the site's instructions disagreeREAD 04the exact page and rule checked
FINISHED EXAMPLE REPORTNOT A LIVE SCAN
200page response
AI checks
H1 headings
crawler content
Source: www.notion.so · captured 09 Aug 2026
Stored production scan. This example does not run again when the page loads.

The common collateral damage is simpler.

In our measured sample, sites that blocked AI often blocked search and training together. CrawlProof shows the cost plainly: you blocked AI training and also removed yourself from AI search — then gives you the configuration that keeps search access while staying out of training.

OPENAI / entirely blocked incl OAI-SearchBot: 8/75
ANTHROPIC / entirely blocked: 13/75
TRAINING ONLY / OpenAI: 10/75 · Anthropic: 9/75
75 sites · root-path robots sample · CrawlProof measurement · not a web-wide prevalence estimate.

One limits table. No hidden units.

These are the current product limits. A page audit, sitemap sample, monitor and client-site batch are different capacities.

CapacityFreeSoloAgency

Free scans remain one-page checks. Paid single-page audit capacity is separate from sitemap sampling.

Make your page easier to use.

Plain answers about what your site says, what AI visitors receive, and what a buyer can understand quickly.

01

Check what is allowed

We show which AI visitors your site welcomes, which it turns away, and where two instructions disagree.

02

Check what arrived

We compare the page a normal visitor sees with the page an AI-style request receives.

03

Read the page like a buyer

We look for clear answers, useful headings, real text, structured data, and sections that make sense on their own.

04

Give you the proof

Each result points to the page, response, or instruction behind it. No guessed score and no promise of a ranking.

HONEST LIMITATION: we test a request from CrawlProof, not a verified crawler from a vendor's network. A refusal proves what our request received, not what every real AI visitor receives.

What you actually get.

Small checks, assembled into a report an owner or agency can explain to a client.

01

ChatGPT is being told your page is empty

We compare the page a browser sees with the page each crawler receives. A 200 response can still be an empty app shell.

  • content size and headings
  • title and structured data
browser / crawler
same URL
real difference
02

Your robots file and your server give opposite orders

We join robots rules, noindex headers, page tags, canonicals, sitemaps, and redirects instead of listing them separately.

robots / headers
page tags / canonical
sitemap / redirect
03

Find the blocks robots.txt cannot show

We also check page tags and server headers that can say “do not list this page” even when the main rules look fine.

server header
page tag
hidden noindex
04

Catch the empty app shell

A page can look fine in a browser but send almost nothing to a text-only visitor. We show that difference instead of calling it healthy.

browser / AI visitor
page size
real difference
05

Questions people actually ask

We flag headings that sound like slogans and show page-derived prompt ideas. They come from your words; they are not predictions of what AI will say.

headings / questions
page-derived prompts
not predictions
06

Sections that make sense alone

We look for short passages that start with “it”, “this”, or “they” and may lose their meaning when copied without the rest of the page.

stand-alone passages
clear subjects
extractable text
07

Structured data and trust signals

We check useful page types, names, dates, sources, numbers, definitions, and comparisons without claiming they guarantee visibility.

FAQ / Article / Product
dates / sources
consistent name
08

Find the sites that need help

Agency scans check a list of client or prospect sites, sort the real issues first, and give you a report you can hand to a client.

50-site list
severity first
client-ready report
09

Keep a dated record

Paid users can save a site, run it again no more than once a day, compare the result with the last run, and review alert records when access, findings, or structured data change. No email is sent.

saved history
daily maximum
alert records
10

Sourced crawler rules, bounded probes

We keep a sourced, dated registry and label its entries honestly: a fixed subset is live-probed, while the rest are rules-only checks. More registry rows do not create an unbounded request bill.

dated sources
live vs rules-only
40-request ceiling
“20 minutes per client per week manually Googling their brand in ChatGPT and Perplexity, screenshotting results, and pasting them into a Google Doc. That is not a workflow. That is desperation.” — Freelance SEO managing nine client sites, Indie Hackers user report; source

Same anxiety. Different bill.

These are the settled comparison figures from our research. Third-party captures remain labelled as such.

ToolPrice observedWhat it doesWhat CrawlProof does instead
Profound$99–$399/moLLM visibility and citation trackingOne-time technical evidence, no model inference
Peec AI~$95+/mo, third-party captureAI search monitoringAI visitor access and page-content checks
AthenaHQ~$295/mo, third-party captureAnswer-engine analyticsRobots and HTTP evidence without a subscription
Semrush AI Visibility Toolkit$99/mo per domainAI visibility and an AI-readiness site auditFocused deterministic audit, $69 once
Ahrefs Brand Radarfrom $199/moBrand visibility in AI answersNo citation tracking, no recurring model cost
ZeroRank AI$89 LTD observedAI-search citation monitoringTechnical half of GEO, not citation monitoring
CrawlProof Solo$69 onceAI access checks, content differences, hidden blocks, fix filesWhat they do that we don't: citation tracking.

Free with purchase. Built, not promised.

Paid access is capacity and deliverables, not a metered scan. The ten included materials are practical files you can use after the audit.

BONUS 01 / PRIVATE GUIDE

AI Crawler Access Playbook

A practical guide to keeping search access while controlling training. Delivered through the activation page after purchase.

BONUS 02 / WORKING CODE

Crawler Agent Kit

Working code that logs real crawler hits on your site, checks published IP ranges where possible, and includes a CLI and fix-file generator.

BONUS 03 / CLIENT HANDOFF

Client-ready explanation

A plain-language handoff that explains what the evidence means, what it does not prove, and what the client should fix first.

BONUS 04 / VENDOR REFERENCE

Crawler reference sheet

A dated reference for the sourced crawler registry, its purposes, and the boundary between a live probe and a robots-only assessment.

BONUS 05 / PROSPECTING

Agency prospecting scripts

Short, evidence-led outreach and intake scripts for turning a real blocked or unreadable page into a useful conversation.

BONUS 06 / LOGS

Crawler log parser

Parse your own server logs and separate what the records can confirm from User-Agent-only claims; compare that evidence with CrawlProof's supplied-request results.

BONUS 07 / CI

Continuous checks kit

Runnable examples for putting robots and response checks into a deployment workflow without adding model calls.

BONUS 08 / TESTS

Robots test pack

Small repeatable tests for allow, disallow, wildcard, and conflicting-rule cases.

BONUS 09 / CDN AND WAF

Protection checklist

A checklist for distinguishing robots policy from CDN, WAF, challenge, and response-header behavior.

BONUS 10 / DIFF VIEWER

Robots change viewer

A simple way to review what changed between two robots files before a crawler-access regression reaches production.

BONUS 11 / DELIVERY

Evidence-to-ticket pack

CSV fields and issue templates that carry the exact URL, evidence, owner and acceptance test into an agency workflow.

BONUS 12 / REGRESSION

Before/after regression harness

A local command that compares status, robots, findings and answer-unit changes without assigning a visibility score.

BONUS 13 / EDITORIAL

Answer-unit worksheet and claim ledger

Original templates for rewriting weak passages and assigning every important claim a dated owner and source.

BONUS 14 / IDENTITY

Crawler-log identity workbook

Reverse DNS, forward confirmation, ASN and published-range fields for evidence stronger than a User-Agent string.

BONUS 15 / INTEGRITY

Report integrity manifest

A local SHA-256 manifest for saved reports, registry versions, assumptions and answer-unit calibration.

Lifetime means no quiet downgrade.

No per-scan credits or CrawlProof-funded model meter. Published page, site, monitor, safety and platform limits apply. No tier remapping or current feature moved behind a later paywall; third-party crawler behavior and future external services remain outside our control.

A note from the founder

I built CrawlProof around the part of GEO that can be honest as a lifetime product: fetching public pages, comparing crawler purpose, and reporting deterministic evidence. It measures access and extraction conditions — not whether an answer engine retrieves, cites, recommends, or selects your site.

FIELD NOTEOne-time access keeps the evidence legible: no recurring visibility bill, no mystery score.REFUND TERMS
See the current merchant-of-record terms at checkout.

Start with the evidence.

Re-check a fix now instead of waiting 30 days

What you're getting

One complete page auditfree
Saved history and daily monitoringpaid tier
Customer-log crawler evidence with bounded live DNS checkspaid capacity
Solo capacity$69 once
Agency capacity$149 once
Per-scan credits$69 once
$69
Solo, one-time payment

Solo is $69 once, with no per-scan credits or recurring model meter. Published entitlement and temporary platform safety limits apply.

CrawlProof measures supplied-request access evidence, extraction evidence, customer-owned log records and deterministic schema checks. When a log includes client IPs, it can perform bounded reverse-DNS and forward-confirmation checks against a sourced, dated vendor registry: up to 12 unique IPs per request, with continuation when more remain. User-Agent strings remain claims; stale or unavailable evidence is reported rather than treated as verified. It does not measure engine retrieval, citations, recommendations, source selection or AI visibility.

  • Solo limits
  • Monitoring limits
  • Log limits
  • Agency limits
  • alert records are stored; no email is sent
  • 15 materials delivered in two gated downloads on activation

Paid tiers include an Access Evidence Pack with finding IDs stable only while the finding key, scope and evidence location remain unchanged; passive marker detection for AIPREF/RSL/Web Bot Auth without adoption or signature validation; bounded Common Crawl archive-presence evidence; and deterministic monitoring diffs. Each archive capture means only that Common Crawl captured the URL on the recorded date—it does not prove model training, retrieval or citation.

Questions worth asking before you buy.

No inflated promise, no fake customer wall, no mystery score.

Isn't this free elsewhere?

Yes. Several free tools check more AI visitors than we do. We show what the result means, point to the exact rule or page tag, give you the corrected file, and let paid users compare and scan a client portfolio. You keep the report instead of getting an unexplained score.

Will this get me cited by ChatGPT?

No. CrawlProof cannot promise citations. Schema's causal effect is unsettled, and answer engines change. We show whether a supplied request was allowed, blocked, challenged, or served an almost-empty page. That is evidence for remediation, not a citation guarantee.

Does llms.txt make me visible?

We can report whether it exists, but adoption by major providers is unconfirmed. It is not the headline feature and we do not claim it is a proven ranking lever.

Does a refusal prove every real AI visitor is blocked?

No. We test a request from CrawlProof, not a verified crawler from the vendor's network. Treat it as strong proof of what our request received, then use the report to find the rule or server setting responsible.

What does lifetime mean?

No per-scan credits or CrawlProof-funded model meter. Published entitlement and safety limits apply to the current deterministic CrawlProof product; third-party crawler behavior and future external services remain outside our control.

How does the refund work?

60 days, stated plainly. If the report is not useful for your site, request a refund within 60 days through the merchant-of-record checkout.

Refund terms at checkoutMerchant of Record: Dodo PaymentsOne-time checkout, not a subscriptionNo CrawlProof-funded model call in the scan

Find the block before the buyer finds it.

Run the free scan first. If the evidence is useful, choose Solo or Agency access for repeat audits.

Run a free scan Get lifetime access
PUBLIC PAGES ONLYNO MODEL CALLEVIDENCE, NOT A VISIBILITY PROMISE