Crawlability Scan
Your robots.txt says one thing. Find out what the crawlers got
From $99 USD · priced by URL count · import and review free
Reads your sitemap, then goes and looks. Every URL you approve is checked against the AI crawlers’ declared access, and a sample of your pages is requested as those crawlers to see what they were actually served. Declared access and observed serving are different questions, and only one of them is answerable by reading your own configuration.
Two things have to be true before an AI system can use your page: you have to permit it, and your stack has to serve it. The first is a declaration in a file you own. The second happens in your CDN, your WAF and your edge rules, which never read that file. This scan measures both and reports them side by side.
What the scan returns
Six artifacts, delivered together in your dashboard. Every one of them is a request we made and a rule we applied — an observation you can check, not a score you have to trust.
One row for every URL you approved
The grid is the deliverable, and it is complete by construction: a run that hits a bot wall or runs out of clock still returns a row for every remaining URL, saying which. A report shorter than the list you approved would be a silent truncation of the thing you paid for.
Always one row per approved URL
Declared crawler access, per URL path
Your robots.txt read the way the standard says a crawler must read it — per path, longest match, wildcards honoured — for every named AI crawler that publishes a robots token. Four answers per crawler per URL: allowed, blocked, not mentioned, or unknown. "Not mentioned" and "unknown" are kept apart, because a file we could not read is not a site that said nothing.
30 crawlers × every approved URL
Declared versus observed
The part a declaration cannot answer. A sample of your pages is requested as the crawlers themselves — seven AI crawler identities plus a browser control — and each identity's answer is recorded next to the browser's: the status it got, the bytes it received, how similar the content was, what format came back, and whether it was challenged. Pages whose answers all matched store the match rather than a table of identical rows.
Per-identity detail, sampled by page template
Index directives, page by page
Whether each page we fetched carries a noindex or nofollow, from your markup and from your response headers, and which of the two it came from. The two are read the way a search engine reads them — the more restrictive wins, so a header noindex is not cancelled by a meta index. What it reports is a directive’s presence, never a verdict on whether the page is indexed.
Meta tags and X-Robots-Tag headers
What your sitemap left out
Entries you list that are dead. Same-site pages your approved pages link to that the sitemap never lists — fetched out of your band’s unused headroom, because you already paid for that capacity and a page you missed is usually the point. Which of your URLs your own llms.txt actually names. And whether your pages have markdown twins, by content negotiation, by .md address, or because your llms.txt says so.
Discovery runs inside the band you bought
Coverage, stated as a number
Every URL the scan probed is published, and every sample says how much of the site it covered — how many page templates we found, how many we sampled, which pages the markdown check reached. A URL we could not measure is reported as unmeasured, with our reason for not measuring it, and never as a finding about your site.
Both directions — a positive result is sampled too
Shape only, not a real run. The delivered grid carries fifteen fixed page columns — title, canonical, headings, link counts, security headers, render dependence and the rest — plus one column for every crawler in the roster, and it scrolls sideways; this shows four crawler columns and one page column so the pattern is legible. The last row is a page the scan could not measure, which is reported as unmeasured rather than as a finding.
You see your URL inventory before any money is involved
Payment gates the crawl, not the inventory. Importing your sitemap and reviewing what came back is free and unmetered — you find out what the instrument sees before deciding whether the crawl is worth buying.
Your URL inventory, before any money
Point it at your sitemap and it reads the whole thing — sitemap indexes included, one level deep, as the protocol allows. You get the list of URLs it found, the documents it read them from, and every cap it hit spelled out, so a cut inventory can never read as a complete one.
Three findings you can act on for free
Entries whose last-modified date is over a year old, entries that carry no last-modified date at all, and duplicates — the same page listed twice under two spellings. All three come out of parsing your sitemap, so none of them needs a crawl.
Nothing is fetched from your site
The free stage reads sitemap documents and your robots.txt. It never requests a page your sitemap lists. That is a pinned contract, not a courtesy: the crawl is the thing being sold, so it cannot happen before the sale.
You choose what gets scanned
Deselect anything you do not want crawled and the price follows the count you approve — the band is resolved from your approved list, so deselecting can move you into a cheaper one. You can also list pages to skip if the scan finds them by following your links.
The two findings a crawl is needed for — what the crawlers were served, and the pages your sitemap never lists — are named on the review screen and withheld until you buy. We would rather tell you what you are not getting than quietly leave it out.
How the scan works
Run the free Diagnostic
The scan reads the domain it measures from a completed Diagnostic rather than a checkout field, so the free run comes first. Five to ten minutes, no credit card, and you keep the ten-pillar report either way.
Import your sitemap
Free, and nothing on your site is fetched. You get your URL inventory, the stale and missing last-modified dates, the duplicates, and every cap the import hit.
Approve what gets crawled
Deselect anything you want left alone. The band is resolved from the count you approve and the price follows it, so you see exactly what the scan will cost before you pay.
Read the grid
The scan runs unattended, paced politely, and stops if your host asks it to. The receipt tells you the longest it can take for the number of pages you approved, and we email you again when the report is delivered — so you can close the page.
The scan identifies itself honestly while it works — one self-describing user agent pointing at our published methodology, unauthenticated GETs only, a pacing floor between requests, and a hard stop with no retry if your host signals that it wants us gone. Read the methodology →
What the scan will not tell you
The boundary is the product. A measurement that quietly widens its own claims is worth less than one that states where it stops.
It is one observation, not a trend line
A scan records what your site served on the day it ran. Your past scans sit on the same shelf so you can re-read them, but there is no scan-over-scan comparison, no drift chart and no alerting. Run it again after you change something and read the two.
It does not look at every page equally
Every URL you approve gets fetched and gets a row. The two deeper checks sample instead, by page template and largest templates first: the markdown-twin check takes one page per template, and the declared-versus-observed walk takes two or three, escalating where a template’s answers start to differ. Both publish exactly which pages they reached, and unchecked is never reported as absent.
It tells you what differed, not why
When a crawler receives something a browser does not, the report names the difference — the status, the size, the format, how much of the content matched. It does not diagnose the cause. Your CDN rules, your WAF configuration and your framework are yours to read; a guess at which one did it would be a finding we cannot support.
One site means your apex and www
Other subdomains are refused rather than guessed at, which means a site that legitimately splits content across them is under-reported. Widening that is registered work, not a silent behaviour we might already have.
There is no model anywhere in it
No LLM reads your pages and nothing is scored by one. Every finding is a request we made and a rule we applied, which is why the scan costs us nothing in vendor spend and why running it twice on an unchanged site returns the same answer. A judgement about how well your content answers a question is a different product and is not in this one.
Pages to skip are exact, and per scan
The skip list matches a URL exactly. No wildcards and no prefixes — typing /admin does not cover /admin/users, and the review screen says so rather than letting you find out from the invoice. The list belongs to that order; it is not remembered for your domain.
Where this sits next to a free checker
Plenty of tools will read your robots.txt, and some of them are free. Here is the split, stated the way we would want it stated to us.
What a robots.txt checker tells you
What your file says, and whether it parses. That is a real job and worth doing — a syntax error in robots.txt can cost you a whole crawler. Several tools do it well and some of them do it free, including at scale across a thousand URLs.
What only a request can tell you
Whether the crawler actually received your page. A declaration lives in a file you control; serving happens in your CDN, your WAF, your edge rules and your framework, none of which read that file. This scan asks for your pages as the crawlers and records what came back, which is the one question you cannot answer by reading your own configuration.
What the scan sits inside
Crawler access is one reading of one pillar — Discoverability. The free Diagnostic scores ten, and every one of them carries the perception gap between what you rated yourself and what we independently detected. The scan lands on the same shelf as that report, for the same domain, so an access finding reads against the infrastructure underneath it.
We score the tools in this space — and our own row — on one published rubric, from public surfaces, with retrieval dates on the evidence. See how the Discoverability tools score on The Readiness Bench →
Priced by the number of URLs you approve
One payment, one scan, no subscription. The band comes from the count you approve on the free review screen, so the price is settled before you pay — and deselecting pages can move you down a band.
A site with more than 1,000 URLs to scan is a conversation rather than a checkout — the largest band is a ceiling, not a starting point, and we would rather scope it with you than round you up into it. Get in touch →
No plan includes a Crawlability Scan — every scan is a one-off order, and a Free-plan account with a completed Diagnostic can buy one. Where it fits in the rest of the ladder: pricing and the Audit stream.
Start with your URL inventory
Tell us your site and we take it from there. The scan reads the domain it measures from a completed Diagnostic, so the free run comes first if you have not done one — five to ten minutes, and you keep the ten-pillar report whether or not you buy a scan afterwards. The sitemap import and the review that follows are free too; you only pay when you approve the crawl.