Crawl Foundry
Site Audit

Site Audit

Site Audit turns a bounded crawl into connected evidence about health, issues, URLs, internal links, structure, history, automation, and page-level technical facts.

Bounded crawl

Keep property, authorization, run, and crawl coverage visible.

Traceable evidence

Move from summaries to rules, URLs, links, assets, and raw facts.

Comparable verification

Confirm fixes with another crawl using an equivalent scope.

Use Site Audit as an evidence loop

1
Select the exact site and decide whether you need a verified full crawl or a restricted public Quick check.
2
Review the URL cap, depth, path rules, robots policy, rendering choice, estimate, and available balance.
3
Start the run, then distinguish admission, worker collection, secondary processing, and report publication.
4
Use Overview and Statistics to choose a question, then inspect Issues, Pages, Internal Links, Structured Data, Performance, or one URL dossier.
5
Record the intended change and the issue instances or page set it should affect.
6
Run a comparable crawl and check the saved evidence before calling the work complete.

Choose the view that owns the question

ViewQuestion it can answer
SitesWhich saved property, authorization state, and latest run are in scope?
OverviewWhere does the selected run deserve investigation?
IssuesWhich recorded rules affect which URLs or resources?
PagesWhich crawled URLs match a reproducible technical filter?
StatisticsWhich measured distributions point to a template or section pattern?
Structured DataWhat markup, validation evidence, and feature profiles were recorded?
PerformanceWhich sampled or covered URLs have field, lab, or diagnostic evidence?
Internal LinksWhich source-target relationships, occurrences, and graph signals were retained?
CompareWhat changed between two published completed or partial runs?
CrawlsDid collection and report publication finish, and where did work fail?
AutomationWhich recurring run and alert policies are active, skipped, or paused?
SettingsWhat data contract will future crawls use?

Keep site, run, and evidence date together

Most views are run-scoped. The site selects the saved property; the run selects an immutable crawl snapshot. A new crawl does not refresh an old run, and opening an older run must not be reported as current site state.

Compare, Automation, Settings, and the Sites portfolio have their own scopes. Follow the site and run switchers instead of assuming every visible number shares one timestamp.

Authorization changes what the crawler may do

Verified and admin-approved sites can use the saved full-crawl policy. Unverified sites are limited to a public probe of at most 100 URLs and depth 2, with robots.txt respected, rendering disabled, conservative pacing, and a separate monthly quota. A blocked site cannot start a run.

Verification proves control of the saved host through a current challenge. It does not prove that every subdomain, login area, third-party resource, or environment belongs in the crawl.

Read collection quality before Site Health

A small issue total can mean a healthy site, a narrow crawl, blocked discovery, failed fetching, unfinished processing, or compacted historical detail. Check requested, discovered, admitted, completed, failed, skipped, and publication state before interpreting the result.

Site Health is published only when required categories, materialization, terminal state, and confidence meet the score policy. A missing score is an evidence state, not a score of zero.

Collection and report publication are separate stages

A run moves through request, dispatch, discovery, fetching, analysis, aggregation, and terminal states. After primary crawl work ends, link projections, findings, issue summaries, link metrics, aggregates, Site Health, AI-readiness evidence, and alerts can still be publishing.

A completed or partial label without secondary completion is shown as Processing. Wait for the published report before treating issue counts, health, comparisons, or alerts as final for that run.

Treat Site Health as a versioned summary

The current score combines seven categories: crawlability and availability, indexability and canonicalization, content and metadata, links and architecture, structured and international signals, performance and resources, and security and protocol. Category prevalence and severity shape the burden; verified blockers can cap the published result.

The score stores score, ruleset, eligibility, scope, coverage, confidence, exactness, and truncation context. Trend comparisons require matching versions and scope, measured confidence in both runs, exact materialization, no truncation, and no material coverage gap.

Move from summary to the owning evidence

  • Issues groups recorded rules and movement, then opens affected evidence.
  • Page Explorer answers bounded URL-set questions with server filters and pagination.
  • Internal Links separates relationships from repeated occurrences and marks count exactness.
  • The URL dossier keeps one URL and run across issues, links, assets, markup, changes, and technical facts.
  • Structured Data and Performance expose their own coverage and missing-evidence states.

Visible rows can be a bounded projection

Tables, graphs, issue samples, URL details, histories, and exports have explicit read or retention limits. Counts can be exact even when samples are paginated; other views label lower bounds, truncation, incomplete catalog state, or missing detail.

Do not turn a loaded page, top list, or graph viewport into a sitewide claim. Use the displayed total, exactness flag, next page, saved filter, or export that belongs to the question.

Use history and automation for repeatable observation

Compare works from two published completed or partial runs and can retain URL, metadata, finding, link, resource, segment, and tree changes. Mapping rules are explicit for crawl-to-crawl, migration, or authorized staging-to-production work.

Automation can schedule daily, weekly, or monthly runs and evaluate active alert rules against a previous comparable run. A due schedule, accepted run, generated alert event, queued email, and delivered email are different facts.

Know when detailed evidence expires

The newest ten terminal runs and every run no older than 30 days keep full detail. Older runs can be announced for compaction; from roughly six to twelve months only one completed or partial monthly anchor survives in compact form. Most runs older than twelve months are deleted, while the newest terminal result for a dormant site is protected.

Up to five finished runs per site can be pinned with a reason against intermediate compaction, but a pin does not override the twelve-month rule. Active comparisons, active exports, and unsettled billing also defer retention work. Billing evidence remains separate from crawl-detail retention.

Close the work with comparable evidence

A code change, ticket status, or higher score is not enough. Reuse an equivalent property, authorization, scope, rendering policy, and completion standard; check the exact issue or URL evidence; and look for regressions in related indexability, linking, rendering, performance, and structured-data signals.