Crawl Foundry
Site Audit

Site Audit URL dossier

The URL dossier keeps page-level signals tied to one URL and crawl run, so an investigator can move from an orientation score to the evidence without losing context.

Bounded crawl

Keep property, authorization, run, and crawl coverage visible.

Traceable evidence

Move from summaries to rules, URLs, links, assets, and raw facts.

Comparable verification

Confirm fixes with another crawl using an equivalent scope.

Confirm URL, site, run, and retention state

The header identifies the saved inventory URL and crawl context. The stat strip shows the page score, issue count, HTTP status, response time, word count, inlinks, and outlinks where retained. These values summarize one URL snapshot, not the live page today.

A summary-only historical run can retain aggregate evidence while page detail has expired. Treat missing detail after compaction as unavailable, not clean.

Use the seven dossier tabs

TabEvidence
OverviewReadiness groups, severity counts, and recent change handoffs.
IssuesPage findings with severity, rule evidence, filters, and pagination.
LinksCrawl path, inbound and outbound relationships, anchors, and target states.
AssetsImages, resources, and hreflang references observed for the page.
SchemaRetained markup items, types, feature profiles, and validation issues.
ChangesSaved page-level differences against a previous comparable run.
TechnicalHTTP, metadata, robots, performance, content, and response facts.

Overview is a router, not a diagnosis

Overview groups signals such as crawlability, indexability, content, links, schema, security, and performance and links to the deeper tab. The displayed score is a page heuristic based on status, issue count, and URL depth; it is not the site-level Health score or a ranking probability.

Issues shows the page's loaded finding set

Findings can be filtered by severity and loaded in bounded pages, currently ten rows per incremental request by default. Compare total count with loaded rows before copying evidence. A resolved or ignored status belongs to the saved finding record and does not update the live page.

Links keeps relationships and totals apart

The tab shows crawl path, inbound and outbound rows, anchors, placements, follow state, response state, redirects, canonicalization, and indexability context. Summary totals can be exact while loaded relationship samples remain paginated. Inspect both endpoints before changing a navigation or contextual link.

Assets includes more than images

Images, scripts, stylesheets, fonts, documents, media, iframes, manifests, and other retained references can carry element, attribute, placement, alt, loading, render-blocking, response, and issue evidence. Hreflang references are shown alongside these page-bound resources but have their own semantic rules.

Schema shows bounded markup snapshots

The page joins retained JSON-LD, Microdata, and RDFa item snapshots with validation issues by item index. It can show types, properties, Google feature profiles, required or recommended gaps, parse errors, and truncation.

A URL can have at most 25 retained item snapshots, with per-item and total JSON budgets. The original page source remains the final reference when a stored value was shortened.

Changes needs an earlier comparable page

The tab compares retained URL facts across runs and groups bounded rows by change type. If there is no previous comparable run or the URL was not matched, no delta is inferred. A recorded technical change still does not prove a traffic or ranking effect.

Technical keeps raw facts and performance evidence together

Technical exposes response time, HTML size, content density, word count, HTTP and redirect state, indexability, canonical, robots, metadata, headings, cache and response headers, and available performance evidence. Missing CrUX field data, missing lab data, and a good field rating are distinct states.

Performance can mix field, lab, and crawl clocks

CrUX can be URL- or origin-scoped field evidence. PageSpeed or Lighthouse measurements are strategy-specific lab observations and can be sampled. Crawl response time is the crawler's own request observation. Compare source, scope, strategy, capture time, and coverage before combining them in one explanation.

Previous and next follow the stored investigation set

Page Explorer can store up to 100 inventory IDs for URL-page navigation in one browser-session context, with up to ten retained contexts. Previous and next move through that bounded set, not the entire crawl. Browser history preserves linked-page investigation separately.

Copied agent context is deliberately bounded

Agent prompts limit root causes, target URLs, evidence findings, array items, fields, lines, and characters. They help carry checked evidence into another workflow. They do not authorize automatic production changes or prove that omitted rows do not exist.

Investigate one URL without losing context

1
Confirm the URL, run time, publication state, and detail-retention state.
2
Use the summary and Overview to choose a question.
3
Open the owning evidence tab and load enough rows for the claim.
4
Follow related source or target URLs when the finding is relational.
5
Check live source, rendered output, server behavior, or Search Console when the crawl alone cannot answer the cause.
6
Verify the intended fix in the same URL dossier on a comparable later run.