Crawl Foundry
Site Audit

Issues, Page Explorer, and Internal Links

A technical finding becomes useful when the rule, affected pages, exactness, and relationships that produced it remain attached to the same crawl run.

Bounded crawl

Keep property, authorization, run, and crawl coverage visible.

Traceable evidence

Move from summaries to rules, URLs, links, assets, and raw facts.

Comparable verification

Confirm fixes with another crawl using an equivalent scope.

Read the issue catalog before the count

The Issues workbench groups findings by code, label, category, source, severity, affected unit, current count, previous comparable count, movement, and exactness. Its catalog can be complete, partial, or failed, and the run can still be unpublished.

Critical and high do not remove the need to inspect business intent. Low and informational evidence can still expose a broad template error.

Use movement only when stored runs are comparable

The workbench can label an issue new, regressed, fixed, improved, persistent, baseline, or indeterminate. A movement needs a previous comparable run and stored aggregates. Missing history, partial catalog coverage, lower-bound counts, or an unpublished run must remain visible instead of becoming a made-up zero.

Open the affected URLs or assets

An issue group can represent URLs, resources, or another bounded unit. Open several affected rows, attach the rule and evidence to any handoff, and distinguish unique affected units from repeated finding occurrences.

Agent copy is a bounded evidence packet from loaded records. It does not create a diagnosis, approve a fix, or prove that no additional affected rows exist.

Build a reproducible Page Explorer set

Page Explorer defaults to internal URLs and reads 50 rows per client page. Search, task status, discovery source, indexability, response class, content kind, issue presence, origin, structured data, canonical, hreflang, reachability, template, segment, path, depth, click depth, word count, and link-count ranges can narrow the server query.

The URL column always stays visible. Other columns and sort order belong to the saved view, not to the crawl snapshot itself.

Save the question, not a copy of the data

A saved Page Explorer view stores filters, selected columns, and sort for reuse. The current default limit is 50 saved explorer views per workspace. Reopening one applies its query to the selected run, so the row set can change when the run changes.

Use structure as a path aggregate

Structure rolls retained pages into path and directory groups. Counts and issue pressure describe the selected run and its admitted scope. They are useful for finding template clusters, but they are not a complete information-architecture map or a statement about how users navigate the site.

Separate relationships from occurrences

Internal Links keeps source, target, anchor, placement, status, nofollow, redirect, canonical, noindex, and broken-target context. Several occurrences with the same source, target, and placement can be compacted into one row, while relationship totals count unique source-target pairs.

Use exactness and pagination before comparing the loaded graph with the sitewide relationship total.

Read both endpoints and the placement

  • A broken target needs the source page, final response, redirect, and intended destination.
  • A canonicalized or noindex target can be intentional, but linking policy may still need review.
  • Body, header, navigation, sidebar, footer, and unknown placements have different editorial meaning.
  • Orphan and not-linked labels depend on the admitted crawl, sitemap evidence, and retained link graph.

Treat click depth and internal authority as projections

Click depth can be exact reachable, exact unreachable, estimated, or missing. Internal authority is calculated from the retained internal graph and can be excluded for ineligible pages. Neither value is a Google metric or proof of ranking impact.

A link opportunity still needs editorial relevance

Link-opportunity analysis can use manual or Keyword Database hints, page roles, graph gaps, topic tokens, and bounded candidate scoring. Its output remains a recommendation over the selected run. Confirm source context, anchor wording, destination intent, and whether the page should be indexable before editing.

Use totals, pages, and exports for different jobs

Issues, page rows, link rows, graph nodes, affected evidence, and URL-detail collections use different page sizes and read budgets. An export can cover a larger bounded set, but a queued export is not a complete file until its worker reaches a terminal state and the retained artifact is available.

Move from finding to verified change

1
Confirm site, run, publication state, rule, severity, affected unit, count exactness, and movement basis.
2
Inspect representative URL dossiers and any source-target link evidence.
3
Measure the template, directory, or page-set scope with a saved filter.
4
Name the owning code, template, CMS rule, or editorial process.
5
Apply the narrowest safe change and preserve intentional exceptions.
6
Use a comparable published crawl to check resolved instances and regressions.