Crawl Foundry
Site Audit

Crawl settings and domain verification

Settings define the evidence contract for future crawls. Verification proves the allowed relationship to the site; neither should be hidden when runs are compared.

Bounded crawl

Keep property, authorization, run, and crawl coverage visible.

Traceable evidence

Move from summaries to rules, URLs, links, assets, and raw facts.

Comparable verification

Confirm fixes with another crawl using an equivalent scope.

Settings define future crawl evidence

The current Settings view edits maximum URLs, maximum depth, include paths, exclude paths, rendering policy, and robots.txt behavior for the selected site. The saved base URL remains the property identity. Older runs keep their original policy and evidence.

Set a cap and depth that answer the question

The URL cap is bounded by the current plan and the hard 5,000-page ceiling. Depth is bounded by current policy with a server ceiling of 25. A higher number can improve reach, but it also changes cost, duration, admitted scope, issue totals, and comparison quality.

Test include and exclude patterns before comparison

Include rules narrow eligible paths and exclude rules remove paths from admission. They can create a clean investigation or accidentally hide a template. Record the reason for a rule and inspect discovered-but-not-admitted scope after the next crawl.

Choose fetched or browser-rendered evidence

Disabled keeps fetched HTML. Eligible renders pages selected by the rendering policy. Required requests browser rendering for every admitted page. Rendering affects the estimate and can change retained metadata, links, resources, markup, and findings.

Treat robots policy as part of the data contract

Verified full crawls can save the selected robots.txt behavior. Public probes always respect robots.txt regardless of the form. Changing the full-crawl setting affects reach and must be disclosed when comparing issue or coverage totals.

Use the proof method you can maintain

MethodUse it when
CMS meta tagYou can edit the homepage head or an SEO plugin.
HTML fileYou can deploy a named file at the site root.
DNS TXTYou manage the domain's DNS and want host-level proof.

Challenge creation and verification are different actions

Creating instructions stores a method and random token. The token normally expires after seven days. Check proof sends the method, token, property, workspace, and organization through the protected server verification route; the browser alone cannot mark a property verified.

Unverified checks remain restricted

An unverified property can run only the public probe policy: at most 100 URLs, depth 2, robots respected, rendering disabled, low host concurrency, and conservative pacing. Schedules require a trusted authorization at dispatch and are skipped while the site remains unverified.

Review the new estimate before the next run

Changing cap or rendering changes the estimate. Saving settings does not reserve balance and does not start a crawl. Reservation happens only when a manual or scheduled run is admitted.

Management access and crawl authorization are separate

Viewing Site Audit, managing sites, starting runs, monitoring, and exporting are capability checks. Domain verification controls crawler trust. An authorized user cannot bypass a blocked site, effective plan limit, organization balance, quota, or active-run capacity.

Document every scope change

1
Record the previous cap, depth, patterns, robots policy, rendering policy, and authorization state.
2
Explain which investigation requires the change.
3
Save the settings and review the updated estimate.
4
Start a new run and confirm admitted scope plus publication.
5
Do not label the first score or issue delta as improvement unless the evidence is comparable.