Crawl Foundry
Site Audit

Sites, verification, and crawl setup

Sites and setup establish which property Crawl Foundry may crawl, how control is proven, and which evidence a future run is allowed to collect.

Bounded crawl

Keep property, authorization, run, and crawl coverage visible.

Traceable evidence

Move from summaries to rules, URLs, links, assets, and raw facts.

Comparable verification

Confirm fixes with another crawl using an equivalent scope.

Create one property per real crawl boundary

A Site Audit site belongs to one workspace and organization. It stores the normalized base URL, host, authorization state, crawl settings, schedules, alert rules, and run history. Reuse the same active host entry when you edit its configuration; use a separate property when host ownership or operating scope is materially different.

  • Current plan defaults allow 3, 5, 10, or 50 active Site Audit sites per organization.
  • The Sites view combines each property with its latest retained run; those rows can have different crawl dates.
  • A blocked property cannot crawl, and an unverified property remains a restricted probe target.

Select the site and run before reading evidence

The site controls the property. The selected run controls the saved evidence in run-scoped views. Starting another crawl creates another snapshot and never rewrites the earlier one.

A site-level setting can change after a run. The run keeps its own crawl policy, requested cap, estimate, counters, and timestamps so later investigation can reconstruct what actually happened.

Use the five-step setup flow

1
Choose My own website or Competitor or public URL.
2
Enter a public homepage or domain and confirm the normalized host.
3
For an owned site, create and complete an ownership challenge.
4
Set URL cap, depth, include and exclude patterns, rendering, and robots policy.
5
Review the effective scope and reservation estimate before saving or starting.

Complete one current ownership challenge

Crawl Foundry offers a CMS meta tag, an HTML file at the site root, or a DNS TXT record. The generated random token is tied to the selected property and method and normally expires after seven days. Creating instructions does not verify the domain; Check proof triggers the server-side verification path.

A failed check records its time, error code, and a bounded message. Confirm protocol, host, deployed HTML, DNS propagation, challenge method, and expiry before generating another explanation.

Understand the public Quick check

A public probe needs no ownership proof, but the server caps it at 100 URLs and depth 2, respects robots.txt, disables rendering, limits concurrency, and applies slower pacing. It has its own monthly organization quota.

Use it for a small public observation. It is not a substitute for the reach, rendering, monitoring, or client-delivery evidence of a verified full crawl.

Define the full-crawl boundary

InputWhat it changes
Base URLSeed, host identity, and the starting point for discovery.
Maximum URLsThe admission ceiling for this run, also bounded by the plan and the hard 5,000-page ceiling.
Maximum depthHow many discovery levels can be admitted, with a current server ceiling of 25.
Include patternsPaths eligible for admission after normalization.
Exclude patternsPaths removed from the eligible crawl set.
RenderingDisabled, eligible-only, or required browser rendering for admitted pages.
robots.txtWhether the saved full-crawl policy respects published crawl directives.

Read effective limits, not only form values

The current plan defaults cap a crawl at 500, 1,000, 2,500, or 5,000 pages. The server also applies organization, site, concurrency, cooldown, pattern, and worker-safety limits. A saved value above a newly effective limit is normalized when the next run or schedule is admitted.

Discovered URLs beyond the admitted cap remain a coverage fact. They are not failed pages, and they should not be silently added to the crawled denominator.

Choose rendering with evidence and cost in mind

Disabled uses fetched HTML only. Eligible rendering reserves an estimate based on 25 percent of the URL cap; required rendering estimates every admitted URL. The run records actual rendered-page units and settles against actual auditable work, releasing unused reservation value where applicable.

Rendering can change metadata, links, resources, structured data, and findings. Keep the policy equal when comparing runs, and do not claim browser-rendered evidence for a run whose policy disabled it.

Separate estimate, reservation, collection, and settlement

The review step shows the current server estimate and spendable organization balance. Admission reserves the estimated charge before dispatch. Completion uses audited and rendered page units to calculate the final provider cost and customer charge; zero-work or eligible failure paths can release the reservation.

A saved request or dispatch record is not proof that pages were collected. Use the run ledger and terminal billing state for that claim.

Start one bounded run deliberately

  • Only one full run, including report publication, can be active for a site.
  • The organization also has an active full-run capacity and a per-site request cooldown.
  • Current default monthly quotas are 15 full crawls and 10 public probes per organization, subject to effective plan or control-plane policy.
  • Completed full crawls count. Failed, cancelled, or partial full crawls count only after at least ten completed pages; balance-skipped runs do not count.

Confirm the worker and the published report

After admission, watch dispatch, worker heartbeat, page counters, resource and external-link work, materialization, publication phase, terminal reason, and billing. A terminal primary crawl can still be publishing findings and aggregates.

Plan around the retention tiers

The newest ten terminal runs and runs no older than 30 days keep detail. Older evidence is announced before compaction. Between roughly six and twelve months, only one completed or partial monthly anchor remains; most runs past twelve months are deleted, except the newest terminal result for the site.

A finished detailed run can be one of up to five pins per site. Pins need a reason, protect intermediate detail, and do not protect an otherwise deletable run past twelve months.

Archive only after checking active work

A property cannot be archived while a full run is active. Archiving also archives its schedules and alert rules. After the current seven-day recovery grace, archived-site retention can delete crawl runs regardless of their normal age tier, while billing records remain governed separately.