Crawl Foundry
Site Audit

Crawl automation and alerts

Automation creates a recurring opportunity to collect evidence, not a promise that every due run will start. Alerts are useful only when their baseline, threshold, delivery, and owner remain visible.

Bounded crawl

Keep property, authorization, run, and crawl coverage visible.

Traceable evidence

Move from summaries to rules, URLs, links, assets, and raw facts.

Comparable verification

Confirm fixes with another crawl using an equivalent scope.

Schedule evidence someone will review

A schedule belongs to one saved site and can run daily, weekly, or monthly in a named time zone. It stores local hour and minute, weekly day or monthly day, URL cap, run policy, run cap, monthly budget, failure behavior, balance behavior, jitter, next due time, and window counters.

The current default limit is five configured schedules per site. More schedules do not create better monitoring when they measure the same scope without a distinct owner or decision.

Due is not the same as started

The dispatcher checks that the site still exists and is verified, no site or organization capacity conflict exists, the organization quota allows another full run, the schedule has not reached its run or budget cap, and the organization can reserve the estimate. It then creates a run and dispatch outbox record.

A duplicate occurrence reuses its existing run. A due time, accepted database row, dispatched event, claimed worker, completed crawl, and published report are separate proof points.

Read visible skip reasons

  • Site unavailable or unverified
  • Another site run active or organization capacity full
  • Organization monthly quota reached
  • Schedule monthly run cap reached
  • Schedule monthly budget cap reached
  • Insufficient organization balance

Run caps and budgets use the schedule time zone

Each schedule keeps a monthly window, number of runs started, and estimated cost reserved in that window. The effective run cap is the lower of the schedule value and the current server policy. The budget check uses the next estimate before reservation; final settlement still follows actual units on the run.

Pause, resume, edit, or delete deliberately

Pausing clears the next due time and can store a reason. Resuming calculates a new due time from the current moment. Editing an active schedule recalculates its next due time and current budget window. Deleting soft-deletes it and removes the due time; it does not erase historical runs.

Choose one of the recorded alert questions

Current alert kinds cover new broken links, new noindex, robots changes, canonical changes, title changes, meta-description changes, structured-data regression, sitemap regression, Site Health drop, and bot anomaly. A rule can add severity, JSON thresholds, a suppression window from 0 to 720 hours, a digest rhythm, and optional schedule scope.

Alerts need a comparable run

Active rules are evaluated after a completed or partial comparable run. The current policy uses a previous successful comparison baseline, optionally scoped to the same schedule. Raw finding and inventory evidence is read through a bounded scan, and the run has an event cap.

No previous run, incomplete coverage, or a suppressed duplicate is a different state from no change.

Send digests to people who own the response

A rule can target workspace users or explicit email addresses, with a current cap of ten recipients per site. Digest choices are per run, daily, or weekly. The rule creator is used as the default recipient when no explicit list is supplied.

Event and email delivery have separate histories

An alert event stores the matched rule, run, severity, summary, counts, and bounded samples. Delivery then has its own queued, claimed, retry, sent, delivered, bounced, complaint, opened, clicked, or error evidence depending on provider events. A generated event is not proof that every recipient received the email.

Use suppression and digests without hiding regressions

The default suppression window is 24 hours. Site-level delivery is also rate-bounded. Use these controls for repeated fingerprints and expected fluctuations, but keep run history available so suppressed notifications do not become missing technical evidence.

Respond from the run, not from the subject line

1
Open the exact site, alert rule, run, and previous comparison baseline.
2
Confirm terminal publication and comparable coverage.
3
Inspect the recorded sample, then open the full issue, page, link, or structured-data evidence.
4
Classify the result as action, observation, accepted exception, collection problem, or delivery problem.
5
Assign an owner and record the intended verification.
6
Check the next comparable published run before closing the change.