Crawl automation and alerts
Bounded crawl
Traceable evidence
Comparable verification
Schedule evidence someone will review
A schedule belongs to one saved site and can run daily, weekly, or monthly in a named time zone. It stores local hour and minute, weekly day or monthly day, URL cap, run policy, run cap, monthly budget, failure behavior, balance behavior, jitter, next due time, and window counters.
The current default limit is five configured schedules per site. More schedules do not create better monitoring when they measure the same scope without a distinct owner or decision.
Due is not the same as started
The dispatcher checks that the site still exists and is verified, no site or organization capacity conflict exists, the organization quota allows another full run, the schedule has not reached its run or budget cap, and the organization can reserve the estimate. It then creates a run and dispatch outbox record.
A duplicate occurrence reuses its existing run. A due time, accepted database row, dispatched event, claimed worker, completed crawl, and published report are separate proof points.
Read visible skip reasons
- Site unavailable or unverified
- Another site run active or organization capacity full
- Organization monthly quota reached
- Schedule monthly run cap reached
- Schedule monthly budget cap reached
- Insufficient organization balance
Run caps and budgets use the schedule time zone
Each schedule keeps a monthly window, number of runs started, and estimated cost reserved in that window. The effective run cap is the lower of the schedule value and the current server policy. The budget check uses the next estimate before reservation; final settlement still follows actual units on the run.
Pause, resume, edit, or delete deliberately
Pausing clears the next due time and can store a reason. Resuming calculates a new due time from the current moment. Editing an active schedule recalculates its next due time and current budget window. Deleting soft-deletes it and removes the due time; it does not erase historical runs.
Choose one of the recorded alert questions
Current alert kinds cover new broken links, new noindex, robots changes, canonical changes, title changes, meta-description changes, structured-data regression, sitemap regression, Site Health drop, and bot anomaly. A rule can add severity, JSON thresholds, a suppression window from 0 to 720 hours, a digest rhythm, and optional schedule scope.
Alerts need a comparable run
Active rules are evaluated after a completed or partial comparable run. The current policy uses a previous successful comparison baseline, optionally scoped to the same schedule. Raw finding and inventory evidence is read through a bounded scan, and the run has an event cap.
No previous run, incomplete coverage, or a suppressed duplicate is a different state from no change.
Send digests to people who own the response
A rule can target workspace users or explicit email addresses, with a current cap of ten recipients per site. Digest choices are per run, daily, or weekly. The rule creator is used as the default recipient when no explicit list is supplied.
Event and email delivery have separate histories
An alert event stores the matched rule, run, severity, summary, counts, and bounded samples. Delivery then has its own queued, claimed, retry, sent, delivered, bounced, complaint, opened, clicked, or error evidence depending on provider events. A generated event is not proof that every recipient received the email.
Use suppression and digests without hiding regressions
The default suppression window is 24 hours. Site-level delivery is also rate-bounded. Use these controls for repeated fingerprints and expected fluctuations, but keep run history available so suppressed notifications do not become missing technical evidence.