Site Crawler¶
Site Crawler is Experimental and off by default. Enable it in Settings → Features → Site Crawler to show it in your workspace sidebar.
Site Crawler saves desktop and mobile screenshots of a website, with descriptions of page elements and recorded navigation. Use it to inspect pages, compare captured states, and replay the interactions that were saved.
Current scope: a crawl creates reusable captures. Synthetic-shopper experiments do not yet run from these captures. Replaying a crawl does not visit the live site or place an order.
Start a crawl¶
- Select the property you want to work with in the property picker.
- Click Site Crawler in the sidebar.
- Click New crawl.
- Choose Default Property to use the property's current URL, or Variant to enter another public address in Variant URL. Choosing Variant creates a separate crawl; it does not create an experiment variant.
- Set Maximum pages and Time limit (minutes) using the limits below.
- Click Review crawl to check the credit reservation.
- Click Start crawl to reserve the credits and open the Base checkpoint. Every fresh crawl starts its own family.
| Setting | Default | Allowed values |
|---|---|---|
| Maximum pages | 100 | 1–500 unique page URLs, shared across desktop and mobile |
| Time limit (minutes) | 60 | 5–120 minutes, starting when capture begins |
| Devices | Desktop and mobile | Both are included in every crawl |
The crawler follows pages on the starting site after any initial redirect. It can record menus, dialogs, product choices and cart interactions. After adding an item to the cart, it opens the checkout page once and saves it as a final view without entering or submitting anything; payment and order steps are never opened, and purchase and personal-information submission controls are described without being used. That saved checkout view is what a 100-shopper experiment with a checkout goal verifies its conversion target against. A page or interaction may remain uncaptured if it is blocked, unavailable or outside the crawl's limits.
Expand a checkpoint¶
Open a finished checkpoint and select a saved view. You can branch from any completed, partial, stopped or failed checkpoint with a usable view after it has finished saving and returning unused credits.
- Select an unexplored button or link in the screenshot or interaction inspector, then choose Continue exploration to capture its result and continue along that branch.
- Choose Explore from this view to explore direct actions first, followed by deeper paths while limits allow. Already captured paths can lead to unexplored descendants.
- Choose Guided exploration and describe what to explore. The browser follows permitted observed actions, stops when the request is satisfied, and reports what it completed or could not do. Where it's available, turn on Allow form filling to let it type short values from your instructions into fields such as a search box (for example, "search the menu for pizza"). It only submits search and filter forms; contact, sign-up, login and checkout forms are never submitted, and it never types real personal details.
Give the new checkpoint a unique name, then review the limits and credits. Names are 1–80 characters after trimming spaces and must be unique across the whole family, ignoring capitalization. Base is reserved for the first crawl.
| Setting | Continuation behavior |
|---|---|
| Maximum new views | Defaults to 20; allowed range 1–500. Counts newly saved exploration views only. |
| Maximum pages | Inherits the parent's setting; 1–500 URLs. The starting URL counts. |
| Time limit (minutes) | Inherits the parent's setting; 5–120 minutes, including restoration. |
| Starting view/device | The selected saved view and its desktop or mobile device. |
The new checkpoint includes its parent's complete map plus the results it adds. Earlier checkpoints and sibling branches remain unchanged. Its progress shows the added views (against Maximum new views), explored URLs, remaining actions and charges beside the cumulative map.
Find and switch checkpoints¶
- In Site Crawler history, a Base crawl with continuations shows N checkpoints (and how many are still running). Click it to list them as a tree: each checkpoint sits under the one it continued, with its mode (Target, View or Guided), device, new views, age and status.
- On a results page, the path above the title (All crawls / Base / …) links to each checkpoint it continues. Click Checkpoints beside the status to open the whole family as a tree and switch to another checkpoint.
- A checkpoint's results page summarizes what it set out to do: its mode and device, the checkpoint it continues, the instructions of a guided checkpoint, where it started and how it ended (Completed, Partly done or Could not finish, with the explanation). Click the starting view to select it on the map.
- On the map, views the checkpoint added carry New in this checkpoint and the restored starting view carries Verified start. Turn on Only this checkpoint to hide the views it inherited.
The crawler must restore and verify the selected view on the live site before continuing. A step it cannot restore or complete is noted in the crawl log and the checkpoint carries on: a View checkpoint moves to its next queued action, and a Target or Guided checkpoint tries again up to three times (a Guided checkpoint asks the model for another action). A checkpoint runs until you stop it, it reaches its limits, or it has nothing left to explore, and everything it saved stays available either way. It never repeats an uncertain cart change to reconstruct a view. Start a fresh Base when you need a new snapshot of changed content.
Access and launch limits¶
Workspace members with permission to view the property can inspect its saved crawls. Owners, admins and members can start and stop crawls. Viewers cannot launch or stop them.
Only one crawl can be active in a workspace at a time. A demo, suspended or unapproved workspace cannot launch. If New crawl is unavailable, read the message on the history page; launch access may also be disabled while the feature is being prepared for your workspace.
Credits and refunds¶
A Base crawl reserves Maximum pages × 2 credits, covering one desktop and one mobile capture per URL. The default reservation is 200 credits. Reduce Maximum pages if you want a smaller reservation; reducing the time limit alone does not reduce it.
The final charge is 1 credit per successfully saved page and device, including its screenshot and completed component descriptions. Additional dialog, menu and other interaction states for that page and device are included. A screenshot with incomplete descriptions can remain visible without being charged.
A continuation reserves the smaller of Maximum pages and Maximum new views for its selected device. It charges one credit per URL/device delivering new, fully described exploration results, even when that URL was visited in a parent. Inherited screenshots, restoration-only work and duplicate outcomes are free.
Unused credits return automatically when the crawl finishes, fails or is stopped. For example, a crawl that reserves 20 credits and delivers six chargeable page/device captures charges six and returns 14. A crawl with no chargeable captures receives a full refund.
Crawls use the workspace's shared credit balance but do not count toward the monthly experiment limit. See Credits and tests for experiment pricing.
Follow progress¶
The results page adds screenshots as they are saved. It shows captured pages, pending pages, captured states, desktop/mobile state counts and the credit reservation. State counts include dialogs and interactions, so they can be higher than the page count.
The site size is not known in advance. Counts show the work discovered and saved so far, not a percentage of the entire website. You can leave the results page and return through Site Crawler; history lists Base crawls newest first, each with its checkpoints.
To stop an active crawl, click Stop crawl on its results page. Saved captures remain available while the crawl finishes cleanup and returns unused credits.
Inspect the map¶
- Select Desktop or Mobile. Desktop is selected first; the tabs show the states saved for each device.
- Enter a page title or URL in Find a page or dialog to narrow the map.
- Drag the map to pan, or use the zoom and fit controls to adjust the view.
- Select a screenshot to open its details in the inspector on the right. Click Focus selected page to center it, or the panel button at the end of the toolbar to hide the inspector and give the map the full width.
The inspector names the selected view (Page capture, dialog, drawer, interaction, cart or checkout) with its title, URL (Copy URL copies it), device and capture time, then splits its details into four tabs:
- Overview — Capture notes explain in plain words anything that cut the capture short (a height limit, a blocked page, incomplete descriptions, a popup that could not be closed), most serious first; At a glance counts components, actions and observations (select one to open its tab); Details lists the device, screenshot size and capture time.
- Components — every recorded element, filterable by name, role or description and by Interactive or Fields. Select one to outline it on the screenshot and read its description, role, link, field, observations and recorded actions.
- Observations — measured and inferred page behavior, each linked to its component.
- Actions — the view's navigation and interactions grouped by outcome: Captured (replayable), Unexplored, Blocked, Failed and Uncertain, each with its reason.
Each page is labeled Page title – URL. Dialogs and other interaction states have their own screenshots and state labels. A captured dialog can therefore appear beside the underlying page after it was closed.
Map screenshots show the first screen of each capture at the device's proportions (desktop, tablet or phone), and the label says how many screens tall the page is. To see the rest, right-click a screenshot (or press the context-menu key or Shift+F10 on it):
- View full screenshot shows the whole page at once. Choose Fit width to read it at full width and scroll.
- View scrollable screenshot shows a window the size of the device's screen that you scroll through as a visitor would; Previous screen and Next screen move one screen at a time, and the position shows which screen you are on.
The inspector's Screenshot: Full · Scrollable links open the same views for the selected capture, and Open image opens the original in a new tab.
Arrows show navigation and interaction links. Solid links represent captured actions; dashed links represent discovered navigation. A discovered link alone does not prove its destination can be replayed.
Component descriptions¶
- Enable Component viewer to show outlines around the recorded elements.
- Select an outline, or a row in the inspector's Components tab, to read its description. A selected component stays outlined even with the viewer off.
- Read the Observations tab and the Capture notes on Overview for the supporting observations and any limits.
Measured observations come from what the crawler recorded, such as an element's position at several scroll offsets. AI inference identifies an interpretation of the captured content. A scroll behavior marked unknown was not established by the observations. Descriptions do not alter the screenshot pixels.
Replay the captured site¶
- Select the page or dialog where you want to begin.
- Click Replay captured site.
- Click a recorded element in the screenshot, or choose a captured action on the inspector's Actions tab, to follow its saved destination.
- Click Back to return to the previous captured state.
- Click Back to map to leave replay.
Replay uses saved states only. If an action was not captured, it displays the recorded reason instead of opening a live page or inventing a result. A cart state follows only destinations recorded with that same cart context. Verified restoration links connect saved views to their newly restored browser context. Switching devices starts a separate replay history.
Partial or missing results¶
Partial, failed and stopped crawls keep the results they saved. Read the reason under the progress counts, which explains in plain words why the crawl stopped, then select an affected screenshot and read its Capture notes.
| What you see | What it means |
|---|---|
| A page limit or time limit reason | The crawl reached a configured boundary; some pages or interactions may remain uncaptured. |
| A cropped or incomplete page capture | Only the recorded portion of the document is available. |
| Incomplete component observations | The image was saved, but some descriptions could not be completed. |
| A blocked or unexplored action | There is no captured destination for that action. |
| No pages were captured | The crawl ended before a page state could be saved. |
Capturing a page does not mean every interaction on it was captured. Check both page coverage and the available actions before using the result. Continue from a usable view to expand missing coverage, or start a new Base to capture changed content. A checkpoint never replaces an earlier one.
Read the crawl log¶
Click Crawl log on a results page to see what the crawler did, in order, and why anything came out partial. It updates while the crawl runs.
- The summary says how the crawl ended and lists Why coverage is incomplete: each reason in plain words (a page limit, a blocked page, a page that took too long, incomplete descriptions…), how often it happened, and what it means.
- The timeline shows every saved view, each page's outcome and, for a checkpoint, each step: restoring the starting view, the actions it explored and, for guided exploration, every decision with the model's reasoning.
- Filter to Issues or Agent decisions, pick a device or search the text. Show in map closes the log and selects that view.
- Download log saves it as a text file (times are offsets from the start of the crawl); JSON saves the structured version, including the underlying codes, for support.
Download the capture record¶
Click Download site-data.json on the results page to save the structured capture record. It describes pages, states, component bounds and observations, recorded actions, screenshot references and missing coverage. A download made while the crawl is active contains only the results saved so far; download again after it ends for the final record.
The JSON file is not a self-contained image archive. Screenshots remain private and require authorized access through Squoosh. Downloading or sharing the JSON does not make its referenced images public.