How crawling works
When you add a website, Records Labs reads its pages on a schedule and keeps the copy in your knowledge up to date. This article explains how the crawler behaves and what to do when a site blocks it.
How the crawler identifies itself
Every request carries the user agent RecordsLabsIngestion/1.0 (+https://recordslabs.ai/ingestion). If a site refuses that identity with a 403, the crawler retries the page once with a standard browser user agent.
The crawler waits about one second between requests to the same site.
robots.txt
The crawler reads a site's robots.txt once and caches it for 24 hours. Pages the file disallows for the crawler are skipped with the reason "Skipped because robots.txt disallows this URL for the crawler." If robots.txt cannot be reached at all, pages are retried rather than read. There is no setting to ignore robots.txt.
Robots rules apply in Whole site mode. Exact pages and sitemap-only reads do not check them, although Skip these paths still applies.
What is captured
The main text of each page, converted to Markdown. Navigation, footers, scripts, and forms are dropped.
Pages that look like a JavaScript app shell (under 100 words of plain HTML) are rendered in a headless browser and read again, up to 40 pages per run.
With Understand images on pages on, up to 20 meaningful images per page are described by the AI. Logos, icons, and images under 160 by 120 pixels are skipped.
With Also read linked PDFs on, PDFs the pages link to are read too, up to 25 per page and 250 per run, on the same host.
Pages under 100 words are skipped, but their links are still followed.
Images, stylesheets, scripts, fonts, archives, and executables are never crawled as pages.
When a site blocks the crawler
A 401 or 403 response counts as Access blocked. Once at least 20 pages have been checked and 90% or more of the most recent ones were refused, a Whole site refresh stops early so the site is not hammered. Pages already in your library stay searchable, and nothing is removed. The person who started the run (or your Admins) gets an in-app notification and an email: "{host} is blocking our crawler." While the site keeps refusing, the schedule only sends a small check, daily for the first week and weekly after that, and resumes normally once the site lets the crawler in.
The run page and the source show Cannot access website: {host} blocks our crawler. People who can add sources see a Fix access button; everyone else is asked to find someone who can. You have two routes:
Grant website access
Use this when a page opens in your browser but the site blocks Records Labs.
Open the run and click Fix access, or expand Does this site block robots or need a sign-in? in the wizard.
If asked, enter the Page or PDF URL. From a run, the blocked page is filled in for you.
Click Test and save access. Records Labs tests the URL with your browser's request settings and, if it works, saves that access for this website only.
If the site needs a sign-in, a cookie banner, or a security check, expand Website requires login or a security check? and click Open secure browser instead. Sign in inside the remote browser window, then click I can view it. The session lasts 20 minutes.
Back on the run, the alert reads Access saved for {host}. Click Try again.
Saved access is encrypted, scoped to one website, and used only for ingestion. You can review, Renew, Revoke, or Remove it on the Website access card under Knowledge → Sources → Connectors.
Ask the site owner to allow the crawler
If saved access still does not get through, the run shows Fix access didn't get through. Click Copy request to copy a short message you can send to the site owner. It asks them to allow the user agent RecordsLabsIngestion through their firewall or bot protection (for example Cloudflare). When they confirm, click Try again.
If they need our IP addresses, or you want more help, the same alert offers two options. Contact support sends the website and run to Records Labs support from the app, and we reply by email. Chat with support opens the Records Labs support chat, already signed in as you with your organization, plan, and role, and tells us which website was blocked. It shows only when the support chat is available.
A 429 ("The website asked us to slow down") is retried automatically with backoff.
Refreshes and change detection
How fresh sets how often the site is checked: Daily, Weekly, Monthly, or Never. Only Admins can change it after the source is created. Admins can also click Refresh now on a completed website source to check for new and changed pages immediately.
Refreshes are designed to be cheap:
The crawler sends
If-None-MatchandIf-Modified-Sinceheaders, so an unchanged page costs one small request.A page whose extracted text and images hash the same as last time is skipped before any AI work.
Pages that rarely change are checked less often, up to four times the interval. Pages cited in answers in the last 30 days are always checked on schedule.
Every page is fully re-read at least every 60 days.
The crawler also reads
Sitemap:lines inrobots.txt(or/sitemap.xml) to find pages that are not linked. Once a site's sitemap dates prove reliable, a page whose date has not moved can be skipped without a request.
A page that returns 404 or 410 on two checks at least a day apart is moved to Trash. A 403 never removes anything.