RRecords Labs Help Center

Adding documents and websites

The Add knowledge wizard brings files and websites into your organization's knowledge. It walks you through four steps: Source, Configure, Access, and Start. Contributors, Editors, Admins, and Owners can use it.

Open the wizard

  1. Go to Knowledge → Sources.

  2. Click Add knowledge in the page header.

The cards under the header are shortcuts. Upload files and Import media open the wizard on that choice, and More ways opens it at the start. Write an article opens the Notes editor, and Structured data opens the Tables page.

Step 1: Source

Pick what you are adding. Under Bring in your content:

  • Upload files — PDFs, docs, spreadsheets, images.

  • Add a website — crawl a docs site or help center.

  • Import media — YouTube and recordings.

  • Structured data — spreadsheets, entities, or an API. This leaves the wizard for the Tables page.

  • Browser capture — courses and pages behind a login. This leaves the wizard for the capture page.

  • Connect an app — Drive, Gmail, Slack, Notion and more. This leaves the wizard for Connectors.

Under Create it here, Write an article and Add an FAQ open the Notes editor instead.

Step 2: Configure files

Choose Upload files (from your computer) or File URLs (links to files on the web, one per line, with an Import CSV option).

Drop files on Drop files or click to browse. The wizard lists them as Queued and lets you remove any file before you start. Accepted types include:

  • Documents: PDF, DOC, DOCX, RTF, TXT, Markdown, HTML, JSON, XML, YAML, EML, ODT

  • Spreadsheets: CSV, TSV, XLSX, ODS

  • Presentations: PPTX, PPT, ODP

  • Images: PNG, JPG, WebP, GIF, BMP, TIFF, HEIC

  • Audio and video: MP3, WAV, M4A, MP4, MOV, WEBM and similar

Legacy .xls and Outlook .msg files are not supported; convert them to XLSX or EML first. Documents can be up to 1 GB and audio or video up to 2 GB. Files fetched from a URL can be up to 512 MB.

When a file already exists decides what happens to duplicates: Skip existing keeps the stored copy, Replace with new version takes its place. Matching is by source plus filename and content hash.

If you queue audio or video, Keep the original audio & video files appears. Off keeps only the searchable transcript. Video also adds Read video frames with vision, which makes slides, diagrams, and on-screen text searchable.

Step 2: Configure a website

Under What to read, choose Whole site (follows links from the start URL) or Exact pages (reads only the page URLs you paste).

For a whole site:

  • Start URL — where the crawl begins.

  • Understand images on pages — on by default. Meaningful images are described so diagrams and photos become answerable.

  • Also read linked PDFs — off by default.

  • How fresh — Daily, Weekly, Monthly, or Never. The default is Weekly.

  • Advanced — Only pages listed in the sitemap, Include subdomains, Only read under this path, Skip these paths (up to 20, comma-separated), and Page limit (leave blank to read the whole site).

If the site blocks robots or needs a sign-in, expand Does this site block robots or need a sign-in? (Whole site only). Access you save there applies to that website only. See How crawling works.

Step 3: Access

Everything this source brings in gets the same settings:

  • Who can see this? — Everyone, Specific people, or Just me.

  • External Channels — Not shown, Signed-in customers, or Public. This controls Help Center pages and chat widget answers.

  • Classification — Trust, Stays accurate for, Sensitivity, How the AI may use the wording, Owning team, Category, and Tags. Click Fill with AI to suggest values for the fields you have not set; your own picks are kept.

  • Redaction — secrets and IDs are always stripped before indexing. Turn on Also redact names, emails & phone numbers for public or widely shared sources.

The details are in Access, sensitivity, and trust.

Step 4: Start

Review the summary, then click Ingest. The wizard shows Files ingesting or Ingestion started and offers View progress.

What happens next

Each item goes through the same pipeline: fetch, extract, split into passages, embed, index. It becomes answerable as it indexes.

  • Where to watch: the Ingesting now section on the Sources page, or Knowledge → Sources → Activity. Runs show Queued, Discovering, Running, Paused, Completed, Failed, or Cancelled. Individual pages show Queued, Running, Completed, Failed, or Skipped.

  • Titles: uploaded files get a readable title generated from their content. Web pages keep the page's own title.

  • Summary, tags, and questions: for each document the AI writes a short summary, up to six tags (reusing your existing tags where they fit), three to five questions the document answers, and a category. A category you chose in the wizard always wins over the AI's suggestion.

  • Review: new content lands as unreviewed. It is already used in answers, but confidence is capped at Medium until an Editor approves it from Knowledge → Review. Content from a source you set to High trust is the exception: it can reach High confidence before review.

Was this article helpful?
Related articles
Connecting business systemsKnowledge & sourcesHow crawling worksKnowledge & sourcesLibrary searchKnowledge & sourcesNotes and the editorKnowledge & sources