Documents

Your documents are the fuel for your AI Corpus.

Your documents are the fuel for your AI Corpus. 🔋 They are what the Corpus reads,
analyses and draws on to answer your questions with precision. Each Corpus has its own
private, secure document space that you can update at any time.

🔐

Your documents are stored in a space dedicated to your Corpus. It is private and
secure
— your data stays confidential.

Accepted formats 🗂️

FormatExtensionMax size
PDF.pdf100 MB
Microsoft Word.docx100 MB
Microsoft PowerPoint.pptx100 MB
Microsoft Excel.xlsx100 MB
Plain text.txt100 MB
Markdown.md100 MB
Comma separated values.csv100 MB
🔥

Files that exceed these limits are rejected — check the size before uploading.

After an upload, your Corpus needs a few minutes to index and analyse the content before
it can answer questions about it. ❇️

Adding documents ➕

The Upload button at the bottom of the page opens a page of its own, where you build
your selection before anything is sent. Nothing leaves your machine until you press
Upload files.

  1. Choose your files. Drag and drop them onto the zone at the top of the page, or
    click Browse files to pick them from your machine. You can do both, several times:
    each selection is added to the list instead of replacing it, and any row can be
    removed with its ✕ button. Up to 20 files at a time.
  2. Tag them if you want — see Tags below.
  3. Choose Simple or Advanced. Simple is the default and is the right choice almost
    every time. Advanced is covered below.
  4. Press Upload files. Your documents are sent one after the other and the page shows
    which one is in flight.

Cancel takes you back at any point, including during the upload. Files already sent
keep going; the one in flight and the ones not started are dropped and cleaned up
automatically — nothing partial ever reaches your Corpus.

🖱️

Drag and drop works in your browser and in the desktop app. On phone and tablet,

use Browse files.

A file is left out and reported to you when it is too large, in a format we don't
accept
, or already in the list.

Tags 🏷️

Tags are free labels — a team, a year, a project — that you attach to your documents so
you can find and filter them later.

Type a tag in the upload page and press Enter, or separate several with a comma. Tags
apply to every file of the selection, which is what makes sending a batch together
worthwhile. A document accepts up to 32 tags of 64 characters each.

Once a document is in your Corpus, its tags appear under its name in the file list. When
a document carries more than five, the extra ones are folded into a +n chip — hover it
to read them.

Tags are not fixed at upload time: choose Edit in a file's menu to attach new ones or
take existing ones off, one file at a time. What you set there replaces the tags the
file carried, so leave in place the ones you want to keep.

Advanced options ⚙️

Most people never need these: Simple mode already sends everything your Corpus needs,
plus your tags. Advanced adds three sections for the cases that do.

  • Provider — a label recording where these documents came from. Defaults to user.
  • Metadata — free key and value pairs stored alongside each document.
  • Chunking — how each document is cut up before it is indexed.

Chunking, in short ✂️

Your Corpus does not read a whole document at once: it reads chunks. Bigger chunks give
the AI more context per answer but retrieve less precisely.

OptionWhat it does
StrategyBy title starts a new chunk at each section heading — right for reports, contracts and manuals. Basic ignores the structure and fills chunks to the limit — better for flat prose or transcripts.
Maximum charactersThe hard cap. No chunk ever exceeds it. Defaults to 10000.
New chunk afterA soft cap. Set it below the maximum for more even chunks.
OverlapCharacters repeated from the previous chunk, so an idea split across a boundary is not lost.
Overlap every chunkImproves recall across boundaries, at the cost of text duplicated in your embeddings.
Combine sections underBy title only — merges consecutive small sections up to this size.
Sections across pagesBy title only — lets a section span a page break.
ℹ️

Leave a field empty to keep the platform default. Sizes are in characters, not

words.

Previewing a document 👁️

Not sure a file is the one you were after? Open its preview and read it without leaving
the app — no download, no local viewer.

Click a file's row in your document list, or pick Preview in its ⋮ menu. Both
open the same page.

  • The page you are reading fills the middle of the screen.
  • Every page of the document is listed down the right-hand side as a thumbnail
    labelled Page 1, Page 2 and so on. Click one to jump straight to it.
  • The arrows under the page move you back and forward one page at a time, and the
    counter between them tells you where you are — Page 3 of 24.

Thumbnails load as you scroll, so a long document opens just as fast as a short one.

The header of the preview carries two buttons:

  • ℹ️ Details — what the document is about, and what it is. See
    below.
  • ⬇️ Download — save the original file to your machine.
⏳

A document has to be indexed before it can be previewed. Preview appears in the menu

once processing is done. If the pages are still being rendered, the preview tells you so —
come back in a few minutes.

💡

The pages you see are images of the document, rendered when it was indexed. To edit

the file, download it.

The details of a document ℹ️

Details, in a file's ⋮ menu, opens everything the platform knows about it — as does the
ℹ️ button in the preview header.

At the top, the record — with the first page of the document beside it, so you can tell at a
glance which file you opened:

Uploaded / UpdatedWhen the file arrived, and when its record last changed.
File sizeThe size of the file you sent.
Storage usedWhat the document really occupies: the file plus everything built from it — the page images, the text conversion, the summary and the embeddings. It is several times the file size, and it is the figure your usage totals are built from.
ChunksHow many passages the document was cut into. See below.
PagesHow many pages it has.
TokensWhat reading the document cost: the tokens spent writing its summary. Part of the token totals on your dashboard.
ChunkingThe settings the document was cut with — strategy, max_characters and the rest of what you chose under Advanced options. A document uploaded without them reads Ingested with the platform default option set.
MetadataAnything you attached to it at upload time.

Under it, the summary your Corpus wrote while indexing the document.

⏳

Storage used, Chunks, Pages and Tokens are computed while the document is

being indexed, and read Not computed yet until that finishes. The first page appears once the
page images have been rendered.

The chunks of a document 🧩

View chunks, at the bottom of the details, opens the passages your document was cut into.

Your Corpus does not answer from a whole document: it answers from
chunks. This page shows you those passages — in the order they appear in
the file, with the text exactly as it is stored — so you can see what your assistant actually
reads.

It is a troubleshooting view. Reach for it when an answer quotes a document and gets it
subtly wrong, and you want to see the passage that produced it.

Each card shows one passage: its rank in the document, the pages it covers, its text,
and the metadata attached to it. A chunk can cross a page boundary, so a span such as
Pages 2–4 is normal, and the summary of the document is stored as a chunk of its own,
belonging to no page.

From the ⋮ menu of a chunk you can:

  • ✏️ Edit — correct the text, the pages it covers, or its metadata.
  • 🗑️ Delete — take the passage out of the index.
⚠️

Editing does not re-index the passage. The chunk keeps being found on the wording it

had, and answers with the wording you gave it. That is what you want for a typo or a name to
remove. For a genuine rewrite, replace the file instead — that re-cuts
and re-indexes the whole document.

🗑️

Deleting a chunk is final. The passage stops being used in any future answer. The

document, its file and its other chunks are untouched, and only re-sending the file brings the
passage back.

🚨

A chunk shown in red has lost its stored text. It still comes up in searches and then

adds nothing to the answer. Delete it, or type its text back in.

Managing your files 💁‍♀️

Seven actions let you keep your Corpus's knowledge base up to date:

  • ℹ️ Details — choose Details in the file's menu to see its record, its summary and the
    passages it was cut into. See above. It is the one entry offered
    whatever state the file is in — including a file whose processing failed, where the record is
    what tells you how far it got.
  • 👁️ Preview — click the file's row, or choose Preview in its menu, to read it page
    by page without downloading it. See above.
  • ➕ Add — click the Upload button at the bottom of the page to open the upload page
    described above, where you stage your files and their tags before
    sending them.
  • ✏️ Edit — choose Edit in the file's menu to change its name and the
    tags attached to it.
  • ♻️ Replace — choose Replace in the file's menu to push a new version of the file.
    It keeps its name and its place, only its content changes, and your Corpus relearns it
    straight away. The new file must be of the same format, so only files of that format are
    offered to you.
  • ⬇️ Download — choose Download in the file's menu to save the current version to
    your machine.
  • 🗑️ Delete — choose Delete in the file's menu to remove a file.
⚠️

Deletion is final. Once deleted, a file cannot be recovered, and your Corpus loses

all knowledge of its content.

♻️

Replacing drops the previous content. The old version is not kept, and answers

that cited the file lose their link to it. Replacing is also how you recover a file whose
processing failed — no need to delete it and start over.

Going further with connectors ⚡

Rather than uploading files one by one, you can synchronise an Corpus with your existing
document sources — Google Drive, SharePoint, OneDrive, Confluence and more — so new
documents are indexed automatically. See the Pro Edition guide for the
full list of connectors.


Did this page help you?