Not everything you want your agent to know lives on a web page.

A returns policy PDF, a wholesale price list in a spreadsheet, terms you keep in a Word doc: none of that has to become a web page first. Upload it directly and your Helpforge agent answers from it exactly the way it answers from your crawled site, with the same citation, the same guardrails, and no rewrite required.

Text, Markdown, CSV, PDF and Word (.docx) today. Files are chunked and embedded the same way your website is, and they survive every re-crawl.

Uploaded, not crawled

What's the lead time on a 12-unit case of the peppermint blend?

The 12-unit case of Peppermint Herbal ships with a 5 business day lead time.

Grounded in: Wholesale pricing (uploaded file)
Why it matters

"Just put it on your website" is not always the answer.

A lot of the content that actually settles a support question was never written to be a public page, and shouldn't have to become one just so a bot can read it.

Some documents aren't public

A wholesale price list, a wait-list agreement, an internal returns exception policy: businesses have real documents they want their agent to know that they would not publish as a page. A crawl can only ever see what is already live on the site.

If it's not a URL, a crawler cannot find it, no matter how well it's written.
Rewriting a PDF into a web page is real work

Most policies already exist as a finished PDF or Word document. Turning that into formatted HTML just to get it indexed is busywork nobody asked for, and it creates a second copy of the truth that can drift from the original.

The document you already have should be enough.
A spreadsheet answers a different kind of question

"What's the lead time on the 12-unit case" is a lookup across a row and a column, not a sentence anywhere on your site. A price list or spec sheet in a CSV holds exactly that shape of answer, and prose was never going to say it as precisely.

Structured data deserves to be read as structured data.
Supported today

Five formats, the ones businesses actually hand us.

Add a file from the dashboard yourself, or send it to us and we upload it for you.

PDF

Text is extracted directly. A scanned, image-only PDF with no real text layer is refused with a clear reason rather than silently indexing nothing, since that needs OCR, which this does not do.

DOCX

Word's modern format. Paragraphs and table rows are read out as prose. The old binary .doc format is refused by name with instructions to re-save it as .docx.

CSV

Every row is rendered with each value labelled by its column header, so a two-column lookup like a lead time next to a case size reads correctly instead of as a stray line of commas.

TXT / MD

Plain text and Markdown are read as-is. The fastest path if you can already paste the content directly into the dashboard.

How Helpforge is built

A document is not a second product bolted onto the crawler.

It joins the same pipeline your website already goes through: chunked, embedded and ranked by the one retrieval step that decides what the agent gets to read.

Ranked, not reserved

An uploaded document competes for a slot in the agent's context on the same score as a crawled page, no fixed share set aside for either one. Upload a document and it changes which content the model reads, never how much, so a business with no documents gets exactly the ranking it always had.

The best match wins, whichever source it came from.
Survives every re-crawl

A re-crawl replaces your website's own pages, and a document you handed us directly is not part of that sweep. It stays in place until you remove or replace it yourself, so refreshing the site's content can never quietly erase a policy PDF you uploaded weeks ago.

A crawl updates what changed on your site. It has no opinion on a file you handed us.
Re-uploading replaces, it doesn't duplicate

Uploading a file under a title you already used replaces that document rather than stacking a second, possibly conflicting, copy of it. Fixed the price list, corrected a typo in the policy? Upload it again with the same title and the old version is gone.

One title, one current version, never two competing answers.
Same guardrails as everything else

A document's text is untrusted content in the prompt, exactly like a crawled page, and the same decline-never-keeps-a-source and stay-in-scope rules apply to an answer grounded in a file. Read the full detail on our hallucinations page.

A PDF you uploaded gets no more trust than a page it did not write.
Limits, on purpose

A file upload is still an untrusted input, so it has ceilings.

A crawler already has to survive a hostile website. A document upload gets the same discipline: fixed bounds on size, length and processing time, checked before any work starts.

Fixed ceilings, not best-effort

Every uploaded file is capped at 15MB, extracted text is capped at 300,000 characters, and a PDF that claims more than 300 pages is refused outright rather than partially processed. A page count is checked before extraction begins, closing a real gap where a malformed PDF could otherwise claim millions of pages and exhaust memory on a file that never had that much content.

A time ceiling, not just a size one

PDF extraction runs under a hard 10-second wall-clock limit, because a small file can still be slow to parse: a dense, highly compressed page can cost real processing time regardless of its byte size. A file that runs past the ceiling is stopped and refused with a clear reason instead of tying up the upload indefinitely.

Simple pricing

One flat build fee, one flat monthly.

Document uploads are part of the agent we build for you, not a higher tier. Upload as many policy files as you need from day one.

$2,500 one-time build
+ $400 per month

Free demo on your real site first. Send us a policy PDF or price-list CSV and ask your demo a question only that file can answer.

Start with a free demo
FAQ

Document uploads, answered honestly.

Can I hand my AI chatbot a PDF instead of putting it on a web page?

Yes. Upload the PDF from your dashboard, or send it to us and we upload it for you. Its text is extracted, chunked and indexed the same way a crawled page is, and your agent can answer from it directly. A scanned, image-only PDF with no real text layer is refused with a clear reason, since that needs OCR, which this does not do.

What file types can I upload?

PDF, Word (.docx), CSV, plain text and Markdown. The old binary .doc format is refused by name with instructions to re-save it as .docx, since it is not a format this reads.

Does an uploaded document get priority over my website's own pages?

No, and that is deliberate. A document competes for a slot in the agent's context on the same score as a crawled page, with no fixed share reserved for either source. The better match for the visitor's question wins, whichever source it came from.

Will re-crawling my website delete a document I uploaded?

No. A re-crawl replaces your website's own pages, and a document you handed us directly sits outside that sweep. It stays in place until you remove or replace it yourself.

What happens if I upload a corrected version of a file?

Uploading a file under a title you already used replaces the old version rather than adding a second, competing copy. There is never more than one current version of a given document.

Is there a size limit?

Yes, on purpose. Files are capped at 15MB, extracted text at 300,000 characters, and a PDF over 300 pages is refused outright. PDF extraction also runs under a 10-second time ceiling, since a small but densely compressed page can still be slow to parse. These are fixed bounds checked before any processing starts, not best-effort limits.

Does a document you uploaded get the same source citation a web page gets?

It grounds the answer exactly as well as a page does and is never counted as an unanswered gap, but most documents have no public web address for a citation link to point to, so no clickable source appears for that part of the reply. It still shows up correctly in your own analytics as a grounded answer. Read more on our source citations page.

Get started

Send us the document, we'll show you the answer.

Give us your website URL, and if you have a policy PDF, price list or Word doc you want the agent to know, send that along too. We build a working demo, usually within 2 business days, with a private link to test it yourself before any commitment. No call, no credit card.

No call required. We reply by email.

Got it, we are on it. Your demo link will land in your inbox within 2 business days.