How does an internal knowledge assistant answer from your documents without making things up?

It searches an agreed set of your documents for the passages that match the question, gives the model only those passages, and tells it to answer from them and cite the source. When nothing relevant is found, it says so and points to the person who owns that topic instead of guessing. This pattern is called retrieval-augmented generation, or RAG.

It reduces made-up answers; it does not remove them. What keeps it honest is a visible citation on every answer, a clear "I don't know", documents that are current, and a set of real questions checked before launch and after each document change.

Updated Sep 28, 2026

How it works

  1. Choose the documents and an owner for each

    Policies, manuals, price lists, procedures. Remove old versions so there is one copy of each policy, and name who updates it.

  2. Split and index them

    Each document is split into passages, and each passage is turned into an embedding, a list of numbers that captures its meaning, stored in a search index. Many setups add keyword search for names and codes.

  3. Retrieve before answering, within access rules

    The question finds the closest passages. Access rules apply at this step, so a question never pulls a passage the person asking could not open.

  4. Answer only from the passages, with a citation

    The model is told to answer from the retrieved passages, name the document and section, and say it does not know when they do not cover the question.

  5. Test with real questions

    Collect questions staff actually ask, each with the right answer and its source. Check them before launch and again whenever documents change.

  6. Keep the documents current

    When a document changes, re-index it. An out-of-date policy in the index is cited as confidently as a current one.

What it costs to run

Costs come from three places: turning documents into embeddings, storing and searching them in a vector database, and the model that writes each answer. These are list prices from the vendors' pages; a small document set may fit within a free tier for the first two, and the model cost grows with the number of questions.

ItemList priceSource
Pinecone Starter plan (vector database)Free; up to 2 GB of storage, 2 million write units and 1 million read units per monthPinecone, PricingRead on Sep 28, 2026
Pinecone Standard planUS$50 per month minimum usage; storage US$0.33 per GB per month; reads US$16 to US$18 per million read units, depending on cloud and regionPinecone, PricingRead on Sep 28, 2026
Embeddings with multilingual-e5-large on PineconeUS$0.08 per million tokens on Standard; 5 million tokens a month included on StarterPinecone, PricingRead on Sep 28, 2026
Claude Sonnet 5 through the Anthropic API (writes the answers)US$2 per million input tokens, US$10 per million output tokensAnthropic, Claude API Docs, PricingRead on Sep 28, 2026
Claude Haiku 4.5 through the Anthropic APIUS$1 per million input tokens, US$5 per million output tokensAnthropic, Claude API Docs, PricingRead on Sep 28, 2026

Third-party list prices, read on the date shown. They are not Betterlane prices.

What we learned building it

From our own demos on invented data, September 2026: a document-based chat for a fictional coworking space (Cedar Workspace) and a regression test suite for a research agent's reports on two fictional companies. Neither is a client system.

Show the passage, not just the answer
In the Cedar Workspace demo, every supported answer highlights the document passage it came from, and the actual document can be opened beside the chat.
Change the document, change the answer
A test edits the source document and checks that the next answer changes. That is how you know the assistant reads the document rather than repeating something fixed.
A missing fact gets an explicit reply
The document itself says only what it states is approved, and the demo says plainly when a question is not covered. It uses pattern matching and section selection, not a language model, so it is stricter than a model would be.
Fail any claim without a source
Our research-agent test suite fails a report with readable reasons, such as a gap with no evidence or an email called verified with no source. The same check fits a knowledge assistant: an answer without a citation fails.

When it is not worth it

  • If staff ask a handful of questions a week and the documents are few, a well-organized shared folder with search does the job.
  • If the documents contradict each other or are out of date, the assistant will cite the wrong one with confidence. Clean them up first.
  • If the answers depend on live data in another system (stock, bookings, balances), that is an integration with that system, not a document search.

Questions

Does the model have to be trained on our documents?

No. The documents are retrieved when each question is asked. Whether the model provider keeps or uses the data sent to it depends on its terms, so read them before connecting sensitive documents.

Can different staff see different documents?

Yes, if access rules are applied when passages are retrieved. Then a question cannot pull a passage its author could not open.

Can it answer HR, legal or medical questions?

It can point to the relevant policy passage. Decisions, and anything about a specific person, go to the responsible colleague.

How do we keep it up to date?

Give each document an owner, re-index when it changes, and remove old versions from the collection.

Want an assistant that answers your team from your own documents and shows where each answer came from?

See the service: Internal knowledge assistantsLet’s talk

All guides