Ask any operations manager where the company's knowledge lives and the honest answer is: everywhere. The HR handbook is a PDF on someone's laptop. The current price list is in a spreadsheet emailed last quarter. The standard supplier contract has four versions on the shared drive, and only the legal officer knows which one is current. So staff do the natural thing. They paste the documents into ChatGPT and ask it questions.
That habit is what our hub article, Your Staff Are Already Using ChatGPT: Why Nigerian Companies Need Private AI, is about. This piece covers the better version of the same idea: letting staff chat with company files through an AI that runs on a server you own, inside your office, where the documents never leave the building. The technique is called retrieval-augmented generation, or RAG, and for most businesses it is the single most useful thing a private AI server does.
What RAG Actually Is, Without the Jargon
A language model on its own only knows what it learned in training. It has never seen your leave policy, your client list or last month's board pack. RAG fixes that without retraining anything. When a staff member asks a question, the system first searches your own document library for the passages most likely to contain the answer, then hands those passages to the model along with the question. The model writes its answer from your material and, in a properly built setup, tells you which document it used so the answer can be checked.
Three pieces make it work:
- The document index · your files are split into short passages and converted into numerical "embeddings" held in a vector database, so the system can find passages by meaning rather than by exact keyword.
- The model · an open model such as Llama, Qwen, Mistral or Gemma, served by an engine like Ollama or vLLM on the AI server.
- The chat interface · a private, ChatGPT-style web page such as Open WebUI, opened in an ordinary browser over the office network. Staff log in with their own accounts and install nothing.
None of this requires sending a single file to a foreign data centre. The index, the model and the chat history all sit on one machine in your server room.
What Staff Actually Use It For
The use cases that land first are rarely glamorous. They are the questions that currently cost someone ten minutes of searching, or an interruption to a colleague. Illustrative examples of the kind of questions companies want answered:
- HR and policy · "How many days of compassionate leave do we give, and who approves it?"
- Sales and pricing · "What payment terms did we agree in the current distributor framework?"
- Legal and contracts · "Which of our supplier agreements renew automatically, and what is the notice period?"
- Operations · "What is the documented procedure for a generator changeover at the warehouse?"
- Onboarding · new hires asking the handbook the questions they would rather not ask their manager.
What It Can and Cannot Do
RAG is genuinely good, but it is not magic, and anyone who tells you otherwise is setting you up for a disappointing pilot. This is our candid summary.
| Task | How well local RAG does it | Notes |
|---|---|---|
| Answer factual questions from policies and manuals | Strong | The core use case; answers can cite the source document |
| Summarise a long report or contract | Good | Noticeably better on larger models (AI Professional and up) |
| Find a clause across many contracts | Workable | Needs careful document preparation and testing |
| Read scanned or photographed paper | Weak until processed | Scans need text recognition (OCR) before indexing |
| Calculate totals across spreadsheets | Unreliable | Use the spreadsheet; the model can misread numbers |
| Replace a lawyer's or accountant's judgement | No | It drafts and finds; a qualified person decides |
Two honest caveats. First, local open models trail the best frontier cloud models on the hardest reasoning tasks. For document question-answering the gap is much smaller, because the answer is usually in the retrieved passage and the model's job is to read and summarise it well. Second, RAG quality depends more on your documents than on your hardware. Duplicate, outdated and conflicting versions produce confident answers drawn from the wrong version. Cleaning up the library is half the project, and it is the half nobody budgets for.
Permissions: The Part Most Pilots Forget
A chat interface that can see every file will happily tell a junior officer what the managing director earns. The fix is structural, not a warning in the staff handbook. Good setups use separate knowledge collections per department, each mapped to a user group: HR documents visible to HR, the board pack to directors, the general handbook to everyone. Open WebUI supports knowledge bases and user groups for exactly this; the design work is deciding the groups before anything is uploaded.
Remember too that the chat logs are themselves sensitive. Questions people ask reveal what they are working on. Decide how long conversations are retained, who can review them, and say so plainly in your acceptable-use policy. Our article on NDPA 2023 and generative AI covers why this matters when personal data is involved.
Getting Your Documents Ready
Most companies hold a mix of born-digital files (Word, PDF, Excel, exported email) and a long tail of scanned paper. Start with the born-digital material, because it indexes cleanly and gives the pilot early wins. Scanned archives can follow once they have been run through text recognition and spot-checked.
- One source of truth · nominate a shared folder per collection that the system re-indexes on a schedule, rather than letting staff upload ad-hoc copies.
- Retire the old versions · archive superseded policies outside the indexed folder so they stop answering questions.
- Name an owner · each collection needs a person who keeps it current, usually the department head or their delegate.
What Hardware It Takes
RAG adds less load than people expect. Building the index is a periodic job; the day-to-day load is the chat model generating answers. What RAG does change is prompt length: every question arrives with several retrieved passages attached, so each conversation uses more GPU memory than plain chat. That pushes document-heavy teams toward more VRAM sooner. These are the Sephora AI Series tiers we deploy for private document chat:
| Tier | Price (inc. VAT) | Core spec | Fits roughly |
|---|---|---|---|
| AI Research | ₦8,800,000 | Core i7-14700K, 64GB DDR5, RTX 4070 Ti Super 16GB | Pilot or small team, roughly 5–15 light concurrent users; 7–14B models |
| AI Professional | ₦18,600,000 | Core i9-14900K, 128GB DDR5, RTX 4090 24GB | Department scale, roughly 15–50 staff; 14–32B models quantised |
| AI Lab | From ₦25,000,000 | Threadripper / Xeon W, 256GB ECC, 2× RTX 4090 or RTX A6000, 1600W redundant PSU | Company-wide; 70B-class models; roughly 200 staff means AI Lab or multiple servers, by consultation |
For the full sizing logic, see Sizing an Office AI Server for 10, 50 and 200 Staff. If you have developers who want to build custom retrieval pipelines rather than use Open WebUI's built-in tools, our developer-focused piece on workstations for RAG and vector database development goes deeper, and our Ollama hardware guide explains how model size maps to memory.
One Nigerian-specific point: an outage halfway through a re-index can leave the vector database in a mess. Every server we install goes in with UPS and inverter planning, automated clean shutdown, and the index on fast NVMe storage so a rebuild, if one is ever needed, takes hours rather than days.
A Sensible Hybrid Policy
You do not have to choose one AI for everything. The policy we usually recommend is simple: private RAG for anything containing client data, staff records, contracts or unpublished financials; an approved cloud AI tool for public information, generic drafting and non-proprietary coding help. That gives staff the frontier models where the data is harmless and keeps the sensitive material in the building. Our piece on shadow AI in Nigerian workplaces explains why a sanctioned option, rather than a ban, is what actually changes behaviour.
How Sephora Sets It Up
- Discovery · which documents, which teams, and who should see what.
- Build · an AI Series server sized to your user count and document volume, assembled and tested in Abuja.
- Install · on-site setup anywhere in Nigeria: office network, user accounts, knowledge collections, UPS and shutdown automation.
- Handover · staff training, an owner per collection, and ongoing support.
If you want staff to chat with company files without those files leaving your building, book a private-AI consultation and we will scope the documents, the permissions and the right server with you. You can also ask Kitan, our site assistant, or message us on WhatsApp at +234 707 096 6669.