The idea is simple. Every member of staff gets a ChatGPT-style assistant they can open in a browser, log into with their company account and use for drafting, summarising, analysis and questions about company documents. The difference from the public tools is that it runs on a server your company owns, inside your building, and nothing typed into it leaves your network.
This article explains how that is built in practice, layer by layer, and what decisions you need to make along the way. If you want the wider business case first, start with Your Staff Are Already Using ChatGPT: Why Nigerian Companies Need Private AI.
The architecture in one picture
A company-wide private AI assistant has five layers. Each is a separate, replaceable component, which means you are never locked into one vendor's choices.
| Layer | What it does | Typical choice | Decision to make |
|---|---|---|---|
| Hardware | Runs the models; GPU memory sets model size and concurrency | AI Series server with UPS | Tier, based on headcount and model size |
| Models | The actual language models staff talk to | Llama, Qwen, Mistral, Gemma (open weights) | One general model or several specialist ones |
| Inference engine | Loads models onto the GPU and answers requests | Ollama (simple) or vLLM (high concurrency) | Start simple or build for scale from day one |
| Web interface | The ChatGPT-style front end staff use | Open WebUI | Login method, groups, retention settings |
| Knowledge (optional) | Lets the assistant answer from company documents | RAG document collections | Which documents, and who may search them |
Step 1: Choose and size the hardware
The server is a GPU workstation or server built for continuous operation. What matters most is GPU memory: it decides how large a model you can load and how many people can use it at once. Large models are usually run quantised, meaning compressed to use less memory with a modest quality trade-off, which is how a 32B model fits on a single 24GB card.
| AI Series tier | Core specification | Price (inc. VAT) | Rough fit for private AI |
|---|---|---|---|
| AI Research | Core i7-14700K, 64GB DDR5, RTX 4070 Ti Super (16GB VRAM) | ₦8,800,000 | Pilot or small team: roughly 5–15 light concurrent users of 7–14B models |
| AI Professional | Core i9-14900K, 128GB DDR5, RTX 4090 (24GB VRAM) | ₦18,600,000 | Department scale: roughly 15–50 staff on 14–32B models, quantised |
| AI Lab | Threadripper or Xeon W, 256GB ECC, 2× RTX 4090 or RTX A6000, 1600W redundant PSU, server chassis | From ₦25,000,000 | Company-wide, 70B-class models; consultation only |
Not everyone in a company uses the assistant at the same moment, which is why a department-scale machine can serve a larger headcount than its concurrent capacity suggests. Our detailed guide, Sizing an Office AI Server: Users, Models and Hardware for 10, 50 and 200 Staff, works through the numbers. For the hardware fundamentals, see Ollama and LM Studio hardware requirements and AI inference servers for business applications.
The sensible path for most firms is a pilot on the AI Research tier at ₦8,800,000, then a move to AI Professional at ₦18,600,000 or an AI Lab from ₦25,000,000 once real usage is known. The pilot machine does not go to waste; it often becomes a dedicated document-search or development server.
Power and placement
Put the server in a cool, locked room or a ventilated cabinet, not under someone's desk. Size the UPS for the GPU under full load, not for a typical office PC, and plan how it ties into your inverter, solar or generator so a grid outage does not take the assistant down mid-afternoon. Configure automatic safe shutdown if backup power runs low. Our guide to UPS units for workstations with extended runtime covers the principles.
Step 2: Choose the models
Open-weight models are downloaded once and run entirely on your server. The main families, Llama, Qwen, Mistral and Gemma, each come in several sizes. Practical guidance:
- One good general model · For most companies, a single capable general model in the largest size your GPU handles comfortably covers the bulk of office work.
- A small fast model · Useful for quick tasks such as rewriting an email, where speed matters more than depth.
- Specialists where justified · A coding model for the IT team, or a model strong in French for cross-border correspondence.
- Test on your own work · Model rankings on public leaderboards are a guide, not a verdict. Run your staff's actual tasks through two or three candidates before choosing.
Be candid with staff about capability. Local models are very good at everyday office tasks, but the best frontier cloud models still lead on the hardest reasoning and coding. A hybrid policy, private AI for anything sensitive and an approved cloud tool for non-sensitive heavy lifting, avoids pretending otherwise.
Step 3: Choose the inference engine
- Ollama · Very easy to install, manage and update models with. Excellent for a pilot and for small teams. Handles several simultaneous users, but is not designed for heavy concurrency.
- vLLM · Built for production serving. Uses techniques such as continuous batching to serve many users efficiently on the same GPU. More configuration, better throughput at department and company scale. See our vLLM hardware guide.
Both expose an API that the web interface connects to, so switching later is a configuration change, not a rebuild.
Step 4: Deploy the web interface
Open WebUI is the most widely used open-source front end for this. It looks and feels like ChatGPT, which keeps training to a minimum. The settings that matter:
- Individual accounts · Every user logs in personally. Where possible, connect it to your existing directory or single sign-on so accounts are created and removed with employment.
- Groups and permissions · Control which teams see which models and which document collections.
- Retention · Decide how long chat history is kept, and configure it. This is a policy decision with data protection implications, covered in our NDPA and generative AI article.
- Branding · Give it a name staff will remember. "Ask Legal-AI" or the company's own name gets more use than a generic product name.
Step 5: Keep it on the network, not the internet
The assistant should be reachable at an internal address such as ai.yourcompany.local, served over HTTPS with an internal certificate, and accessible only from the office LAN or the company VPN. It should not be published to the public internet. The models need no internet connection to answer questions; the server only needs outbound access, under IT control, when downloading model or software updates.
Step 6: Add company knowledge
Once the basic assistant is working, load document collections so staff can ask questions of the company's own material: policies, templates, product manuals, past proposals. This is retrieval-augmented generation, and it is the feature that usually turns a curiosity into a daily tool. Do it per team, with permissions that mirror who is allowed to see the source documents. We cover the method and its pitfalls in Chat With Your Company Files Privately: Local RAG for Nigerian Businesses.
Step 7: Policy, training and rollout
A tool without a policy recreates the shadow AI problem in a new form. Publish a one-page rule on what data goes to the private AI and what may go to approved cloud tools. Run short sessions with each team showing real examples from their own work. Start with one department, watch usage for two weeks, then expand. Our 30-day private AI rollout plan sets this out week by week, and our shadow AI article explains why the sanctioned tool needs to be at least as convenient as the unsanctioned one.
What you end up with
A browser bookmark on every desk that opens a capable AI assistant, with company logins, company retention rules, company knowledge and no data leaving the building. The hardware is a one-time purchase from our AI Series, protected by a UPS, and the software stack is open, so you are never dependent on a single vendor's pricing or terms.
Sephora Systems designs, builds, installs and supports private AI assistants for Nigerian companies, from Abuja with nationwide delivery. Book a private AI consultation and we will size the system to your headcount and workload. You can also ask Kitan, our site assistant, or WhatsApp us on +234 707 096 6669.