The first question every company asks us about private AI is "how big a server do we need?" The second is usually "we have 50 staff, so we need something for 50 people, right?" The honest answer to the second question is no. Headcount is the wrong number to size an office AI server by. What actually decides the hardware is how many people are generating answers at the same moment, which model you want them to use, and how long their prompts are.
This guide walks through those three numbers and then gives worked recommendations for offices of 10, 50 and 200 staff. If you are new to the idea of a company-owned AI, start with our hub article, Your Staff Are Already Using ChatGPT: Why Nigerian Companies Need Private AI.
The Three Numbers That Decide the Hardware
1. Concurrent users, not headcount
Office AI use is bursty. People ask a question, read the answer, go back to work. At any given second, only a fraction of staff are waiting on the server. For planning, we assume between 10% and 20% of headcount active during a busy hour, with fewer than that generating at the same instant. That is a planning assumption, not a survey result, and we refine it during a pilot by looking at real usage logs.
2. Model size
Open models come in size classes measured in billions of parameters. Bigger models write better, follow instructions more reliably and reason more carefully, but need more GPU memory. Quantisation (storing the model at lower precision, typically 4-bit) shrinks that memory need dramatically with a modest quality cost, which is why almost every office deployment runs quantised models.
3. Context length
Every active conversation holds its history and any attached documents in GPU memory. Plain chat is light. Document chat, where each question arrives with several retrieved passages (see our guide to local RAG), is heavy. Teams that work with long contracts or reports need more VRAM per user than teams asking short questions.
How Model Size Maps to GPU Memory
These are approximate figures for the model weights alone at 4-bit quantisation. Real deployments need extra headroom on top for conversations.
| Model class | Approx. VRAM for weights (4-bit) | Typical office use | Fits on |
|---|---|---|---|
| 7–8B (e.g. Llama 8B, Qwen 7B) | About 5–6GB | Fast everyday chat, drafting, simple Q&A | 16GB card with plenty of room |
| 12–14B (e.g. Qwen 14B, Gemma 12B) | About 8–10GB | Better drafting and document Q&A | 16GB card, fewer simultaneous long contexts |
| 27–32B (e.g. Qwen 32B, Gemma 27B) | About 17–20GB | Strong summarisation and reasoning | 24GB card |
| 70B-class (e.g. Llama 70B) | About 40GB or more | Closest to frontier quality locally | Two 24GB cards or a 48GB workstation card |
Our Ollama and LM Studio hardware guide explains these figures in more depth, and running Llama 70B locally covers the top row specifically.
The Serving Engine Matters as Much as the GPU
- Ollama · simple to install and manage, excellent for pilots and small teams, and happy to switch between several models.
- vLLM · built for throughput, batching many users' requests together so the GPU stays busy. As concurrency climbs, vLLM gets far more out of the same hardware. See our vLLM serving guide.
- The chat front end · a private ChatGPT-style interface such as Open WebUI sits in front of either engine, so staff see the same thing whichever we choose.
Worked Recommendations: 10, 50 and 200 Staff
The table below applies the planning assumption above to three office sizes and maps each to a Sephora AI Series tier. Prices include VAT.
| Office size | Planning busy-hour users (10–20%) | Sensible model class | Recommended tier | Price (inc. VAT) |
|---|---|---|---|---|
| 10 staff | 1–2 | 7–14B | AI Research (i7-14700K, 64GB, RTX 4070 Ti Super 16GB) | ₦8,800,000 |
| 50 staff | 5–10 | 14–32B quantised | AI Professional (i9-14900K, 128GB, RTX 4090 24GB) | ₦18,600,000 |
| 200 staff | 20–40 | 32B to 70B-class | AI Lab (Threadripper / Xeon W, 256GB ECC, 2× RTX 4090 or RTX A6000) or multiple servers | From ₦25,000,000, by consultation |
10 staff: AI Research
A ten-person office, whether a consultancy, a small law practice or a finance team, is comfortably served by AI Research, which suits roughly 5–15 light concurrent users. A 16GB RTX 4070 Ti Super runs a 7–8B model quickly for everyday chat and a 14B model for better document answers. 64GB of system memory leaves room for the document index and the chat interface on the same machine. This is also the tier we use for pilots in larger companies.
50 staff: AI Professional
At 50 staff you are at the top of the AI Professional range, which suits roughly 15–50 staff. The 24GB RTX 4090 runs quantised 27–32B models, which is where answers start to feel noticeably more capable, and 128GB of memory handles larger document collections. If most of those 50 people do heavy document work all day, or headcount is growing, talk to us about AI Lab or a second server rather than running one machine at its limit.
200 staff: AI Lab or multiple servers
At around 200 staff the decision is architectural rather than a single box. AI Lab, with two RTX 4090s or an RTX A6000, 256GB of ECC memory and a 1600W redundant power supply in a server chassis, can serve a 70B-class model to a whole company. Some organisations are better served by two or three servers, for example one per site or one per sensitivity level. That is why AI Lab is consultation-only: the right answer depends on your sites, network and departments. Our piece on AI inference servers for business applications covers the wider design questions.
Everything Else in the Box
- System memory · holds the document index, the chat interface and, in a pinch, model layers that don't fit on the GPU. Don't skimp.
- Storage · NVMe SSDs for models and the vector database. Models are large files and you will keep several.
- Network · wired gigabit to the server. Staff can connect over Wi-Fi; chat traffic is small.
- Backups · user accounts, settings and document collections backed up off the server. The models themselves can be re-downloaded.
Power: Size the UPS With the Server, Not After It
A Nigerian office AI server needs UPS and inverter planning on day one. The planning figures below are approximate peak draws for the whole system, used to size protection; real draw at idle is far lower.
| Tier | Approx. peak system draw (planning figure) | Power protection we plan for |
|---|---|---|
| AI Research | Around 600–700W | Online UPS with headroom, inverter backup, automated clean shutdown |
| AI Professional | Around 800–1,000W | Online UPS with headroom, inverter or generator backup, automated clean shutdown |
| AI Lab | 1600W redundant PSU class | Dedicated circuit, online UPS sized to the redundant supply, generator integration |
The UPS rides through grid-to-generator switchovers and, if power doesn't come back, shuts the server down cleanly so the document index isn't corrupted. Heat matters too: a busy GPU server in a warm, unventilated store room will throttle. Our guide to AI workstation power and cooling in Nigeria covers both.
Four Sizing Mistakes We See
- Buying for headcount · paying for capacity that sits idle, or the reverse, under-buying because "only the managers will use it".
- Ignoring context length · a server that handles short chat fine can struggle once everyone starts attaching 40-page contracts.
- Chasing the biggest model · a fast 14B model that answers in seconds often gets used more than a slow 70B model that keeps people waiting.
- Leaving power until later · one hard power loss during a re-index costs more in downtime than the UPS would have.
The honest way to size an office AI server is to pilot, measure and then scale, and we are happy to do all three with you. Book a private-AI consultation and we will size the server to your staff, models and documents. You can also ask Kitan, our site assistant, or message us on WhatsApp at +234 707 096 6669.