What Server You Need for a Local LLM with Ollama (2026)
Quick answer: The practical rule is twice as much RAM as the model’s size. A small 3-billion-parameter model wants about 4 GB; a 7-8B one, about 8 GB; a 13B, about 16 GB. It runs without a graphics card, but slowly: fine for background work, not for live chat. And if your goal is saving money against a paid API, it almost never wins: you do this for privacy, not price.
Running a language model on your own server has become straightforward. The hard part is sizing the machine, because here a mistake doesn’t mean “a bit slow” — it means “won’t start” or “the system killed the process halfway through”.
🔎 This article contains affiliate links. If you sign up through them we earn a commission at no extra cost to you. Read our affiliate policy.
The RAM rule
A model occupies memory just by being loaded, before answering anything. How much depends on two things: how many parameters it has, and how heavily it’s compressed (what’s called quantisation).
| Model size | RAM to load it | Recommended server RAM |
|---|---|---|
| ~1 billion parameters (1B) | ~1 GB | 4 GB |
| ~3 billion (3B) | ~2 GB | 4-8 GB |
| ~7-8 billion (7B/8B) | ~5 GB | 8-16 GB |
| ~13 billion (13B) | ~8 GB | 16 GB |
| ~30 billion or more | 20 GB+ | 32 GB+ |
The second column is what the model occupies; the third is what the machine should have, because you also need memory for the operating system, for the conversation context, and for whatever else runs there. Sizing by the second column is the classic error: the model loads, and the server runs out of air as soon as a long request arrives.
Models downloaded by default usually arrive already quantised, which is what lets something needing 16 GB uncompressed run in 5 GB. You lose a little quality in exchange — usually not much.
Do you need a graphics card?
You don’t need one, but it changes the experience completely.
Without a GPU, the model runs on the processor. It works: it answers, and the answer is just as good. What changes is speed — from a few words per second to a comfortable reading pace.
That defines what each setup is for:
- Without a GPU: background tasks. Classifying incoming email, summarising documents overnight, tagging tickets. Nobody is waiting at a screen.
- With a GPU: anything conversational where a person is waiting.
Most budget VPS plans have no GPU, and the ones that do cost several times more. Before buying one with a card, be sure your case is the second one.
What it actually costs
For the realistic case — a 7-8B model doing background work, no GPU — you need a server with at least 8 GB of RAM. At Hostinger that’s the KVM 2 plan: $8.99/mo on an annual term, renewing at $14.99/mo. If you’ll also run n8n or other services on the same machine, step up to 16 GB.
Pricing verified 6 September 2026. The full plan breakdown is in which VPS you need.
Now the part almost no guide says: that is not cheaper than paying for an API. Fifteen dollars a month of a commercial API buys an enormous number of requests, with a better model, and nothing to maintain. If your motive is saving money, run the numbers first; it usually loses.
So why do it?
For privacy, and it’s an excellent reason in some sectors. If the data you’ll process is a medical record, a case file, or information about minors, keeping it on a machine you control is an argument worth more than the price difference.
The other two legitimate motives: predictable fixed cost at enormous volume, and working without internet access in isolated environments.
If your case is “I want to try AI in my processes”, start with a paid API. Migrate later, if the reason appears.
What to run on top
Ollama on its own is an engine with no interface: you talk to it via command line or API. The usual move is adding a usable layer:
- Open WebUI to get a ChatGPT-like chat with users and permissions. Add about 2 GB of RAM.
- n8n to connect the model to your processes: read email, summarise, write the result somewhere. Add another 2 GB.
Add it all up before picking a plan. Ollama with a 7B, plus Open WebUI, plus n8n on one machine wants 16 GB for comfort, not 8. The project directory lists the RAM for each one so you can do the arithmetic.
FAQ
Which model should I start with? One from the 7-8B family in its quantised version. It’s the balance point: fits in 8 GB and gives reasonable results for summarising, classifying and drafting.
Can I run it on my laptop instead of a server? Yes, and for testing it’s the most convenient option. A laptop with 16 GB handles a 7B fine. What you can’t do is keep it available 24/7 for other processes to call; that needs a server.
Is it legal to use these models in my company? It depends on the model: each has its own licence and some restrict commercial use. Check the model card before putting it into production.
Will it be as good as ChatGPT? No. A model that fits in 8 GB doesn’t compete with the large commercial models. For narrow tasks — summarising, classifying, extracting data from text — it’s more than enough; for complex reasoning, the gap shows.
How slow is it without a GPU? On the order of a few words per second with a 7B on a modest server. Acceptable for something automatic; uncomfortable for live chat.
🔎 This article contains affiliate links to Hostinger. Read our affiliate policy.
