For most Iranian organisations, using large language models through foreign services runs into three obstacles: the organisation's data has to leave the country, access to those services is unreliable, and paying for them and even opening an account is a struggle. A local model, one that runs on the organisation's own server, removes all three.
But a local model is not a full replacement for the largest commercial models. The right decision starts with knowing what you want it for.
Where it works well
- Answering from internal documents — policies, technical documentation, contracts. The model answers only from the text it's given and shows where the answer came from. This is called RAG.
- Classifying and extracting — working out what an email or ticket is about, pulling the amount and date out of an invoice, summarising a long report.
- Drafting — a first reply to repetitive requests, reviewed by a person before it goes out.
Where to be careful
- Long, multi-step reasoning — smaller models go wrong sooner in long chains of reasoning.
- Persian — models differ widely in Persian quality and are usually weaker than in English. Before choosing, test on real samples of your own documents, not on a leaderboard.
- The final decision — in anything that moves money, contracts or access, the model's output is a suggestion and a person approves it. See human approval in automation.
What hardware it needs
Memory is the main limit. Models are usually run compressed (quantised, for example to 4 bits), and then every billion parameters takes roughly half to two thirds of a gigabyte, plus the memory the input text needs.
| Model size | Approximate memory at 4 bits | Good for |
|---|---|---|
| 7 to 8 billion parameters | About 5 to 6 GB | Classifying, extracting, summarising |
| 13 to 14 billion parameters | About 9 to 10 GB | Answering from documents, with better quality |
| About 70 billion parameters | 40 GB and more | Harder work, on a strong GPU or several |
Small models also run on a CPU and ordinary server memory, but slowly. For interactive use by several people at once, a GPU with enough memory is what makes the difference. llama.cpp and Ollama are the usual starting points; for serving many users on GPUs, vLLM.
Local is not the same as secure
- The model should only reach documents the person asking is allowed to see. A RAG system that searches every document for everyone is a data leak in itself.
- Text from outside — an email, an attachment, a web page — can carry hidden instructions for the model. Don't give a model that reads such text permission to do anything sensitive.
- Log the questions and answers, both to improve quality and to be able to follow up.
Where to start
Pick one specific, bounded job — for example, "answer staff questions from the HR policy". Collect fifty real questions, write down their correct answers, and test two or three models on those fifty. That small test tells you more than any comparison table about whether a local model is enough for your work.