What does running AI locally mean?
Running AI locally means the language model works on your own hardware: a laptop, a workstation or a server you manage. Your questions, documents and customer data stay on that machine and do not go to an external AI provider.
With cloud AI, such as ChatGPT or Claude in the browser, you send every question to the provider's servers. When you run AI locally, you download an open model and run it yourself. You are then your own provider, with the control and the maintenance that come with it.
Why do businesses run AI locally?
Businesses run AI locally for four reasons: control over data, predictable costs, availability and independence from a single provider.
Control over data. A local model processes your input on your own machine. For that processing there is no external AI provider receiving your data. In healthcare, legal services and finance that carries a lot of weight. Read more in AI and privacy.
Predictable costs. You pay for cloud AI per user or per use. A local model costs you hardware, power and maintenance, and after that no fee per question.
Availability. A local model works without an internet connection once the model files are on your machine. LM Studio describes this in its own documentation as operating entirely offline.
Independence. A cloud provider's prices, terms and model versions can change. A model you downloaded yourself keeps working the way you tested it.
What do you need to run AI locally?
To run AI locally you need three things: hardware, a model and software that runs the model.
Hardware. Memory is the deciding factor: the model has to fit in the memory of your machine or your graphics card. The makers of LM Studio list these requirements (as of October 2026):
| System | What LM Studio asks for |
|---|---|
| Mac | Apple Silicon (M1 or newer), macOS 14 or newer, 16 GB of RAM or more recommended |
| Windows | Processor with AVX2 or an ARM system (Snapdragon X Elite), at least 16 GB of RAM recommended, graphics card with at least 4 GB of dedicated memory recommended |
| Linux | x64 or ARM64, Ubuntu 20.04 or newer |
According to LM Studio it also works on a Mac with 8 GB, with smaller models and a modest context. LM Studio does not support Intel Macs. For larger models and several users at once you end up with a workstation or server with a heavy graphics card.
A model. You choose an open model that you are allowed to download and run yourself. Well-known families are Mistral from the French company Mistral AI, Llama from Meta, Gemma from Google and Qwen from Alibaba. OpenAI also offers open models with gpt-oss. Each family has its own license: read it before you use a model commercially. You can find the families side by side in our comparison of open-source LLMs for business.
Software. A program loads the model and makes it usable. Two widely used options:
- Ollama: open-source software that starts a model with a single command, for example Gemma 4. Ollama has a REST API, so your own applications and agents can call the model.
- LM Studio: a desktop program with a chat window in which you download and use models, suited to starting without a terminal.
Local or cloud: what is the difference?
The difference between local and cloud lies in where your data is processed, how you pay and who does the maintenance.
| Aspect | Local | Cloud |
|---|---|---|
| Data | Stays on your own machine | Goes to the provider |
| Costs | Hardware, power and maintenance | Per user or per use |
| Models | Open models that fit your hardware | Also the provider's largest models |
| Availability | Works without internet | Depends on internet and provider |
| Maintenance | With you or your IT partner | With the provider |
The largest models from providers such as OpenAI, Anthropic and Google run in their own cloud and cannot be downloaded. A local model is smaller. So test with your own real-world cases whether the quality is sufficient for the task you have in mind.
When is running AI locally the right choice?
Running AI locally fits your business if one of these four situations applies:
- You process sensitive data, such as medical, legal or financial files
- Your sector or your clients set requirements for where data is stored and who can access it
- You use AI so intensively every day that costs per use add up
- You want a core process to keep working if a provider changes its prices or terms
If you use AI now and then and without sensitive data, a business cloud subscription is simpler. Combining is also possible: work with sensitive data locally, the rest in the cloud. What you put in writing for cloud AI is covered in using AI under the GDPR.
How do you start running AI locally?
You start with a trial on a computer you already own and only scale up once the trial succeeds.
Step 1: install a runner. Download Ollama or LM Studio on a machine that meets the requirements.
Step 2: pick one task. Take work that comes back often and whose result is easy to judge, such as summarizing long documents or drafting replies.
Step 3: test with your own examples. Give the model ten real examples and compare the result with what you would write yourself. Try a second model if the quality disappoints.
Step 4: decide on the next step. If the model is good enough, look at whether more colleagues will work with it and whether a central machine is needed. How that works is covered in AI on your own server.
A trial on one computer costs you nothing in software: Ollama is open source and many open models are free to download. Your investment is time and, when you scale up, hardware.
Frequently asked questions
Is a local AI model worse than ChatGPT? The largest models run in the cloud and are stronger at complex reasoning tasks. For well-defined tasks such as summarizing, classifying and drafting text, an open model can be sufficient. Test that with your own examples before you choose.
What does running AI locally cost? The software and many open models are free to download. You pay for hardware, power and maintenance. A trial on an existing computer with 16 GB of RAM only costs time.
Do I need a graphics card to run AI locally? For Windows, LM Studio recommends a graphics card with at least 4 GB of dedicated memory (as of October 2026). Macs with Apple Silicon use their shared memory. Larger models need more memory.
Does the GDPR still apply if I run AI locally? Yes. Running locally removes the transfer of data to an external AI provider. For everything you do with personal data the GDPR rules still apply, such as a legal basis, security and retention periods.
Which software do you use to run AI locally? Ollama and LM Studio are two widely used options. You start Ollama from the terminal and it has a REST API for integrations. LM Studio is a desktop program with a chat window.
Further reading and help getting started
Read how AI on your own server works, look at our AI agent implementation on your own infrastructure or book a no-obligation call. If you want ongoing guidance in setting up and maintaining your AI use, look at 1-on-1 AI coaching: €399 per month excl. VAT, as an annual subscription with the first month as a trial month.
