Running an AI model locally means the model runs on a computer you own, and your prompts, documents and answers are processed on that machine rather than on a provider’s servers. You need three things: an open-weight model file, free software such as Ollama to run it, and a machine with enough memory. A capable setup today is a single workstation, and the data it handles never leaves your building.
Most business owners have heard the phrase and few can picture the mechanics, which makes the whole idea sound harder than it is. AI use itself is already mainstream: Roy Morgan counts 13.6 million Australians aged 14 and over, 58 per cent of them, using AI tools in an average four weeks (Roy Morgan), and almost all of that happens in a browser connected to someone else’s cloud. We track that adoption picture in our source-verified index of Australian AI and digital indicators. The local alternative uses the same underlying technology in a different place: yours.
What does running an AI model locally actually mean?
A local AI model is a file of numbers, called weights, that encodes everything the model learned during training. Running it locally means loading that file into your computer’s memory and letting your own processor generate each answer. There is no account, no per-seat licence, and nothing is transmitted to a third party. The model even works with the network cable pulled out.
The models come from the same industry that builds the subscription chatbots. Meta’s Llama, Alibaba’s Qwen, Mistral’s models and Google’s Gemma are all released as open-weight files anyone can download and, under most of their licences, use commercially. The cloud tools remain the default way Australians meet AI: the same research found 10.5 million Australians, 45 per cent, using ChatGPT alone (Roy Morgan). A local deployment is not a rejection of that technology. It is a decision about where the technology runs, and we have set out the case for local models in detail before.
What hardware do you need to run AI models?
The one constraint that matters is memory. A model has to fit in your machine’s RAM, and it runs fastest when it fits in the memory attached to a graphics chip. Small models run on an ordinary office laptop. The mid-sized models that handle real business work run on a single high-memory workstation: a serious equipment purchase, but a workstation, not a data centre.
Model capability per gigabyte has also improved quickly, which keeps pulling serious work down onto ordinary hardware. Simon Willison, an independent researcher who has run local models since 2023, made the point plainly in his end-of-year review:
“That same laptop that could just about run a GPT-3-class model in March last year has now run multiple GPT-4 class models!”
Simon Willison, Things we learned about LLMs in 2024 (simonwillison.net)
The practical sizing rule is honest and unglamorous: buy for the model you intend to run, not the biggest one you can imagine. Open models are published in sizes from a few billion parameters up to hundreds of billions, and quantisation, a standard compression technique, lets each of them run in a fraction of the memory the raw file suggests. For most business assistants, a mid-sized model on one workstation is the sensible target.
What software actually runs a local model?
You run a local model with free, open-source software that loads the weights and gives you a chat window or an API. Ollama is the common starting point: install it, pull a model by name, and you are chatting within minutes. Tools like LM Studio wrap the same job in a desktop interface, and llama.cpp is the engine working underneath many of them.
That first chat window is genuinely a quick win, and it is worth doing just to demystify the technology. It is also where the easy part ends. A model on its own knows nothing about your business. The value for a firm comes from connecting it to your own documents through retrieval, giving staff a sanctioned way to use it, and deciding what it may and may not touch. The software that runs the model is free; the thinking that makes it useful is the actual work.
What does local AI cost compared with a subscription?
Local AI swaps recurring per-seat fees for a one-off hardware purchase and some setup work. The model weights cost nothing, the software that runs them costs nothing, and the marginal cost of a query is electricity. Whether that trade wins on dollars depends on how many seats you would otherwise license and for how long. For a five-person office it is often line-ball; for a fifty-person firm the arithmetic shifts fast.
For most of the businesses we work with, though, the deciding factor is not the subscription line item. It is where the data goes. A confidentiality-bound practice, in law, health or accounting, pays for cloud AI twice: once in fees and once in the compliance exposure of client data leaving the building. We have mapped the four levels of AI data security before, and a local deployment is the only level where the data question disappears rather than being contracted around.
Where does a DIY install stop and a business deployment start?
An afternoon with Ollama proves the concept. A tool your staff rely on every day is a different job: retrieval over your real documents, access controls, backups, monitoring, a plan for model updates, and a named owner for the system. The install is easy precisely because it carries no obligations. A business deployment is dependable precisely because someone has done the unglamorous work of adding them.
This is the same distinction we draw in automation: reliability comes from deterministic plumbing, and reasoning is added only where it earns its place. Our private AI builds exist for the second half of that job, the part between a working demo and a system a firm can trust with client records. And the honest qualifier applies as always: if a cloud tool under the right commercial terms already covers your obligations, that is what we will tell you.
Owned hardware, owned data and owned capability reinforce one another. Each one alone is a nice-to-have; together the control they give you compounds into something a subscription cannot offer: an AI capability that belongs to the business. If you want to find out what that would take for your firm, start a conversation or see how we scope and price a build.
Frequently asked questions
Can I run an AI model on a normal office computer?
Yes, within limits. Small open-weight models run on an ordinary office laptop, and they are capable enough for drafting, summarising and simple questions. The mid-sized models that approach cloud-chatbot quality need a workstation with substantial memory, ideally on a dedicated graphics chip. The model must fit in memory; that single constraint decides what a given machine can run.
Do local AI models work without an internet connection?
Yes. Once the model file is downloaded, everything runs on your own machine: the prompt, the processing and the answer. No connection is required and nothing is transmitted. That is the core privacy property of local AI, and it is why confidentiality-sensitive businesses use it for work they would never paste into a cloud chatbot.
Is local AI cheaper than paying for ChatGPT subscriptions?
Sometimes. Local AI replaces recurring per-seat fees with a one-off hardware cost plus setup, so the answer depends on headcount and time horizon. A large team on paid AI plans can reach break-even quickly; a sole trader rarely will on cost alone. Most businesses that go local do it primarily for data control, with the subscription saving as a secondary benefit.
Are open-weight models good enough for real business work?
For most day-to-day tasks, yes. Open-weight models now handle drafting, summarising, extraction and internal question-answering well, and independent researchers have documented models of GPT-4-class capability running on a single laptop. The largest cloud models still lead at the hardest reasoning tasks, so the practical approach is matching the model to the job rather than assuming either side wins everything.
Read more
SEO + GEO · 10 Sept 2026
How to Measure Your AI Search Visibility
Most businesses have no idea whether ChatGPT or Google's AI Overviews mention them. Here is how to measure AI search visibility, and when a tool is worth it.
8 MINAUTOMATION · 07 Sept 2026
Automation Bias: When Trusting the System Costs You
Automation bias is the tendency to over-trust automated output. Here is how it costs businesses, and how to design workflows that catch it early.
8 MINWEB · 03 Sept 2026
When a Custom Software Build Beats Off-the-Shelf SaaS
Most businesses should buy software, not build it. Here are the four signals that a custom build will pay for itself, and what AI has changed.
9 MINLet's compound
Tell us where growth stalls.
One team that connects your website, marketing and operations, so the results compound. Pick whichever way is easiest to start.
Australian-based · Founder on every project