Skip to content
AI

Self-Hosted AI for Business: 3 Ways to Run AI Yourself

Self-hosted AI gives you control over data and cost, but it doesn't make you compliant on its own. Three ways to start, what each costs, and how to test first.

Kerstin Dallinger9 min read
An open steel bank vault in a dark concrete room, with a glowing blue wireframe sphere on a plinth inside - AI kept on hardware you control.

Key takeaways

  • Self-hosted AI runs on hardware you control: the computer you have, a rented GPU or a machine you own. It keeps prompts away from third-party APIs, but it doesn't make you GDPR compliant on its own.
  • Small models handle narrow, repetitive tasks grounded in your own data well. They are a poor fit for broad general knowledge.
  • Test before you buy. Rerun a workflow you already have on a rented GPU for an afternoon and compare the answers.

Most businesses that use AI send every request to someone else's server. That works until the price changes, the model starts answering differently, or a client asks where their data went. Self-hosted AI is the alternative: you run the model on hardware you control.

In the first episode of intelligence.md, the Ninja Partners series on making AI work in small and mid-sized businesses, I talked to Ali Karpuzoglu about how to do that in practice. Ali has a master's in autonomous systems, was head of development at an AI education startup where he did a lot of self-hosting, and runs his own models at home. This article covers what he showed, with his numbers from the day of recording, 8 October 2026. You can watch the full session on YouTube or listen on Spotify.

Why run AI yourself at all

Ali gave three reasons, and privacy weighed most for him.

Privacy. If the model runs on your machine, you don't have to hand client data to an external provider to get a task done. Ali's example was a law firm sitting on piles of unstructured files. A local model can pull out keywords and find the relevant documents without any of it going to OpenAI or Anthropic.

Cost risk. Cloud AI is cheap right now partly because, in Ali's words, "you're using a lot of investors' money to subsidise the cost." Nobody knows where those prices go next. One attendee put it well in the chat: the benefit of a local model is that "nobody can take it away tomorrow and make it ten times more expensive."

Predictable behaviour. Ali's team once watched the OpenAI API start "reacting differently, even though you had the version pinned." A model file on your own hardware only changes when you change it. For an automated workflow, that matters as much as the price.

Self-hosted doesn't mean compliant

Running your own model does not make a business GDPR or AI Act compliant on its own. Coming from law and compliance, this was the question I most wanted on record, and Ali's answer was short: "No, sadly not."

Self-hosting keeps the prompts off someone else's API, but that is one part of the data flow. You still have to look at what the system can reach, where its output ends up, who has access, and which rules your sector adds. Education and legal both have their own, he pointed out. The model is also just one component. If it can browse the web or run automations without guardrails, data can still leave the building. As Ali put it, compliance takes "more consideration than just saying we self-hosted so we're fine."

Three ways to run AI yourself

You can run AI on the computer you already have, on a GPU you rent by the hour, or on hardware you own. Ali's advice is to go through them in that order and only move up a level once the one below has shown you what you need.

1. The computer you already have

The easiest way to run AI locally costs nothing. WebLLM Chat downloads a model into your browser and runs it on your own machine, with nothing to install. In the session, Ali ran a seven-billion-parameter model on his MacBook and had it sort example messages using a prompt he wrote himself.

Two things to know before you try it. Small models are unreliable on general knowledge, so give them what they need (examples, categories, your own guidelines) instead of asking open questions. And a browser tool is for experimenting. For an actual business process, Ali suggests a dedicated machine running a model server. Plenty of people build setups around Mac minis. "If you have a laptop with 32 gigs of RAM, you can trial out some pretty cool stuff already."

To check whether a model fits your machine, Ali used Hugging Face, where you can enter your laptop or GPU in your account settings and see which versions of a model will run on it. Those versions are the same model stored at different precision, a technique called quantisation. In his example, a 27-billion-parameter model needed close to 60 GB at 16-bit and about 6 GB at 1-bit, though he called 1-bit quality "abysmal". On a 32 GB MacBook, the 5-bit version was the workable middle. Ollama and LM Studio are the easiest ways to start: both give you a ChatGPT-style chat window after a copy-paste install.

2. Rent the computing power

"Don't just buy hardware for hardware's sake," Ali said, and his advice is to rent first. RunPod bills GPUs by the hour, and European providers such as Scaleway and Hetzner offer similar machines. On the day of recording, an RTX Pro 6000 cost about $2 an hour in RunPod's community cloud. Buying the same card, Ali said, runs to around €20,000 before you build a computer around it.

A rented GPU is just a computer, so you decide what runs on it. On RunPod you can start a model server such as vLLM or Ollama, add a chat interface like Open WebUI, and point the tools you already use at your own server instead of OpenAI's. That gives you a self-hosted LLM your existing tools can call, a private stand-in for the ChatGPT API. For anything more complex, Ali recommends bringing in a technical person.

3. Own the hardware

Owning the hardware, the classic on-premise AI setup, used to mean a five-figure budget. It still can, but Ali pointed to cheaper routes: a used RTX 3090 for around €1,000, or AMD cards at under €2,000 each. Apple machines have a following too, because their high memory bandwidth suits language models. Ali himself runs two DGX Spark-class boxes with 256 GB of memory between them.

Timing is the catch. Hardware prices have been climbing, and Ali showed how the DGX Spark's price has only gone up since its release. "The bang for your buck will change by the week."

One point from our side, not from the session: owning only pays if the machine stays busy. A 2026 preprint measured the same GPU producing output at anywhere from $0.21 to $15.25 per million tokens, depending on how much traffic it handled. An earlier cost model found that owning makes economic sense mainly above roughly 50 million tokens a month, or where data residency rules require it. Neither paper is peer reviewed, so read them as scenarios. Both point the same way: a machine you own is only cheap if it is busy most of the time.

How to test before you commit

The test Ali recommends stays out of your day-to-day work. Take a workflow that already runs on a cloud model, such as sorting incoming messages by team, urgency and whether a person needs to look. Keep its inputs and outputs. Then run the same inputs through a local or rented model and compare. It doesn't have to run in parallel: rent a GPU for a couple of dollars, "let it go through everything at once, and then in two hours you have your results," and switch it off.

In his demo, a 27-billion-parameter Qwen model got 14 of 15 test messages right. The one miss was a hard case. Where the answers differ, improve the prompt or the guidelines, or decide where to escalate. A pattern Ali likes: let the small model handle the routine and pass anything it's unsure about to a frontier model or a person. "You still reduce your cost, because you're using the better models only in small cases."

Plenty of tasks don't need a frontier model at all. "Do I need to ask Fable for what's in my list of guidelines, which is just kind of like a search task?" Ali asked. He expects the setup around the model to matter more and more as small models catch up, and he warned against going from zero to fully automated overnight.

Check the licence

Most local AI models are open-weight rather than fully open source: you get the trained model, not the data it was trained on. Each one also comes with its own terms. "Just because you can technically run them doesn't mean you can do anything commercially," Ali said. His example was a Qwen image model released under a research licence that ruled out commercial use. Read the licence before a model goes into a business process. One attendee admitted in the chat they hadn't known there were terms to check at all.

Keeping it running

Running your own model means maintaining it. Ali's list:

  • Secure the server. Like any service you run yourself, it needs updates and access control.
  • Keep the network private. Ali connects an agent on a Hetzner server to the hardware in his home through Tailscale, a private network, so the two never talk over the open internet.
  • Update services and models. New models arrive constantly, and "the same hardware can perform significantly better six months in the future because the models just keep getting better."

And keep a person checking the output until the results show you can step back.

Where to start

For most small and mid-sized businesses, the first step is small. Pick one narrow, repetitive task, rerun its history on a rented GPU for an afternoon, and look at the gap. A small gap means you've found a candidate. A large one costs you a few dollars and tells you the cloud model is earning its fee.

If you'd rather have someone go through your workflows with you first, that's what our AI Readiness Audit is for: a senior-led diagnostic of where AI pays off in your business, what to leave alone, and what it will actually cost. New episodes of intelligence.md are listed on the series page.

Sources

Frequently asked questions

What is self-hosted AI?

Self-hosted AI means the model runs on hardware you control instead of on a provider's servers. That can be the computer you already have, a GPU you rent by the hour, or a machine you own. Most self-hosted setups use open-weight models such as Qwen, which you download and run yourself.

Is self-hosted AI automatically GDPR compliant?

No. Self-hosting keeps your prompts away from a third-party API, but compliance depends on the whole data flow: what the system can reach, where its output goes, who has access, and which rules apply in your sector. A self-hosted model that can browse the web or trigger automations without guardrails can still move data out.

How much does it cost to run AI yourself?

It ranges from nothing to five figures. A model in your browser costs nothing. On the day of the session, a high-end RTX Pro 6000 GPU cost about $2 an hour to rent on RunPod, while buying the same card was around €20,000 before the rest of the machine. A used RTX 3090 sold for around €1,000. Hardware prices are moving fast, so check current numbers before you decide.

Which tasks suit a small self-hosted model?

Narrow, repetitive tasks grounded in your own information: classifying and routing messages, flagging urgency, pulling keywords out of unstructured documents, finding the relevant file, answering from your own guidelines. Small models are weak on broad general knowledge, so they work best when you give them the facts they need.

Can I use open-weight models commercially?

Not always. Each model comes with its own licence, and some only allow research or personal use. Being able to download and run a model doesn't mean you may use it in your business, so read the licence before a model goes into a business process.

How do I test whether a local model is good enough?

Take a workflow that already runs on a cloud model, keep its inputs and outputs, and rerun the same inputs on a local or rented model. Compare the answers, fix the prompt or guidelines where they differ, and send anything the small model is unsure about to a stronger model or a person.

Kerstin Dallinger, AI Trainer & Strategist at Ninja Partners
Written by

Kerstin Dallinger

AI Trainer & Strategist, Ninja Partners

Legal Counsel by training, AI strategist by choice - the non-developer who ships real AI systems daily. Designs the websites, smart funnels, and agentic automation behind Ninja's growth - systems she scopes, builds, and runs herself.

Connect on LinkedIn

Want the next article by email? Get the Ninja Partners newsletter, one or two emails a month on AI, marketing and growth.

Start here

Not sure where your AI spend is leaking? Start with the AI Readiness Audit.

A senior-led diagnostic of your tools, workflows and data - a prioritised plan for where AI pays back, what to leave alone, and what it will actually cost. No pitch deck.