Open-weight AI models: what they are and what they cost to run
Ali5 min read
Last week we wrote about what ChatGPT, Claude and Gemini do with what you type into them, and how the answer changes depending on which account you're on. There's a third option we skipped: open-weight AI models, where the question doesn't come up at all, because the model runs on hardware you control and nothing leaves the building.
Two years ago that option came with a serious catch. The models you could run yourself weren't good enough to bother. That's changed fast, and it's worth understanding why.
What "open-weight" actually means
A model is a file. A very large file, but a file. Open-weight means the company that trained it published that file, so anyone can download it and run it on their own hardware. Kimi K2.6, DeepSeek V4, GLM-5.2, Llama 4, Mistral and gpt-oss are all open-weight.
We say open-weight rather than open source deliberately. Several of these licences aren't OSI-approved open source: Meta's Llama licence has usage restrictions, and Moonshot's Kimi licence has a clause about competing products. The weights are available. That's not the same thing as open source, and the distinction matters if you're building something commercial on top.
A year ago they weren't close. Now they are.
Artificial Analysis is an independent evaluator that runs the same benchmark suite across every major model, open and closed. We're citing them and not the launch chart each lab publishes about its own model, because independence is the whole point here.
GLM-5.2 and Meta's Muse Spark 1.1 both score 51 on that chart, against 60 for Claude Fable 5 and 59 for GPT-5.6 Sol. A year ago that gap wasn't nine points, it was closer to thirteen, and the best open-weight model of the moment sat at 22 while the leading proprietary models were in the mid-thirties. The direction of travel is the story, not any single score.
For most business work, drafting, summarising, pulling fields out of documents, answering questions against your own records, that's now close enough not to matter. On the genuinely hard end it isn't. On Humanity's Last Exam, a set of expert-level questions across disciplines, the top three open-weight models score between 34 and 36 percent, against 44 for GPT-5.5 and 45 for Gemini 3.1 Pro Preview. On CritPt, which tests research-level physics, open-weight models land between 4 and 12 percent against 27. Anyone telling you the gap is gone entirely is selling something.
None of that capability gap matters much if the reason you're looking at open-weight models is to keep your data off someone else's servers, which is the actual reason most SMBs end up here. So it's worth being clear about what "open-weight" does and doesn't get you on that front.
Using their chat app isn't the same thing as self-hosting
Moonshot, DeepSeek and Alibaba all publish their weights and run their own chat apps and paid APIs, the same way OpenAI does. Typing into Kimi's chat app is not self-hosting. It's using a hosted product, subject to that company's privacy and training policy, with the same personal-versus-business account question we covered last week.
We read three privacy policies for that post, not thirty. The rule of thumb transfers: check what the account you're on actually does with your data before you paste anything sensitive into it. The specifics don't. You'd need to read the policy yourself.
No vendor in the loop
Download the weights, run the model on a machine you own, and there's no vendor policy to track, because there's no vendor in the loop. No account tier to check. No setting to switch off. The data never leaves.
What running one actually costs
Frontier-class open-weight models are enormous, and this is usually where the conversation stops short.
Even a compressed version of Kimi K2.6, shrunk down from the full-precision original, is a file of roughly 594GB. Running it needs four data-centre GPUs at minimum, eight to use its full context window. These are cards that cost thousands of dollars each. That's a data-centre deployment. Not a server in the back office, and not something a ten-person business stands up on a whim.
The file you can download isn't the only size available, though. Smaller open-weight models run on hardware that costs thousands rather than hundreds of thousands, and for a well-defined job (extracting data from invoices, drafting first-pass replies, classifying incoming enquiries) they're often more than enough.
Where this leaves you
Open-weight models are now close to the frontier on most everyday work, meaningfully behind on the hardest, and free to download but not free to run. For a small business, the real decision usually isn't "are they good enough" anymore. It's whether the privacy upside is worth setting up and paying for hardware, versus staying on a hosted tool and being careful about which account you use.
That's a different answer depending on your data, your volume and your budget, and it's the next post. If you'd rather talk it through directly, we offer a 30-minute call, no cost involved, and we'll tell you plainly if we're not the right fit for your business.