Why open-weight models are back on the agenda
Sound familiar?
Your AI feature works well on a proprietary API, but the bill grows with every new customer, legal is asking where the data goes, and someone has read that a free model does the same job.
An open-weight model is one whose trained model files are published, so you can run it on any infrastructure you choose. A proprietary model, such as GPT-6, Claude or Gemini, is available only through its maker's API or cloud partners.
Two things changed in 2026. Open-weight models reached a scale that makes them serious alternatives for everyday business work, and their hosted APIs became some of the cheapest ways to run AI at all. Hugging Face's summer 2026 review of open models found that, in almost every month of 2026, the largest open model from a Chinese lab was larger than any model an American lab released.
- Models built on Qwen, 2.6 times Meta's footprint
- 151,448
- Chinese open releases above 20B parameters under Apache 2.0 or MIT
- 81%
- All-time open-model downloads that go to models under 1B parameters
- 83%
The open-weight models worth shortlisting
| Model | Maker | Licence | Size (total / active parameters) | Context window | Hosted API price per 1M tokens |
|---|---|---|---|---|---|
| DeepSeek-V4-Pro | DeepSeek | MIT | 1.6T / 49B | 1M tokens | $1.32 in / $3.96 out (peak) |
| DeepSeek-V4-Flash | DeepSeek | MIT | 284B / 13B | 1M tokens | $0.30 in / $1.20 out (peak) |
| Mistral Small 4 | Mistral AI | Apache 2.0 | 119B / 6B | 256K tokens | Varies by host |
| Qwen3.6-35B-A3B | Alibaba Qwen | Apache 2.0 | 35B / 3B | 262K native, up to about 1M | Varies by host |
"Active parameters" matters more than total size for running costs. These are mixture-of-experts models: only a small part of the model works on each token, so a 284-billion-parameter model can run faster and cheaper than its size suggests. DeepSeek's off-peak API rates are half the peak rates shown.
Sources: DeepSeek pricing, DeepSeek-V4-Pro, DeepSeek-V4-Flash, Mistral Small 4, Qwen3.6-35B-A3B.
Three ways to run an open-weight model
| Option | What it means | Best for |
|---|---|---|
| Maker's hosted API | Call the model like any other API, e.g. DeepSeek's own platform | Lowest cost and effort, if the provider's data terms suit you |
| Your cloud provider | Run the model on managed GPUs in your own cloud account and region | Data residency and compliance without running servers yourself |
| Your own servers | Download the weights and serve them on GPUs you manage | Very high steady volume, strict data control, or fine-tuning |
Self-hosting needs serious hardware. Mistral lists 4x NVIDIA HGX H100, 2x HGX H200 or 1x DGX B200 as the minimum for Mistral Small 4. Only 6 billion of its parameters are active per token, but all 119 billion still have to fit in GPU memory.
When self-hosting pays off
Hosted open-weight APIs have moved the break-even point a long way. At DeepSeek-V4-Flash's off-peak rate of $0.15 per million input tokens, $5,000 a month buys over 33 billion input tokens. A self-hosted server has to beat that after paying for GPUs running around the clock, the engineers who keep them running, and capacity for your busiest hour.
IfYour volume is modest or unpredictable
ThenUse an API: a hosted open-weight model or a budget proprietary tier.
IfData must stay in a specific country or inside your own cloud account
ThenRun an open-weight model on managed GPUs in that region.
IfYou process billions of tokens a month at a steady rate
ThenPrice self-hosting against your current API bill, including engineering time.
IfYou need the model to learn your formats or terminology
ThenAn open-weight model you can fine-tune, after instructions and retrieval have been tried.
IfYour task is long, high-stakes agent work
ThenTest a top proprietary model alongside the open options. See our AI model comparison.
Want the break-even worked out for your volume?
Get a free consultationThe trade-offs
Pros
- Much lower cost per token through hosted APIs
- Run it where your data has to stay
- Fine-tune it on your own data
- No surprise deprecations: the version you test is the version you keep
- Permissive MIT and Apache 2.0 licences on many leading models
Cons
- Self-hosting needs multi-GPU servers and people to run them
- You own security updates, scaling and uptime
- The top proprietary models still tend to lead on the hardest agent tasks
- Licences vary: some large models carry custom terms
- Hosted APIs differ in where they process data
Read the licence before you build. Hugging Face found that 81% of large Chinese open releases use Apache 2.0 or MIT, while only 29% of comparable American releases do. 41% of those carry custom terms, and 30% state no licence at all.
A practical way to decide
Most teams do not need to choose one side. A common setup in 2026 is a cheap open-weight model for high-volume, simpler steps, and a proprietary model for the hard ones, chosen per request. We explain how in model routing: why one AI model is no longer enough, and the cost side in our guide to AI API pricing.
If you want help testing or hosting open-weight models, see our AI integration services and Hugging Face work.
Key takeaway
Start with a hosted open-weight API to test quality at low cost. Move to your own cloud when data rules require it, and to your own servers only when steady volume makes the maths work.
Frequently Asked Questions
Written by
Vikas Patel
AI & Engineering
