In 2026, businesses that want to run AI inside their own office — with data never leaving the organisation — have more choice than ever: dozens of open (open-weight) AI models from the US, China and Europe can be downloaded, self-hosted and used for commercial purposes. But “open” does not mean “anything goes”: each model family comes with a different licence, and the gap to GPT, Claude and Gemini varies widely from model to model. This article maps the latest open models by origin, spelling out the licence and commercial terms, architecture, strengths and independent benchmark scores against commercial models, then suggests how to choose a model by need and by the RAM of a Mac Studio / Mac mini. Every version, licence and figure is cited to a primary source, checked on 30 Sep 2026.
Quick summary
- China leads on raw capability: on the Artificial Analysis Intelligence Index, the top 3 open models are MiMo-V2.6-Pro (46), GLM-5.3 (45) and Kimi K3 (44) — all Chinese; Claude Opus 5.5 scores 58, GPT-5.6 Sol 47, Gemini 3.8 Flash 41.
- The US leads on small models with clean licences: Gemma 4 (Google), Muse Glimmer 30B (Meta), Granite 4.2 (IBM) and gpt-oss (OpenAI) are all Apache 2.0; Phi-4-reasoning-vision (Microsoft) is MIT. Llama 4 uses its own licence with conditions.
- Europe = Mistral: Mistral Small 4, Large 3, Ministral 3 and Devstral Small 2 are Apache 2.0; Mistral Medium 3.5 and Devstral 2 use a Modified MIT licence and may not be used by companies with revenue above USD 20 million per month.
- Vietnamese: on SEA-HELM, Qwen 3.5 397B (78.15) and Gemma 4 31B (77.09) are close to Claude Opus 4.8 (78.70) and GPT 6 Astra (77.86).
- Running on a Mac: a 32GB Mac mini is enough for the ~27–31 billion parameter class (Qwen3.8-27B, Gemma 4 31B); a 128–256GB Mac Studio opens up the 120–300 billion parameter class (estimate).
- Highest-scoring open models on the Artificial Analysis open-weights ranking: MiMo-V2.6-Pro 46, GLM-5.3 45, Kimi K3 44 (Intelligence Index v4.3.2, checked 30 Sep 2026).
- On the same index: Claude Opus 5.5 scores 58, GPT-5.6 Sol (max) scores 47, Gemini 3.8 Flash (high) scores 41.
- SEA-HELM Vietnamese (18 Sep 2026): Opus 4.8 78.70 · Qwen 3.5 397B 78.15 · GPT 6 Astra 77.86 · Gemma 4 31B 77.09.
- Mistral Medium 3.5 uses Modified MIT: it may not be used if the company’s consolidated revenue exceeds USD 20 million per month — according to the LICENSE file on Hugging Face.
- Kimi K3: 2.8 trillion parameters, 104 billion active; a separate agreement is required to run a model service (MaaS) with revenue above USD 20 million per 12 months — according to the Kimi K3 License.
What does a “commercially usable” open AI model mean?
An open AI model (open-weight) is one whose developer publishes the weights so anyone can download it and run it on their own machine; “commercially usable” means the accompanying licence allows the model to be built into products, services or business processes. The two are separate: weights can be freely downloadable while the licence still prohibits or restricts commercial use. So the first question when choosing a model for a business is not “which model is the strongest” but “what does this model’s licence let us do”.
Among the models Namtech reviewed for this article, licences fall into three groups. Group 1 — permissive licences (Apache 2.0, MIT, and NVIDIA’s OpenMDW): commercial use with no revenue or user threshold, you only need to keep the copyright notice when redistributing. For example, Google writes in the Gemma 4 launch post that the model is released under the Apache 2.0 licence, which permits commercial use. Group 2 — custom licences with conditions: commercial use is allowed but with user thresholds, revenue thresholds or labelling obligations — Llama 4, Kimi K3, Qwen3.8-Max, GLM-5.3, MiniMax M3, Mistral Medium 3.5. Group 3 — no commercial use: for example, the MiniMax-M2.7 licence states that any form of commercial use is prohibited without separate written permission from MiniMax. Group 3 is excluded from this article.
A note on terminology: many articles call these models “open source”. The more accurate term is “open model” or “open-weight”, because what is opened is mainly the weights, while training data and training code are usually not published. Namtech has analysed the licensing question from a deployment angle in detail in our article on choosing open-source models for internal AI; this article focuses on mapping the newest models as of the end of September 2026.
| Origin · Vendor | Model (release) | Licence | Architecture | Main strengths |
|---|---|---|---|---|
| US · OpenAI | gpt-oss-120b / 20b (5 Aug 2025) | Apache 2.0 + gpt-oss usage policy | MoE 117B (5.1B active) / 21B (3.6B); 131K context | Adjustable reasoning, tool calling; the 20b version runs in 16GB |
| US · Google | Gemma 4 E2B/E4B/26B A4B/31B (2 Apr 2026), 12B (3 Jun 2026) | Apache 2.0 | 31B dense; 26B MoE (3.8B active); context up to 256K | Multilingual (more than 140 languages), multimodal, function calling; strong Vietnamese on SEA-HELM |
| US · Meta | Muse Glimmer 30B (Aug 2026) | Apache 2.0 (+ usage policy) | Dense ~29.6B, accepts images; 131K+ context | Agents running on personal machines, tool calling |
| US · Meta | Llama 4 Scout / Maverick (5 Apr 2025) | Llama 4 Community License (conditional) | MoE 109B / 400B, 17B active | Multimodal, very long context (Scout 10M) |
| US · Microsoft | Phi-4-reasoning-vision-15B (4 Mar 2026) | MIT | Dense 15B; 16K context | Maths and science reasoning over images; computer-use tasks |
| US · IBM | Granite 4.2 3B/8B/30B (25 Aug 2026) | Apache 2.0 | Dense 30B; 128K context (extendable to 512K) | Toggleable reasoning, tool calling, enterprise-oriented |
| China · Alibaba | Qwen3.8-27B (Aug 2026) | Apache 2.0 | Dense 27B, accepts images; 262K context (extendable to 1M) | Agentic coding, computer use; strongest in the class that fits a small machine |
| China · Alibaba | Qwen3.8-2.4T-A95B (Aug 2026) | Qwen3.8-Max License (USD 50 million/12-month threshold for MaaS) | MoE 2.4T (95B active) | Qwen’s largest open model |
| China · DeepSeek | DeepSeek-V4.1-Flash (10 Sep 2026) | MIT | MoE 552B, multimodal; 1M context | Coding agents, terminal; on par with commercial models on vendor-published benchmarks |
| China · Z.ai | GLM-5.3 (weights 27 Aug 2026) · GLM-5.3-Flash (Aug 2026) | GLM-5.3 License (only gated at USD 10 billion revenue) · Flash: MIT | MoE ~753B · Flash 320B (18B active) | Coding, agents; Flash is the more compact multimodal version |
| China · Moonshot | Kimi K3 (weights 27 Jul 2026) | Kimi K3 License (USD 20 million/12-month threshold for MaaS) | MoE 2.8T (104B active); 1M context | Reasoning, web browsing, computer use |
| China · Xiaomi | MiMo-V2.6-Pro-RL (Sep 2026) | MIT (stated on the model card) | MoE 1.02T (42B active); 1M context | Highest score among open models on Artificial Analysis |
| China · MiniMax | MiniMax M3 (Jun 2026) | MiniMax Community License (labelling, USD 20 million/year threshold) | MoE ~428B (~23B active); 1M context | Web browsing, agents |
| Europe · Mistral (France) | Mistral Small 4 (16 Mar 2026) | Apache 2.0 | MoE 119B (6.5B active); 256K context | Combines chat, reasoning, coding and multimodal; Vietnamese is in its language list |
| Europe · Mistral | Mistral Large 3, Ministral 3 (2 Dec 2025); Devstral Small 2 (9 Dec 2025) | Apache 2.0 | Large 3: MoE 675B (41B active); Ministral 3B/8B/14B; Devstral 24B | Large: multimodal; Ministral: small machines; Devstral: coding agents |
| Europe · Mistral | Mistral Medium 3.5 (22 May 2026) | Modified MIT (not usable if revenue > USD 20 million/month) | Dense 128B; 256K context | Agentic coding (77.6% SWE-Bench Verified, vendor-published) |
| Europe · EU | EuroLLM-22B-Instruct-2512 (Dec 2025) | Apache 2.0 | Dense ~22.6B; 32K context | European languages, translation (no Vietnamese) |
Release dates are taken from the official vendor blog or model card; for Kimi K3 and GLM-5.3 they are the date of the first weights commit on Hugging Face. A “MoE” (mixture-of-experts) architecture means each token activates only part of the parameters; “dense” means all of them are activated.
The US group: small models, clean licences
What the US group has in common in 2026 is clean licences and moderate size: most are Apache 2.0 or MIT, and most are small enough to run on a workstation or a Mac. The trade-off is that on independent leaderboards, the best US open model sits well behind the leading Chinese group. One notable detail: apart from gpt-oss and Llama 4, the model cards of Gemma 4, Muse Glimmer, Phi and Granite only compare against other open models, not directly against GPT, Claude or Gemini.
OpenAI gpt-oss-120b and gpt-oss-20b
gpt-oss is still OpenAI’s newest open chat model as of the date checked. According to the gpt-oss paper on arXiv, both models are released under the Apache 2.0 licence together with the gpt-oss usage policy; the model card is dated 5 Aug 2025. Both are MoE: the 120b version has 117 billion parameters with 5.1 billion active, and the 20b version has 21 billion with 3.6 billion active. The gpt-oss-120b Hugging Face page states that the 120b version fits on a single 80GB GPU and the 20b version runs within 16GB of memory, with adjustable reasoning effort, function calling, web browsing and Python code execution.
On comparisons with commercial models, OpenAI itself reports that gpt-oss-120b beats o3-mini and approaches o4-mini on standard benchmarks — these are 2025-generation commercial models. On an independent scale, Artificial Analysis scores gpt-oss-120b (high) at 12 points, and on Arena, gpt-oss-120b ranks 208th with 1,352 points. gpt-oss’s real strength today is being light and fast: the 20b version runs on a 16GB Mac mini, suitable for classification, extraction and simple internal assistant tasks.
Google Gemma 4
Gemma 4 is Google’s strongest open model family, released under the Apache 2.0 licence. The launch post dated 2 Apr 2026 introduced the E2B, E4B, 26B MoE and 31B dense versions; the 12B version came later, on 3 Jun 2026. According to the Gemma 4 model card, the 26B A4B version has 25.2 billion parameters in total but only 3.8 billion active, a 256K context, support for more than 140 languages, a thinking mode and native function calling. The model card reports the 31B version at 84.3% on GPQA Diamond and 80.0% on LiveCodeBench v6 — but compares only against Gemma 3.
Google said that at launch, the 31B version ranked 3rd among open models on the Arena leaderboard. Checked on 25 Sep 2026, Gemma 4 31B had 1,453 points on Arena, 73rd overall; Artificial Analysis scores it at 19 points. Gemma 4’s standout feature for Vietnamese businesses is Vietnamese: on SEA-HELM, Gemma 4 31B scores 77.09 — only about 1.6 points behind Claude Opus 4.8, despite being far smaller. This is why Gemma 4 is often the first candidate Namtech tests for a Vietnamese-language internal assistant on a mid-size machine.
Meta: Muse Glimmer 30B and Llama 4
Meta has two very different tracks. The newest is Muse Glimmer 30B, released by Meta Superintelligence Lab in August 2026 under Apache 2.0 — no user threshold, no “Built with Llama” requirement. The Muse Glimmer model card states that the model is intended for both commercial and research use, is dense with about 29.6 billion parameters plus an image encoder, is optimised for autonomous agent tasks on user hardware, and that the 4-bit quantised version runs in about 24GB or 32GB. The model card only compares against Gemma 4 and Qwen; Artificial Analysis scores Muse Glimmer (high) at 17 points, and SEA-HELM Vietnamese at 73.52.
Llama 4 (Scout and Maverick, 5 Apr 2025) is MoE with 17 billion active parameters, 109 billion and 400 billion in total. Meta itself reports in the Llama 4 announcement that Maverick beats GPT-4o and Gemini 2.0 Flash on many popular benchmarks — compared with 2024–2025 commercial models. The licence is the part to read carefully: the Llama 4 Community License requires companies with more than 700 million monthly active users on the release date to request a separate licence, mandates displaying “Built with Llama” and putting “Llama” at the start of derivative model names; and the acceptable use policy does not grant rights to the Llama 4 multimodal models to individuals domiciled in, or companies with their principal place of business in, the EU (end users of products are not affected). For Vietnamese businesses, the 700 million threshold is practically out of reach, but the labelling obligations still apply.
Microsoft Phi and IBM Granite
The newest Phi model with a language core is Phi-4-reasoning-vision-15B, released on 4 Mar 2026 under the MIT licence. The model card lists 15 billion parameters and a 16,384-token context, targeting maths and science reasoning over image data and computer-use tasks; the model card only compares against open models. A 16K context is fairly short for Q&A over long documents, so Phi is better suited to specialised tasks such as reading charts and diagrams.
IBM Granite 4.2 (3B, 8B, 30B) was released on 25 Aug 2026. IBM’s blog on Hugging Face confirms that all of Granite 4.2 is released under Apache 2.0; the 30B model card describes a dense architecture with reasoning, a 128K context extendable to 512K, and tool calling with reasoning. IBM does not compare against any other model in the model card; Artificial Analysis scores Granite 4.2 30B at 15 points. Granite suits businesses that want a “traditional” vendor with a clear licence more than the highest score.
The Chinese group: leading on open-model capability
The Chinese group holds the top spots among open models. Artificial Analysis writes on its open-model ranking page that MiMo-V2.6-Pro and GLM-5.3 (max) are the two most intelligent open models, followed by Kimi K3 (max) and GLM-5.3-Flash. What they share: most are MoE with hundreds of billions to trillions of parameters, a 1 million token context, and model cards that compare directly with Claude and GPT. One caveat: no model card or licence in this group mentions Vietnamese; evidence on Vietnamese comes only from the independent SEA-HELM leaderboard.
Qwen (Alibaba): Qwen3.8
Qwen3.8-27B is a dense 27 billion parameter model that accepts images, with a 262,144-token context (extendable to 1 million) and an unconditional Apache 2.0 licence. According to the Qwen3.8-27B model card, Qwen reports 61.7 on SWE-bench Pro versus 53.4 for Claude Opus 4.6 Max, and 84.3 on OSWorld-Verified versus 72.7; but on GPQA Diamond (89.2 vs 91.3) and Terminal Bench 2.1 (73.0 vs 78.2) it is still behind. Artificial Analysis scores Qwen3.8 27B (xhigh) at 34 points — the highest among open models of this size that Namtech could find, nearly double Gemma 4 31B.
The two larger versions have very different licences. Qwen3.8-2.4T-A95B (2.4 trillion parameters, 95 billion active) uses the Qwen3.8-Max License: businesses running a model service (MaaS) or an AI work assistant with consolidated revenue above USD 50 million over 12 consecutive months must request a separate licence. Qwen3.8-Flash-Next is even stricter: the Qwen Community License 1.0 requires every business running MaaS or an AI work assistant to request a separate licence before commercial use — with no revenue floor. In other words, within the same Qwen family, the 27B version is free to use, while Flash-Next is not, if your business model is selling AI services.
For Vietnamese, the previous generation Qwen3.5 is still worth noting: the Qwen3.5-397B-A17B model card states that support was expanded to 201 languages and dialects, and on SEA-HELM Vietnamese the 397B version scores 78.15 — the highest among open models on the leaderboard. Qwen3.6-35B-A3B (35 billion total, 3 billion active, Apache 2.0) and Qwen 3.6 27B (76.12 on SEA-HELM) are compact options for mid-size machines.
DeepSeek V4.1-Flash and V4-Pro
DeepSeek releases all its weights under the MIT licence. The newest is DeepSeek-V4.1-Flash: the DeepSeek change log records the official release of this model on 10 Sep 2026, and notes that the earlier V4 Flash generation has been retired from the API (the open weights remain available). According to the V4.1-Flash model card, it is a multimodal MoE with 552 billion core parameters and a context of up to 1 million tokens. The vendor’s comparison table at maximum reasoning effort shows V4.1-Flash at 90.6 on Terminal-Bench 2.1, slightly ahead of Opus 5.0 (89.1) and GPT-5.6 Sol (88.8); but on GPQA Diamond it scores 90.9 versus 93.4 and 94.1. Artificial Analysis scores V4.1 Flash at 39 points; on Arena, it ranks 29th.
The large DeepSeek-V4-Pro (1.6 trillion parameters, 49 billion active) has an official 0813 release replacing the preview. For businesses self-hosting, the more interesting version is the first-generation DeepSeek-V4-Flash: 284 billion parameters, 13 billion active, MIT — according to the V4-Flash model card. At this size it is one of the few “near-frontier” open models that fits on a single 256GB Mac Studio, based on the capacity estimate below.
GLM (Z.ai): GLM-5.3 and GLM-5.3-Flash
The GLM-5.3 model card states clearly that GLM-5.3 uses the same base model as GLM-5.2 — all improvements come from post-training; the weights went up on Hugging Face on 27 Aug 2026. Artificial Analysis scores GLM-5.3 (max) at 45 points, 2nd among open models. In the vendor’s table, GLM-5.3 scores 66.9 on DeepSWE v1.1, higher than Opus 4.8 (58.0) but lower than Fable 5 (69.7) and GPT-5.6 Sol (72.7). The “GLM-5.3 License” sounds strict but is in fact very broad: its terms only require businesses running MaaS with revenue above USD 10 billion per 12 months to pass a Z.ai security review.
The compact GLM-5.3-Flash uses MIT, with 320 billion parameters in total and only 18 billion active, and is the first natively multimodal model in the GLM-5 line according to the GLM-5.3-Flash model card; Artificial Analysis scores it at 42 points, 4th among open models. The previous GLM-5.2 is also MIT; on SEA-HELM Vietnamese, GLM 5.2 scores 76.93. Namtech published a dedicated article on GLM-5.2 when it launched.
Kimi K3 (Moonshot AI)
Kimi K3 is the open model with the most parameters in this article: according to the Kimi K3 model card, it is an MoE with 2.8 trillion parameters in total, 104 billion active, a 1,048,576-token context, accepting both text and images. Moonshot reports K3 at 93.5 on GPQA Diamond, versus 92.6 for Fable 5, 94.1 for GPT-5.6 Sol and 91.0 for Claude Opus 4.8. On independent scales, Artificial Analysis scores Kimi K3 (max) at 44 points, and on Arena K3 ranks 16th — the highest among open-weight models Namtech found on the leaderboard on 25 Sep 2026.
The Kimi K3 License allows commercial use but with a threshold: businesses operating a model service (MaaS) with revenue above USD 20 million per 12 months must sign a separate agreement with Moonshot. The licence also waives these requirements for internal use — exactly the internal AI scenario for businesses. In practice, at 2.8 trillion parameters, K3 far exceeds the capacity of one or a few Macs; Namtech analysed this in detail in our article on Kimi K3.
MiniMax M3, Xiaomi MiMo and other models
MiniMax M3 (about 428 billion parameters, 23 billion active, 1M context) allows commercial use with conditions: the MiniMax Community License requires displaying “Built with MiniMax M3”, written permission if the product has revenue above USD 20 million per year (below that, a one-time notice is enough), plus a list of prohibited uses. Artificial Analysis scores M3 at 29 points. Do not confuse it with MiniMax-M2.7 — an older version that prohibits commercial use.
Xiaomi MiMo-V2.6-Pro currently tops the open group on Artificial Analysis with 46 points. The MiMo-V2.6-Pro-RL model card describes an MoE with 1.02 trillion parameters, 42 billion active, a 1M context, accepting text, images, audio and video; the licence is listed as MIT in the model card metadata. Note: at the time checked, the repo had no separate LICENSE file — businesses should keep a snapshot of the model card as evidence. Other Chinese models such as Tencent Hy4-preview (Apache 2.0) or Meituan LongCat-2.0 (MIT) also have open weights, but do not yet have enough independent scores to compare in this article.
The European group: Mistral is almost the only contender
In Europe, commercially usable and competitive open models currently come almost exclusively from Mistral AI (France); other projects mainly serve European languages or research. The most important thing with Mistral is to distinguish two licence groups: Apache 2.0 (free to use) and Modified MIT (blocked for companies with large revenue).
Mistral Apache 2.0: Small 4, Large 3, Ministral 3, Devstral Small 2
Mistral Small 4 was released on 16 Mar 2026; Mistral’s announcement states that the model is released under Apache 2.0. The model card describes an MoE with 128 experts (4 active), 119 billion parameters with 6.5 billion active per token, a 256K context, accepting text and images; a single model combining chat, reasoning (formerly Magistral) and coding (Devstral) modes. Vietnamese is in the model card’s language list. Artificial Analysis scores Small 4 (Reasoning) at 11 points. Namtech has a dedicated article on Mistral Small 4.
Mistral Large 3 and Ministral 3 were released on the same day, 2 Dec 2025. The Mistral 3 announcement states that Large 3 is a sparse MoE with 41 billion active parameters, 675 billion in total, and that all models are Apache 2.0; Ministral 3 comes in 3B, 8B and 14B sizes for small machines. Devstral Small 2 (24B, 9 Dec 2025) is an agentic coding model; the model card says it is light enough to run on a 32GB Mac. According to the table in the Devstral Small 2 model card itself, the model scores 68.0% on SWE-bench Verified, versus 77.2% for Claude Sonnet 4.5 and 76.2% for Gemini 3 Pro (competitor figures taken from public announcements) — about 8–9 points behind but able to run on-site.
Mistral Modified MIT: Medium 3.5 and Devstral 2
Mistral Medium 3.5 (22 May 2026) is a dense 128 billion parameter model with a 256K context. Mistral reports that Medium 3.5 scores 77.6% on SWE-Bench Verified, ranking above Devstral 2 and Qwen3.5 397B A17B. On SEA-HELM Vietnamese, Medium 3.5 scores 68.82 — significantly lower than the Qwen, Gemma and GLM group. Artificial Analysis scores it at 14 points; Arena ranks it 113th.
The crux is the licence. Medium 3.5’s LICENSE file is Modified MIT, which states that you may not exercise any rights under the licence if the worldwide consolidated monthly revenue of your company (or the company you work for) exceeded USD 20 million in the previous month. Devstral 2 (123B) uses the same terms. For most Vietnamese SMEs this threshold is out of reach; but large groups or subsidiaries of multinational groups must check consolidated revenue before use.
Other European models
EuroLLM-22B-Instruct-2512 — an EU-funded project — uses Apache 2.0, is dense with about 22.6 billion parameters and a 32K context, and focuses on European languages and translation; the language list in the model card does not include Vietnamese. Switzerland’s Apertus 1.5 (not in the EU) appears on SEA-HELM Vietnamese with 64.66 points for the 70B version, but Namtech has not been able to read this version’s licence file directly, so it is not on the recommended list. Some other models such as ALIA (Spain) or Velvet (Italy) carry notes restricting intended use in their model cards, while Aleph Alpha uses a licence for non-commercial research only — so Namtech does not include them in the recommended list.
Benchmarks: how do open models compare with GPT, Claude and Gemini?
Short answer: the best open models already match or beat some commercial models on sale, but still sit a clear distance behind the top commercial model. For a fair comparison, Namtech prioritises three independent leaderboards: Artificial Analysis Intelligence Index v4.3.2 (combining 10 evaluations, run under uniform conditions), the Arena leaderboard (users vote for the better answer — it measures preference, not correctness), and SEA-HELM for Vietnamese.
| Model | Type | Artificial Analysis Index | Arena (score · rank) |
|---|---|---|---|
| Claude Opus 5.5 (Anthropic) | Commercial | 58 | 1,509 · rank 1 |
| GPT-5.6 Sol (OpenAI) | Commercial | 47 | 1,483 · rank 19 |
| Gemini 3.8 Flash (Google) | Commercial | 41 | 1,492 · rank 10 (preliminary) |
| MiMo-V2.6-Pro (Xiaomi) | Open · MIT | 46 | 1,480 · rank 23 |
| GLM-5.3 (Z.ai) | Open · GLM-5.3 License | 45 | 1,480 · rank 24 |
| Kimi K3 (Moonshot) | Open · Kimi K3 License | 44 | 1,488 · rank 16 |
| DeepSeek V4.1 Flash | Open · MIT | 39 | 1,477 · rank 29 |
| Qwen3.8 27B (Alibaba) | Open · Apache 2.0 | 34 | 1,438 · rank 95 |
| Gemma 4 31B (Google) | Open · Apache 2.0 | 19 | 1,453 · rank 73 |
| Muse Glimmer (Meta) | Open · Apache 2.0 | 17 | 1,424 |
| Granite 4.2 30B (IBM) | Open · Apache 2.0 | 15 | — |
| Mistral Medium 3.5 | Open · Modified MIT | 14 | 1,426 · rank 113 |
| gpt-oss-120b (OpenAI) | Open · Apache 2.0 | 12 | 1,352 · rank 208 |
| Mistral Small 4 | Open · Apache 2.0 | 11 | — |
Artificial Analysis scores use the strongest configuration shown on each model’s page (for example “max”, “high”, “Reasoning”); pages may not be updated at the same time. Licence labels in the table are Namtech’s, based on the original LICENSE file, not Arena’s labels.
Table 2 shows three things. First, the gap at the top remains: Claude Opus 5.5 leads the best open model by 12 points on Artificial Analysis. Second, the middle tier already overlaps: MiMo-V2.6-Pro, GLM-5.3 and Kimi K3 score higher than Gemini 3.8 Flash, and MiMo is only one point behind GPT-5.6 Sol. Third, there is a “chasm” between strong models and models that run on small machines: models under 35 billion parameters only reach 17–34 points, and Qwen3.8-27B alone stands well apart from the rest of this group.
| Open model | Benchmark | Open model score | Commercial model score (same table) |
|---|---|---|---|
| DeepSeek-V4.1-Flash | Terminal-Bench 2.1 | 90.6 | Opus-5.0: 89.1 · GPT-5.6 Sol: 88.8 |
| DeepSeek-V4.1-Flash | GPQA Diamond | 90.9 | Opus-5.0: 93.4 · GPT-5.6 Sol: 94.1 |
| Kimi K3 | GPQA Diamond | 93.5 | Fable 5: 92.6 · GPT-5.6 Sol: 94.1 · Opus 4.8: 91.0 |
| GLM-5.3 | DeepSWE v1.1 | 66.9 | Opus 4.8: 58.0 · Fable 5: 69.7 · GPT-5.6 Sol: 72.7 |
| Qwen3.8-27B | SWE-bench Pro | 61.7 | Opus 4.6 Max: 53.4 |
| Qwen3.8-27B | Terminal Bench 2.1 | 73.0 | Opus 4.6 Max: 78.2 |
| Devstral Small 2 (24B) | SWE-bench Verified | 68.0% | Claude Sonnet 4.5: 77.2% · Gemini 3 Pro: 76.2% |
| gpt-oss-120b | Aggregate of standard benchmarks | OpenAI: beats o3-mini, approaches o4-mini | |
| Llama 4 Maverick | Many popular benchmarks | Meta: beats GPT-4o and Gemini 2.0 Flash | |
Vendor-reported figures are useful for seeing where a model is strong, but should be read with three caveats. Each vendor chooses the benchmarks, reasoning effort and harness that favour it; many tables compare against previous-generation commercial models (Opus 4.6, GPT-4o) rather than the latest; and for the same model, scores can differ between tables — for example DeepSeek reports different NL2Repo figures in the model card and the change log. So Namtech uses Table 2 for ranking and Table 3 only to understand each model’s strengths, and always re-measures on the client’s real data before deciding.
Which open models are good at Vietnamese?
According to the SEA-HELM Vietnamese leaderboard — AI Singapore’s independent evaluation, updated 18 Sep 2026 — several open models are already close to the best commercial models in Vietnamese. This is a big difference from the English aggregate leaderboards: the gap in Vietnamese is much narrower, and small models such as Gemma 4 31B sit very close to the commercial models.
| Model | Type | Vietnamese score |
|---|---|---|
| Claude Opus 4.8 | Commercial | 78.70 |
| Gemini 3.1 Pro Preview | Commercial | 78.66 |
| Qwen 3.5 397B MoE | Open · Apache 2.0 | 78.15 |
| GPT 6 Astra | Commercial | 77.86 |
| Gemma 4 31B | Open · Apache 2.0 | 77.09 |
| GLM 5.2 754B MoE (FP8) | Open · MIT | 76.93 |
| Qwen 3.6 27B | Open · Apache 2.0 | 76.12 |
| Muse Glimmer 30B | Open · Apache 2.0 | 73.52 |
| DeepSeek V4 Pro 1600B MoE | Open · MIT | 73.40 |
| Mistral Medium 3.5 128B | Open · Modified MIT | 68.82 |
| Apertus 1.5 70B | Open | 64.66 |
With a confidence interval of about ±1.5 points, the top four models — Opus 4.8, Gemini 3.1 Pro Preview, Qwen 3.5 397B and GPT 6 Astra — are practically tied in Vietnamese. Gemma 4 31B and Qwen 3.6 27B, two models small enough to run on a Mac mini or an entry-level Mac Studio, are only about 1.5–2.6 points behind the leaders. By contrast, Mistral Medium 3.5 and Apertus are clearly weaker, so they are not a priority choice for Vietnamese applications. Note: SEA-HELM does not yet have scores for some new models such as Qwen3.8, Kimi K3 or DeepSeek V4.1-Flash, so for these models you still need to run your own tests.
Choosing an open model by Mac Studio and Mac mini RAM
On a Mac, the deciding factor for which model will run is unified memory capacity, because all the weights — including the inactive experts of an MoE — must sit in RAM. The table below uses the same estimation formula as our article on the real cost of running an LLM on-premise and Namtech’s product pages: usable memory ≈ 75% of RAM, a 4-bit model ≈ 0.5 byte/parameter, with 20% reserved for context. This is a theoretical capacity ceiling, not a speed measurement.
| Machine · RAM | Maximum model size (4-bit, estimate) | Commercially usable candidates |
|---|---|---|
| Mac mini M6 16GB | ~19 billion parameters | Gemma 4 12B, Phi-4-reasoning-vision-15B, Ministral 3 14B; gpt-oss-20b (OpenAI says it runs in 16GB, close to the ceiling) |
| Mac mini 24GB | ~28 billion | Gemma 4 26B A4B, Qwen3.8-27B, Qwen 3.6 27B, Devstral Small 2 24B |
| Mac mini 32GB | ~38 billion | Gemma 4 31B, Muse Glimmer 30B, Granite 4.2 30B, Qwen3.6-35B-A3B |
| Mac mini M5 Pro 48–64GB / Mac Studio M5 Max 48–64GB | ~57–76 billion | The group above with longer context, running 2 models in parallel (for example one chat model + one coding model) |
| Mac Studio M5 Ultra 96GB | ~115 billion | The 30B group with many users; gpt-oss-120b (117B) and Mistral Small 4 (119B) are close to the ceiling — test in practice |
| Mac Studio M5 Max 128GB | ~153 billion | gpt-oss-120b, Mistral Small 4, Mistral Medium 3.5 128B |
| Mac Studio M5 Ultra 256GB | ~307 billion | DeepSeek-V4-Flash 284B; GLM-5.3-Flash (320B) slightly exceeds the ceiling |
| Cluster of 2–4 Mac Studio 256GB | ~614 – 1,228 billion | DeepSeek-V4.1-Flash 552B, GLM-5.2/5.3 (~753B), MiniMax M3 (~428B) |
Estimate, not a benchmark: RAM × 0.75 × 0.8 ÷ 0.5 byte/parameter. MoE models such as Gemma 4 26B A4B or Qwen3.6-35B-A3B take up RAM according to total parameters but run fast according to active parameters. Kimi K3 (2.8T), Qwen3.8-2.4T and DeepSeek-V4-Pro (1.6T) exceed the capacity of the configurations above.
Looking at Table 5, the “sweet spot” for most businesses is the 24–32GB group: Qwen3.8-27B (the highest independent score in the small group) and Gemma 4 31B (the best Vietnamese in the small group) are both Apache 2.0 and fit on one Mac mini. A 96–128GB Mac Studio is worth the investment when you need to serve many users at once, run several models in parallel, or use a 120 billion parameter model. “Near-frontier” models of a few hundred billion parameters and up need a 256GB Mac Studio or a cluster — details in our article on clustering Mac Studios for AI.
Choosing a model by need
There is no best model for everything; the practical approach is to identify the main need, filter by licence, then choose the smallest model that is good enough and measure it on real data. Based on the sources above, Namtech suggests the following starting points.
Internal assistant and Q&A over Vietnamese documents: start with Gemma 4 31B or Qwen 3.6 27B — both Apache 2.0, above 76 points on SEA-HELM Vietnamese, and fit on a 32GB Mac mini. If you have a large Mac Studio, Qwen 3.5 397B gives the highest Vietnamese score among open models. Coding and agents: Qwen3.8-27B for small machines (61.7 on SWE-bench Pro, vendor-reported); Devstral Small 2 if you want a European code-specialist model; DeepSeek-V4.1-Flash or GLM-5.3 if you have a cluster and need capability close to commercial models. Reading images, charts and interfaces: Gemma 4, Muse Glimmer and Mistral Small 4 all accept images; Phi-4-reasoning-vision specialises in reasoning over images. Very long context: DeepSeek, GLM, Kimi and MiMo all announce 1 million tokens, and Qwen3.8-27B extends to 1 million — but long context uses a lot of RAM for the cache, which must be factored into the machine configuration.
Prioritising the cleanest licences: if your legal team wants to avoid any custom terms, limit the list to Apache 2.0 and MIT: Gemma 4, Muse Glimmer, Granite 4.2, gpt-oss, Phi, Qwen3.8-27B, Qwen3.5/3.6, DeepSeek, GLM-5.2, GLM-5.3-Flash, Mistral Small 4, Large 3, Ministral 3, Devstral Small 2. This list already covers almost every internal need at Mac scale. Before loading any model, you should also check the provenance of the repo you download from — see our article on open model provenance.
Legal notes on open model licences
This section summarises terms read from the original licence files so businesses know what to ask — it is not legal advice; specific contracts and business models should be reviewed by a lawyer. There are five common types of terms. User thresholds: Llama 4 (700 million monthly active users, counted on the release date). Revenue thresholds: Mistral Medium 3.5 and Devstral 2 (USD 20 million/month, applying to every purpose), Kimi K3 (USD 20 million/12 months, only for MaaS), Qwen3.8-Max (USD 50 million/12 months, MaaS or AI work assistants), GLM-5.3 (USD 10 billion, only requires a security review), MiniMax M3 (USD 20 million/year). No floor: Qwen3.8-Flash-Next requires every business running MaaS or an AI work assistant to request a separate licence.
Labelling obligations: “Built with Llama” and the “Llama” prefix for derivative models; “Built with MiniMax M3”; Qwen3.8-Max and Kimi require displaying the model name when the product exceeds 100 million monthly active users or USD 20 million in monthly revenue. Regional and purpose restrictions: Llama 4’s usage policy does not grant rights to the multimodal models to companies with their principal place of business in the EU; MiniMax M3 has a list of prohibited uses, including military purposes; gpt-oss and Muse Glimmer come with their own usage policies even though the licence is Apache 2.0.
For the internal AI scenario — a business running a model itself to serve its own staff, not selling AI services externally — most MaaS thresholds do not apply, and the Kimi K3 License even explicitly waives its requirements for internal use. The biggest real-world risk is picking the wrong version: the same family with different licences (Qwen3.8-27B vs Qwen3.8-Flash-Next; MiniMax M3 vs M2.7; Mistral Small 4 vs Medium 3.5). Read the LICENSE file of the exact original repo you download from, keep a copy with the download date, and record it in your deployment file.
How does Namtech deploy open models on Macs?
Namtech deploys internal AI on the principle of measure first, decide later: pick 2–3 candidate models with suitable licences, test them on the client’s real documents and questions, measure Vietnamese quality, speed and RAM usage, and only then settle on the model and machine configuration. The model runs 100% on a machine located at the client’s office or at Namtech, and data is not sent to any external AI service.
Businesses have two ways to start. Buy a machine: see configurations and prices for Mac Studio and Mac mini. Or rent a machine if you are not ready to invest: Namtech’s Mac Studio and Mac mini rental for AI service runs on a 12-month term, the machine is hosted at Namtech, the business accesses the model via API or VPN and Namtech handles administration. This is a good way to test a model such as Qwen3.8-27B or Gemma 4 31B on real data before deciding to buy.
As of 30 Sep 2026, commercially usable open AI models are strong enough for most internal needs: China leads on capability, the US leads on small models with clean licences, Europe has Mistral — and for Vietnamese, Gemma 4 31B or Qwen 3.6 27B running on a single Mac mini are already close to GPT, Claude and Gemini, as long as you pick the right licence and measure on real data.
Frequently asked questions
Is an open (open-weight) AI model the same as open source?
Not exactly. Open-weight means the model weights are published for download and self-hosting. Whether it can be used commercially depends on the accompanying licence: Apache 2.0 and MIT impose almost no restrictions, while custom licences such as the Llama 4 Community License, Kimi K3 License, Qwen3.8-Max License or Mistral’s Modified MIT have conditions on user thresholds, revenue or labelling. Some even prohibit commercial use, for example MiniMax-M2.7.
What is the strongest open model today, and how far behind GPT, Claude and Gemini is it?
According to the Artificial Analysis Intelligence Index (checked on 30 Sep 2026), the leading open models are MiMo-V2.6-Pro (46 points), GLM-5.3 (45) and Kimi K3 (44), all from China. On the same index, Claude Opus 5.5 scores 58, GPT-5.6 Sol 47 and Gemini 3.8 Flash 41. This means the best open models already match or beat some commercial models, but are still about 12 points behind the top commercial model.
Which open models are good for Vietnamese?
On the SEA-HELM Vietnamese leaderboard (updated 18 Sep 2026), Qwen 3.5 397B scores 78.15 and Gemma 4 31B scores 77.09, close to Claude Opus 4.8 (78.70), Gemini 3.1 Pro Preview (78.66) and GPT 6 Astra (77.86). Qwen 3.6 27B scores 76.12. For small machines, Gemma 4 31B and Qwen 3.6 27B are the two candidates worth trying first; you should still test them on the business’s real documents.
Do Vietnamese businesses run into problems using Llama 4 commercially?
For most businesses, the Llama 4 Community License’s threshold of 700 million monthly users is not an issue. But the licence requires displaying “Built with Llama”, putting “Llama” at the start of derivative model names and complying with the acceptable use policy. That policy also does not grant rights to the Llama 4 multimodal models to companies with their principal place of business in the EU. To avoid these conditions, you can choose an Apache 2.0 model such as Gemma 4 or Muse Glimmer.
What size of open model can a Mac Studio or Mac mini run?
Based on the estimate Namtech uses for consulting (usable memory about 75% of RAM, a 4-bit model about 0.5 byte/parameter, 20% reserved for context): a 32GB Mac mini can run models up to about 38 billion parameters such as Qwen3.8-27B or Gemma 4 31B; a 128GB Mac Studio up to about 153 billion such as gpt-oss-120b or Mistral Small 4; a 256GB Mac Studio up to about 307 billion such as DeepSeek-V4-Flash. This is a capacity ceiling, not a speed measurement.
Is there a way to run open models internally without buying a machine?
Yes. Namtech offers Mac Studio and Mac mini rental on a 12-month term: the machine is hosted at Namtech, the business accesses the model via API or VPN, and Namtech administers the machine. See configurations and prices on Namtech’s Mac Studio rental page.
Choose the right open model for your business
Namtech helps filter models by licence, measure Vietnamese quality and speed on real documents, then deploy them on-site on a Mac Studio or Mac mini — bought or rented on a 12-month term.
Book a free consultationNote: This article compiles public sources checked on 30 Sep 2026. The benchmark figures in Table 3 are reported by each vendor under its own conditions; Tables 2 and 4 are independent leaderboards as of the date checked and change frequently. Model sizes by RAM are capacity estimates, not speed measurements. The licence section only summarises terms for reference and is not legal advice.
- Artificial Analysis — Open Source / Open Weights Models (open model ranking) — checked 30 Sep 2026
- Artificial Analysis — Intelligence Index v4.3.2 methodology — checked 30 Sep 2026; per-model scores on the Claude Opus 5.5, GPT-5.6 Sol, Gemini 3.8 Flash pages
- Arena — Text Leaderboard (25 Sep 2026) — checked 30 Sep 2026
- SEA-HELM (AI Singapore) — Vietnamese leaderboard, updated 18 Sep 2026 — checked 30 Sep 2026
- OpenAI — gpt-oss-120b & gpt-oss-20b Model Card (arXiv 2508.10925, 5 Aug 2025) · Hugging Face openai/gpt-oss-120b
- Google — Gemma 4: Byte for byte, the most capable open models (2 Apr 2026) · Hugging Face google/gemma-4-31B-it
- Meta — Muse Glimmer 30B model card (Aug 2026)
- Meta — Llama 4 Community License · Llama 4 Acceptable Use Policy · Meta AI blog — The Llama 4 herd (5 Apr 2025)
- Microsoft — Phi-4-reasoning-vision-15B model card (4 Mar 2026)
- IBM — Granite 4.2 (25 Aug 2026) · Hugging Face ibm-granite/granite-4.2-30b
- Qwen — Qwen3.8-27B model card · Qwen3.8-Max License · Qwen Community License 1.0 (Qwen3.8-Flash-Next) · Qwen3.5-397B-A17B · Qwen3.6-35B-A3B
- DeepSeek API Docs — Change Log (DeepSeek-V4.1-Flash, 10 Sep 2026) · Hugging Face DeepSeek-V4.1-Flash · DeepSeek-V4-Flash
- Z.ai — GLM-5.3 model card · GLM-5.3 License · GLM-5.3-Flash · GLM-5.2
- Moonshot AI — Kimi K3 model card · Kimi K3 License
- MiniMax Community License (MiniMax M3) · MiniMax-M2.7 Non-Commercial License
- Xiaomi — MiMo-V2.6-Pro-RL model card
- Mistral AI — Mistral Small 4 (16 Mar 2026) · Hugging Face Mistral-Small-4-119B-2603 · Introducing Mistral 3 (2 Dec 2025) · Devstral Small 2 24B
- Mistral AI — Mistral Medium 3.5 (22 May 2026) · Mistral Medium 3.5 LICENSE (Modified MIT)
- EuroLLM-22B-Instruct-2512 model card