Blog

The hottest AI news & a perspective on internal AI

The big moves from Claude, ChatGPT, Grok, Copilot, Gemini and NVIDIA — with a perspective for Vietnamese enterprises on data sovereignty, PDPL compliance and private internal AI.

Clustering several Mac Studios for internal AI: RDMA over Thunderbolt 5, MLX and exo
PinnedApple24/09/2026

Clustering several Mac Studios for internal AI: RDMA over Thunderbolt 5, MLX and exo

A Mac Studio cluster for internal AI: RDMA over Thunderbolt 5 since macOS 26.2, distributed MLX and exo, how many cables, one machine vs a cluster and when one machine is enough.

Read more →
Building your own internal AI: why & the roadmap (8 steps)
Internal AI02/07/2026

Building your own internal AI: why & the roadmap (8 steps)

Why build your own internal AI, how it compares to cloud AI, and an 8-step roadmap from hardware to operations.

Read more →
Chat UI & integrating internal AI into your workflow
Internal AI02/07/2026

Chat UI & integrating internal AI into your workflow

A chat UI for staff, an OpenAI-compatible API to plug into your helpdesk/CRM, SSO authentication & permissions — all on-premise.

Read more →
Choosing open-source models & commercial licenses
Internal AI02/07/2026

Choosing open-source models & commercial licenses

Popular model families, how commercial licenses differ, picking the smallest size that works, and how to test candidates for internal AI.

Read more →
Commercially usable open AI models: US, China, Europe — benchmarks and each model’s strengths
Open models30/09/2026

Commercially usable open AI models: US, China, Europe — benchmarks and each model’s strengths

gpt-oss, Gemma 4, Muse Glimmer, Qwen3.8, DeepSeek V4.1, GLM-5.3, Kimi K3, Mistral: which licences allow commercial use, independent benchmark scores vs GPT, Claude, Gemini, and choosing a model by Mac Studio / Mac mini RAM.

Read more →
Departments & access control when rolling out internal AI
Internal AI02/07/2026

Departments & access control when rolling out internal AI

A department map for internal AI and how to set access control (RBAC) so each department only sees the documents it's allowed to — plus access control at the RAG layer.

Read more →
Evaluating quality & tuning internal AI (eval, guardrails)
Internal AI02/07/2026

Evaluating quality & tuning internal AI (eval, guardrails)

Measure quality with a golden set, reduce hallucination with RAG + citations, set guardrails, and when to fine-tune.

Read more →
How does internal AI reason to answer accurately, like Claude?
Internal AI02/07/2026

How does internal AI reason to answer accurately, like Claude?

Same next-token prediction on a transformer, grounded in your documents via RAG, step-by-step reasoning — and an honest answer to 'is it as good as Claude?'.

Read more →
Internal AI system architecture, layer by layer (with diagrams)
Internal AI02/07/2026

Internal AI system architecture, layer by layer (with diagrams)

An on-premise internal AI architecture explained layer by layer, plus a data-flow diagram for a single question.

Read more →
On-premise hardware for internal AI: how to choose
Internal AI02/07/2026

On-premise hardware for internal AI: how to choose

Apple Silicon vs GPU, the memory rule of thumb by model size, and how to size AI Box / AI Pro / AI Enterprise by scale.

Read more →
Operating, monitoring & scaling internal AI
Internal AI02/07/2026

Operating, monitoring & scaling internal AI

Monitor, back up, update models and move up AI Box → AI Pro → AI Enterprise — step 8/8 of the build-your-own internal AI series.

Read more →
RAG: teach your internal AI your own documents
Internal AI02/07/2026

RAG: teach your internal AI your own documents

Let your internal AI read your own documents and answer with citations: the RAG pipeline, vector DBs, multilingual embeddings and chunking.

Read more →
Serving: install & optimize model speed
Internal AI02/07/2026

Serving: install & optimize model speed

Choose Ollama, vLLM or llama.cpp; quantization; batching; an OpenAI-compatible API for internal AI.

Read more →
The security system of internal AI: defense in depth
Security02/07/2026

The security system of internal AI: defense in depth

The defense-in-depth layers for internal AI — network isolation, RBAC, encryption, audit logs, prompt-injection defense, PII, secrets management, physical security.

Read more →
Trending Pool: how internal AI stays current with world knowledge
Internal AI02/07/2026

Trending Pool: how internal AI stays current with world knowledge

A periodic knowledge-update mechanism through a controlled channel — keeping internal AI current while staying isolated, without opening the internet directly.

Read more →
What is an AI token? Data units & the context window
AI basics02/07/2026

What is an AI token? Data units & the context window

A token is the smallest unit an AI model reads and generates. Understand tokens, tokenization, the context window, and how they relate to AI cost.

Read more →
"The year of AI sovereignty": the EU AI Act tightens the rules, why Vietnamese businesses should keep data on-premise
Data Sovereignty23/06/2026

"The year of AI sovereignty": the EU AI Act tightens the rules, why Vietnamese businesses should keep data on-premise

The EU AI Act's high-risk obligations were pushed to 12/2027 and 08/2028, but fines still reach 7% of global turnover. Why consider on-premise internal AI.

Read more →
15 fake JetBrains plugins stole AI API keys from ~70,000 developers: a supply-chain lesson
Security02/07/2026

15 fake JetBrains plugins stole AI API keys from ~70,000 developers: a supply-chain lesson

Plugins posing as AI assistants on the JetBrains Marketplace stole OpenAI/DeepSeek keys. JetBrains removed them 17/06. A supply-chain lesson for businesses.

Read more →
2026 AI chatbot leaks: from McKinsey's Lilli to 300 million Chat & Ask AI messages — why on-premise AI is the way out
AI Security08/07/2026

2026 AI chatbot leaks: from McKinsey's Lilli to 300 million Chat & Ask AI messages — why on-premise AI is the way out

Two major 2026 AI chatbot data leaks — both caused by misconfigured infrastructure, not model flaws. Why on-premise AI is the way out.

Read more →
Anthropic Launches Claude Sonnet 5: Near-Flagship Performance, 1 Million Token Context, $2/$10 Pricing
Anthropic02/07/2026

Anthropic Launches Claude Sonnet 5: Near-Flagship Performance, 1 Million Token Context, $2/$10 Pricing

Anthropic's agentic mid-tier model (30/06): default for Free/Pro, 1M token context by default, $2/$10 per Mtok — the launch promo became permanent on 10/08.

Read more →
ChatGPT gets a new "automated assistant": schedule reminders, monitor the web and apps for you
OpenAI22/06/2026

ChatGPT gets a new "automated assistant": schedule reminders, monitor the web and apps for you

OpenAI launched Scheduled Tasks in ChatGPT — set reminders, recurring jobs and automatic web/app monitoring tasks.

Read more →
Choosing a vector database for on-prem RAG in 2026: pgvector, Qdrant, Milvus, Weaviate, Chroma compared
RAG · Vector database20/07/2026

Choosing a vector database for on-prem RAG in 2026: pgvector, Qdrant, Milvus, Weaviate, Chroma compared

A tooling guide: five open-source vector databases compared by license, version and storage, plus how to pick a multilingual embedding model for Vietnamese.

Read more →
Decree 142/2026: detailed rules under Vietnam's AI Law — what businesses need to know
AI Regulation23/06/2026

Decree 142/2026: detailed rules under Vietnam's AI Law — what businesses need to know

Decree 142/2026/NĐ-CP (effective 01/05/2026) implements AI Law 134/2025/QH15: 3 risk tiers (high, medium, low), conformity assessment, sandbox. What to prepare.

Read more →
DeepSeek unveils V4 preview: an open-source model that "closes the gap" with frontier models
DeepSeek22/06/2026

DeepSeek unveils V4 preview: an open-source model that "closes the gap" with frontier models

V4-Pro & V4-Flash open-source, 1 million token context — DeepSeek claims it closes the gap with frontier models.

Read more →
EU delays high-risk AI rules to Dec 2027 (Digital Omnibus): why sovereign AI still matters
AI Regulation06/07/2026

EU delays high-risk AI rules to Dec 2027 (Digital Omnibus): why sovereign AI still matters

The EU postpones high-risk AI obligations (Annex III) to 2 Dec 2027, but the penalty framework of up to EUR 35M / 7% of turnover stays intact. What Vietnamese firms should do.

Read more →
From pilot to production: in July 2026, enterprise AI money moved to two places — the bridge to production and the data border
Data sovereignty29/07/2026

From pilot to production: in July 2026, enterprise AI money moved to two places — the bridge to production and the data border

Three press releases in July 2026 show where enterprise AI money is going: the bridge from pilot to production, and infrastructure that keeps data inside a border.

Read more →
Gemini 3.5 Flash launches at Google I/O 2026: the Flash model that beats last-gen Pro at coding and AI agents
Gemini22/06/2026

Gemini 3.5 Flash launches at Google I/O 2026: the Flash model that beats last-gen Pro at coding and AI agents

Gemini 3.5 Flash — the default model for coding and AI agents, beating Gemini 3.1 Pro across a range of benchmarks, faster and much cheaper.

Read more →
Gemini 3.6 Flash: 17% fewer output tokens, but ML processing is still global-endpoint only — how Vietnamese businesses should read it
Gemini · AI cost · Data sovereignty22/07/2026

Gemini 3.6 Flash: 17% fewer output tokens, but ML processing is still global-endpoint only — how Vietnamese businesses should read it

Google shipped Gemini 3.6 Flash on 21 Jul 2026: output at $7.50 per 1M tokens (down from $9.00) and 17% fewer tokens. The trade-off: no ML processing region is listed yet.

Read more →
Getty Images teams up with OpenAI: licensed photos go straight into ChatGPT, Getty stock soars
OpenAI23/06/2026

Getty Images teams up with OpenAI: licensed photos go straight into ChatGPT, Getty stock soars

A multi-year display agreement — Getty images appear in ChatGPT (not training). Getty stock soars.

Read more →
GLM-5.2: the world's strongest open-weights model (as of 24 July 2026) now ships under MIT — and it changes the on-premise AI equation for businesses
Open models24/07/2026

GLM-5.2: the world's strongest open-weights model (as of 24 July 2026) now ships under MIT — and it changes the on-premise AI equation for businesses

A near-frontier model you can self-host, MIT-licensed, with a 1M-token context — and it changes the on-premise AI and data-sovereignty equation for businesses.

Read more →
Goodbye Open Llama: Meta Launches Muse Spark — the 'Superintelligence Lab's' First Proprietary AI Model
Meta22/06/2026

Goodbye Open Llama: Meta Launches Muse Spark — the 'Superintelligence Lab's' First Proprietary AI Model

Meta Superintelligence Labs launches Muse Spark on a proprietary path, breaking from the Llama line's open-source tradition.

Read more →
Google launches Nano Banana 2 Lite: AI images in 4 seconds, $0.034 per 1,000 images
Google02/07/2026

Google launches Nano Banana 2 Lite: AI images in 4 seconds, $0.034 per 1,000 images

Gemini 3.1 Flash-Lite Image (30/06): Google's fastest & cheapest image generation model for high-volume enterprises.

Read more →
GPT-5.6 launches like never before: the US government approves each customer allowed to use the most powerful model
OpenAI02/07/2026

GPT-5.6 launches like never before: the US government approves each customer allowed to use the most powerful model

OpenAI announced Sol/Terra/Luna (26/06) — in the initial phase only ~20 partners vetted by the US government can use Sol. A new precedent in AI control.

Read more →
How long does an in-house AI rollout take? The 8–10 week plan, who the client has to assign, and the six variables that decide whether you finish on time
In-house AI05/08/2026

How long does an in-house AI rollout take? The 8–10 week plan, who the client has to assign, and the six variables that decide whether you finish on time

There is no single correct number for every company. There is a readable plan: five phases, 8–10 weeks, and six variables that decide which end of that range you land on.

Read more →
MCP drops sessions: the 28 July 2026 spec makes in-house agents replicable — and here is the list you must fix before upgrading
Internal AI03/08/2026

MCP drops sessions: the 28 July 2026 spec makes in-house agents replicable — and here is the list you must fix before upgrading

MCP moves from a bidirectional stateful protocol to stateless request/response. In-house servers can now run many replicas behind a plain round-robin load balancer — at the cost of a breaking release.

Read more →
Microsoft brings Copilot processing in-country for 15 nations — why "data sovereignty" is pushing enterprises toward internal AI
AI Analysis15/07/2026

Microsoft brings Copilot processing in-country for 15 nations — why "data sovereignty" is pushing enterprises toward internal AI

Microsoft is expanding in-country processing for Microsoft 365 Copilot to 15 nations. Data sovereignty is now mainstream — and why internal / on-premise AI is the strongest level of control.

Read more →
Microsoft builds 7 in-house "MAI" models at Build 2026: less reliance on OpenAI, Copilot becomes a "super app"
Microsoft23/06/2026

Microsoft builds 7 in-house "MAI" models at Build 2026: less reliance on OpenAI, Copilot becomes a "super app"

7 self-developed models for coding/reasoning/image/voice + a Copilot 'super app' — Microsoft cuts reliance on OpenAI.

Read more →
Microsoft launches "Microsoft IQ" at Build 2026: a context layer for AI agents on Copilot, Foundry and Copilot Studio
Copilot22/06/2026

Microsoft launches "Microsoft IQ" at Build 2026: a context layer for AI agents on Copilot, Foundry and Copilot Studio

Microsoft IQ — a context layer that helps AI agents understand both world knowledge and internal data. GA on Copilot, Foundry and Copilot Studio.

Read more →
Mistral launches Mistral Small 4: merging reasoning, multimodality and coding into one open-source model
Mistral22/06/2026

Mistral launches Mistral Small 4: merging reasoning, multimodality and coding into one open-source model

Mistral Small 4 — MoE 119B (6B active), 256k context, Apache 2.0, merging reasoning + multimodality + coding agent; runs on ~4× H100.

Read more →
Moonshot launches Kimi K3: the world's largest open-weight model at 2.8 trillion parameters — closing in on Claude Fable 5 and GPT-5.6 Sol
Moonshot AI17/07/2026

Moonshot launches Kimi K3: the world's largest open-weight model at 2.8 trillion parameters — closing in on Claude Fable 5 and GPT-5.6 Sol

1M-token context, benchmarks trailing only Claude Fable 5 and GPT-5.6 Sol, prices at a third of closed rivals — weights open on July 27, 2026.

Read more →
NVIDIA launches Vera Rubin: a new chip lineup for the era of the "AI factory" and autonomous agents
NVIDIA23/06/2026

NVIDIA launches Vera Rubin: a new chip lineup for the era of the "AI factory" and autonomous agents

Vera Rubin's new chips span from training to agentic inference — the foundation for enterprises to build their own AI infrastructure.

Read more →
NVIDIA unveils the RTX Spark superchip: running AI agents and giant LLMs directly on personal PCs
NVIDIA22/06/2026

NVIDIA unveils the RTX Spark superchip: running AI agents and giant LLMs directly on personal PCs

RTX Spark — a CPU+GPU superchip bringing LLMs up to 120 billion parameters and AI agents to run right on personal PCs, no cloud needed.

Read more →
OpenAI Codex CLI bug silently wears down SSDs: ~640 TB/year, drives worn out in under a year
OpenAI23/06/2026

OpenAI Codex CLI bug silently wears down SSDs: ~640 TB/year, drives worn out in under a year

Codex CLI logs excessively (~37 TB/21 days ≈ 640 TB/year) — enough to drain a 1TB SSD's lifespan in under a year. Cause, the patch and how to fix it.

Read more →
OpenAI retires GPT-5.2 for GPT-5.5 and adds "Active sessions": model lifecycles keep getting shorter
OpenAI23/06/2026

OpenAI retires GPT-5.2 for GPT-5.5 and adds "Active sessions": model lifecycles keep getting shorter

AI model lifecycles keep getting shorter: GPT-5.2 retires, conversations move to GPT-5.5. The stability lesson for businesses.

Read more →
OpenAI unveils its first AI chip "Jalapeño" with Broadcom: designed in 9 months, aiming to slash inference costs
OpenAI02/07/2026

OpenAI unveils its first AI chip "Jalapeño" with Broadcom: designed in 9 months, aiming to slash inference costs

The first inference ASIC designed by OpenAI (24/06) — 9-month tape-out, deployment late 2026. The race for AI hardware independence.

Read more →
Perplexity launches "Brain": AI memory that self-learns overnight to make the Computer agent smarter
Perplexity22/06/2026

Perplexity launches "Brain": AI memory that self-learns overnight to make the Computer agent smarter

Brain builds a living 'context graph' of what the agent has done, then overnight synthesizes it into an 'LLM wiki' so the agent can improve itself — a Research Preview for the Max & Enterprise Max pla

Read more →
ServiceNow data exposure from a misconfigured endpoint: lessons in controlling sensitive data
Security23/06/2026

ServiceNow data exposure from a misconfigured endpoint: lessons in controlling sensitive data

A configuration error let data be accessed beyond permissions — even a large SaaS platform still leaked. Lessons in controlling sensitive data.

Read more →
Sovereign LLM on-premise: why 2026 is the year enterprises "bring AI home"
AI Analysis13/07/2026

Sovereign LLM on-premise: why 2026 is the year enterprises "bring AI home"

A sovereign LLM delivered on-premise for a telecom (1 July 2026), IBM's data on shadow-AI breach costs, and Vietnam's tightening data laws converge into one trend.

Read more →
The EU Cybersecurity & AI Action Plan (7 July 2026): when AI is both the weapon and the shield — how businesses should read it
AI policy27/07/2026

The EU Cybersecurity & AI Action Plan (7 July 2026): when AI is both the weapon and the shield — how businesses should read it

Presented 7 July 2026: evaluating AI models before market entry, a structured-access Blueprint, a secure testing platform — and what it means for businesses outside the EU.

Read more →
The real cost of running an LLM on-premise in 2026: VRAM, GPU or Mac Studio, and electricity
On-premise AI18/07/2026

The real cost of running an LLM on-premise in 2026: VRAM, GPU or Mac Studio, and electricity

A TCO breakdown for businesses: VRAM per open model, GPU and Mac Studio electricity at official Vietnam rates, hardware prices for three internal AI packages, vs cloud APIs.

Read more →
US forces Anthropic to pull Fable 5 worldwide: what happened and lessons for Vietnamese businesses
Anthropic22/06/2026

US forces Anthropic to pull Fable 5 worldwide: what happened and lessons for Vietnamese businesses

Fable 5 & Mythos 5 pulled worldwide after a US government order — the risk of depending on foreign cloud AI.

Read more →
What exactly did you just download? Cisco opens a provenance database of nearly 900 open models — and a standard for what "derived from" means
Security31/07/2026

What exactly did you just download? Cisco opens a provenance database of nearly 900 open models — and a standard for what "derived from" means

Cisco opens a provenance database of nearly 900 open models, plus a standard for what counts as derivation. For anyone running AI in-house, this is the question to answer before a model goes in.

Read more →
wp2shell: WordPress core flaw exploited in the wild — a two-CVE chain, patch to 6.9.5 / 7.0.2 now
WordPress · Security · RCE · Patching21/07/2026

wp2shell: WordPress core flaw exploited in the wild — a two-CVE chain, patch to 6.9.5 / 7.0.2 now

What happened between 17 and 20 July 2026, how the two-CVE chain works, who is affected, and the 24-hour checklist for businesses.

Read more →
xAI brings Grok to Databricks: AI agents reason directly on enterprise internal data
Grok23/06/2026

xAI brings Grok to Databricks: AI agents reason directly on enterprise internal data

Grok is natively available on Databricks (18/06) — agents reason on Lakehouse data, no external pipeline needed.

Read more →
xAI's Grok 4.3 lands on Amazon Bedrock: 1 million token context window, configurable reasoning for the enterprise
Grok22/06/2026

xAI's Grok 4.3 lands on Amazon Bedrock: 1 million token context window, configurable reasoning for the enterprise

Grok 4.3 arrives on AWS Bedrock with a 1 million token context window and configurable reasoning levels — aimed at enterprise workloads.

Read more →