There’s been a lot of noise lately around open models. In this piece I want to address three concerns and three corresponding truths at the same time.
I think what I am about to say is common sense, or at least should be common sense. So why is there so much debate, misleading statements and contrary views?
There is a negative, even nefarious, side to how some open-source models appear to have been created. That problem is model distillation.
Distillation itself is a legitimate and widely accepted technique for fine-tuning language models. The issue arises when it is used at scale to extract the shape, reasoning patterns and capabilities of a frontier model.
This is achived by generating high-volume, high-quality input–output pairs and training on them. In effect, it becomes a form of model theft.
Anthropic has publicly pointed to DeepSeek, Moonshot AI (the company behind Kimi) and at least one other player as having engaged in this kind of behaviour.
Whether or not every allegation sticks, the principle remains: frontier labs have every incentive (and responsibility I guess) to detect and disrupt large-scale query campaigns designed to reverse-engineer their models.
That is the first reality we should keep in mind…
The second concern is more practical and, in my view, under-discussed.
If you use an open-weight model that is hosted by a party you do not fully trust, you have handed over data governance.
When you log into a DeepSeek-hosted or Kimi-hosted endpoint, you typically have no visibility into where the data flows, how long it is retained, or how it might be used.
For enterprises especially, this should be a non-starter.
The ability to run models in an air-gapped environment with complete control over data flow is one of the strongest arguments for open weights.
I’m continually surprised that this point does not feature more prominently in the debate.
If the distillation allegations have any substance, then the case against using hosted instances operated by the accused parties becomes even stronger. You simply do not control the data and they don’t care about data governance and privacy.
Here is the part that often gets missed…
These models have open weights. You can download them. You can run them locally or in a private cloud under full air-gap conditions.
NVIDIA’s Nemotron effort has shown how effectively this can be done: taking open models, fine-tuning them and deploying them under tight control.
Perplexity has done something similar, downloading DeepSeek weights, fine-tuning and deliberately reducing certain political constraints that originated in the original Chinese training regime.
It would be a genuinely sad day for AI progress if regulatory or commercial pressure made this kind of responsible reuse difficult or impossible.
The opportunity is enormous: take highly capable open models…regardless of their origin…and place them in environments where the organisation retains complete sovereignty over data and compute.
No foreign host sees the traffic. No company-specific information leaves the perimeter. The weights become just another downloadable, editable, privately runnable software artefact.
There is a further, quieter benefit.
As more teams build agentic systems, they increasingly orchestrate multiple specialised models, many of them open-weight, assigning different models to different stages of a process.
For development, research and longer-running agent loops, the ability to run capable models locally or on private infrastructure is a genuine advantage.
Frontier APIs will always have their place for the absolute cutting edge, but open weights remove friction and cost for everything else.
I do not claim any special insight.
I simply find the current uproar and the accompanying misunderstandings about open-weight models difficult to reconcile with the practical realities.
Distillation risks exist and should be policed.
Hosted instances from untrusted parties carry real governance risk.
But the open weights themselves remain one of the most powerful tools available for organisations that care about control, cost and capability.
Treating them as inherently dangerous misses the point.
The differentiator for the frontier labs will continue to be the models that stay at the absolute leading edge.
Everything else is increasingly becoming infrastructure that serious teams will want to own and run themselves.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.