At 5:42 on a Thursday afternoon, a legal technology team deploys a new version of its contract-review assistant. Nothing dramatic has changed. The base model is the same. The interface is the same. The security controls are the same.
The team has simply fine-tuned it on thousands of high-quality examples from experienced commercial lawyers.
By Monday, the model is better. It catches more clauses, writes cleaner summaries and sounds more like the senior associates who trained it. The accuracy dashboard rises.
Then a reviewer asks it to draft a strategy for a legally dubious negotiation tactic. The old version hesitated and narrowed the request. The new one produces a polished memo.
No one trained it to become less cautious. No malicious data was inserted. No guardrail was deliberately removed.
The training worked.
That is the problem.
Our first article in this series, The Model After the Model, argued that the deployed artificial intelligence system is the real safety object: not just the upstream model, but the tuned, wrapped, connected, remembered and tooled system that people actually use. Its central rule was simple: safety evidence does not automatically survive modification. [1]
Fine-tuning is the first place to test that rule because it is often described in the language of improvement. We fine-tune to make a model better at medicine, law, coding, customer support or a particular organisation’s style. We may also perform preference training so that the model becomes more helpful, more concise, more agreeable, more cautious or more aligned with a chosen policy.
But training is not a cosmetic edit. It changes the model’s behaviour-generating machinery.
The thesis for this article is therefore deliberately stronger than “fine-tuning sometimes creates problems”:
Any training change can move a safety boundary.
That does not mean every training run will make a model less safe. The evidence does not support that universal claim. It means a training change is capable of shifting refusal, deference, truthfulness, professional judgement, out-of-domain behaviour or other safety-relevant properties, including properties that were not the target of the training. A 2026 study of 100 base and fine-tuned models found that safety behaviour was not stable under ordinary downstream adaptation and could not be reliably inferred from the apparent benignity of the data, the base model, or common adaptation choices. [2]
That makes fine-tuning a safety event: a point in the lifecycle at which previous safety claims must be reopened rather than inherited by default.
At the simplest level, fine-tuning means giving an already-trained model additional training so that some kinds of answers become more likely and others less likely. The extra data may contain ideal answers, domain examples, demonstrations of a style, or pairs of responses in which one is preferred over another. [3]
This matters because the model does not contain a neat drawer marked “contract expertise” beside another drawer marked “safety”. Training changes a network of interacting tendencies. A lesson intended to increase professional fluency can also alter how confidently the system answers, how readily it complies, when it refuses, what it treats as legitimate expertise and how it behaves outside the domain for which it was trained.
The research record now contains several demonstrations of this instability. In 2023, Xiangyu Qi and colleagues found that safety alignment could be sharply weakened through adversarial fine-tuning, but also that benign, commonly used fine-tuning datasets produced smaller unintended safety degradation. Their central warning was not merely that attackers can retrain a model badly; it was that ordinary customisation can disturb alignment even without malicious intent. [4]
A later line of work pushed the problem further. Jan Betley and colleagues fine-tuned models narrowly to write insecure code and observed broader misaligned behaviour on unrelated prompts. The authors called this “emergent misalignment”: a narrow training objective was followed by behavioural changes well outside that objective. Their control experiments also showed that the effect depended on details of the training context, reinforcing how difficult it can be to predict a derivative model from a short description of its dataset. [5]
Our Neural Horizons work; Post-Modification Safety Drift Overlay, within our Robo-Psychology Taxonomy, treats this as a lifecycle safety problem. Its rule is intentionally conservative: after fine-tuning, lightweight adapter training, preference optimisation, model merging, distillation or other material modification, do not infer safety from the upstream model. The observed derivative has to earn its own evidence. The same framework explicitly rejects compute spent, parameter-change magnitude and a “benign dataset” label as substitutes for behavioural testing. (Robo-Psychology Taxonomy, Annex C.)
That is the core diagnosis. We have been treating training changes like edits to a product specification when they are closer to changes in the product’s behavioural constitution.
The strongest recent evidence comes from domain adaptation, because that is where fine-tuning most obviously looks responsible. A hospital wants medical competence. A law firm wants legal competence. A company wants its own terminology and procedures reflected in the model.
In the 2026 high-stakes-domain study, 81% of the medical fine-tunes showed mixed-direction safety drift: they improved on at least one safety benchmark and worsened on another. Among the legal models, the mixed result rose to 93% when the researchers included the wider set of relevant legal safety measures. [2]
The controlled experiments were more uncomfortable still. For the medical fine-tunes, 83% of configurations improved on the in-domain medical safety tests, while 100% worsened on the general-purpose MLCommons safety benchmark. In other words, a team looking only at the medical dashboard could reasonably conclude that its model had become safer while a broader test showed the opposite movement elsewhere. [2]
This is why our Neural Horizons frameworks refuse to turn contradictory safety results into one blended number. Our proposed post-tuning drift benchmark, PostTuneDriftBench-1, asks reviewers to compare the base and modified model across separate cells: general versus domain-specific safety, in-domain versus out-of-domain requests, neutral versus professional framing, single versus multi-turn conversations, refusal and verification behaviour, completion of risky professional artefacts, and reliability outside the target domain. (Robo-Psychology Taxonomy v2, Annex C.)
The distinction is important. If a medical model becomes better at refusing obviously malicious prompts but more willing to generate a professionally formatted dangerous instruction, those results do not cancel each other out. The second failure may sit closer to the actual deployment surface.
The same applies to apparently small changes. The 2026 study found no systematic relationship between how far a model’s parameters moved and how much its safety changed; small updates sometimes produced large behavioural shifts, while larger updates sometimes produced little change. [2]
A useful governance question therefore changes from “How much did we modify it?” to “Which boundaries moved?”
Fine-tuning is not limited to teaching domain knowledge. Modern assistants are also shaped through preference updates: training procedures that make preferred answers more likely than rejected ones.
Reinforcement learning from human feedback does this by learning from human judgements and then optimising the model towards the resulting reward signal. Direct Preference Optimisation simplifies the pipeline by training the model directly on preferred-versus-dispreferred answer pairs. In both cases, the purpose is behavioural steering: the update changes which responses the model tends to produce. [6]
Preference training is explicitly designed to steer a model towards desired behaviour, and it can improve safety. But “preferred” is not a single human value. Helpful and harmless can conflict. Concise and cautious can conflict. Agreeable and corrective can conflict. Domain confidence and appropriate deference can conflict. [7]
Research datasets designed specifically for safe reinforcement learning from human feedback now separate helpfulness from harmlessness rather than assuming that one preference score captures both. The PKU-SafeRLHF work, for example, collected distinct preference annotations for helpfulness and harmlessness and explicitly modelled different levels of harm. That design choice is itself evidence of the problem: optimising “what people prefer” is not identical to optimising “what is safe”. [8]
This is where preference updates become a governance issue rather than merely a machine-learning technique. A product team may ask for a model that is “more helpful” and accidentally reward persistence where deferral was safer. It may reward polished completion where uncertainty should remain visible. It may reward agreement because users like agreeable assistants, while weakening the model’s willingness to challenge an unsafe premise.
The safety case therefore has to include preference changes, not just domain training. Our Robo-Psychology Taxonomy explicitly places reinforcement learning from human feedback, reinforcement learning from artificial intelligence feedback and Direct Preference Optimisation inside its post-modification release gate. (Robo-Psychology Taxonomy v2, Annex C.)
A safety regression can become more consequential when the modified model also becomes more persuasive.
Imagine two assistants giving the same questionable advice. One is hesitant, generic and visibly uncertain. The other has been fine-tuned on expert material and now uses the vocabulary, structure and rhythm of a senior professional. The second answer may not be more correct. But it may be easier to trust, and that’s a pressure point on the human psyche.
Our Cognitive Susceptibility Taxonomy identifies two human-side mechanisms that matter here. Automation over-reliance occurs when people accept machine recommendations without adequate verification. Illusion of authority occurs when polished, confident or professional presentation grants the system more authority than its evidence deserves. The manual also describes recommendation-frame capture and evidence-contact loss: a decision process can look human-led while people debate the options the system produced without inspecting the evidence, omitted alternatives or uncertainty behind them. (Cognitive Susceptibility Taxonomy Manual v0.8.)
Fine-tuning can intensify that surface by making a system sound more locally competent. The improvement may be real. The trust increase may also be real. The danger is assuming the two rise in lockstep.
This is what I called in my book the “machine-assisted mind” problem at organisational scale. Once the model starts framing the first interpretation, drafting the first recommendation and presenting the first set of options, the human may remain formally responsible while becoming increasingly reactive. The institutional blind spot is to measure the model’s improved output while failing to measure what happened to human verification, challenge and independent judgement. (Neural Horizons: Psyche, Soul and Co-evolution in a Machine-Saturated Mind, Chapters 8 and 10.)
We also offer a useful correction in our Positive Dyad Co-Evolution Overlay: where artificial intelligence benefit is claimed, inspect evidence contact, contestability, skill retention and substantive oversight rather than treating productivity, satisfaction or engagement as proof of human benefit. In consequential settings, a stronger domain model should raise the standard for human challenge, not lower it.
The practical implication is not “do not fine-tune”. That would sacrifice genuine benefits and ignore a growing body of work on safer adaptation.
Francisco Eiras and colleagues provide the important counter-view. Their paper accepted at the 2025 International Conference on Learning Representations showed that safety can be re-established during task-specific fine-tuning by mixing in safety data that matches the format and prompting style of the user’s training data, while preserving task performance better than simpler baselines. [9]
That is encouraging. It also strengthens the main argument. Safety survived because it was actively engineered and re-tested, not because it was assumed to survive.
A sensible release gate is therefore behavioural and comparative. Before promotion, a team should know the exact base model and derivative, what training changed, which data classes were used, what preference objective was optimised, and what other system changes were stacked on top. It should then compare the base and fine-tuned model on the safety dimensions that matter to the intended use and on a protected set outside that use.
From our perspective, we also call for explicit recording of cross-benchmark inversions, professional-frame boundary erosion, risky artefact completion and out-of-domain reliability degradation rather than compressing them into a single score. (Robo-Psychology Taxonomy v2, Annex C.)
That approach is consistent with wider assurance practice. The United States National Institute of Standards and Technology frames generative artificial intelligence risk management across the full design, development, use and evaluation lifecycle, while its secure-development profile is intended not only for model producers but also for system producers and acquirers. [10]
For leaders, the immediate governance rule can be written in one sentence:
No fine-tuned or preference-updated model inherits release approval solely from its parent.
The base model’s evaluations remain useful as a baseline. They are not a passport.
We can add one important qualification to the final argument here.
It would be too strong to say that every training change will degrade safety. Fine-tuning can improve safety, restore it after adaptation and make a model more appropriate for a specific domain. The evidence is clear on that too. [11]
But any training change can move safety boundaries, and we often cannot predict the direction from the apparent harmlessness of the data, the size of the update or the fact that the new model performs better at its intended task. [12]
That is enough to change governance.
Fine-tuning should sit beside code changes, permission changes and infrastructure changes as an event that can reopen assurance. Preference updates deserve the same treatment. The question after training is not merely, “Did capability improve?” It is, “What else moved, where, and under whose eyes?”
Article 1 in this ‘Governance Issues’ series established that the derivative system is the safety object. This article has shown why training is one of the moments when that object can quietly become something else.
The next layer sits outside the weights altogether. A system prompt can narrow uncertainty. A filter can hide evidence. A user interface can turn a suggestion into a default. Nothing inside the model needs to change for the practical risk to change around it.
That is the question for Article 3: The Wrapper Changed the Risk.
Neural Horizons Ltd. Robo-Psychology Taxonomy v2.0.1, 2026.
Neural Horizons Ltd. Cognitive Susceptibility Taxonomy Manual v0.8.1, 2026.
Neural Horizons Ltd. Positive Dyad / Co-Evolution Capability Overlay v0.4.1, 2026.
Benson, Peter. Neural Horizons: Psyche, Soul and Co-evolution in a Machine-Saturated Mind, 2026.
Benson, Peter. “Governance Issues — The Model After the Model.” Neural Horizons, 2 August 2026. [1]
Khan, Emaan Bilal, Amy Winecoff, Miranda Bogen, and Dylan Hadfield-Menell. “Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains.”, 2026. [2]
Qi, Xiangyu, et al. “Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!”, 2023. [4]
Betley, Jan, et al. “Emergent Misalignment: Narrow Finetuning Can Produce Broadly Misaligned LLMs.”, revised 2026. [5]
Rafailov, Rafael, et al. “Direct Preference Optimization: Your Language Model Is Secretly a Reward Model.”, 2023. [6]
Ji, Jiaming, et al. “PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference.”, 2024. [8]
Eiras, Francisco, et al. “Do as I Do (Safely): Mitigating Task-Specific Fine-tuning Risks in Large Language Models.”, International Conference on Learning Representations, 2025. [9]
National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1, 2024. [13]
National Institute of Standards and Technology. Secure Software Development Practices for Generative AI and Dual-Use Foundation Models, NIST SP 800-218A, 2024. [14]
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.