Jargon
Fine-tuning changes a habit far more readily than it adds a fact
Continuing to train an existing model on new material is the most commonly recommended and most commonly misunderstood adjustment available, largely because of what people expect it to accomplish.
By Manish Trivedi3 min read

The mechanism is just more training, later
A fine-tuned model is one that was trained normally, then trained further on a smaller and more specific collection of examples, with the parameters starting from where the first process left them rather than from noise. There is no separate algorithm. It is the same procedure, resumed with different data.
The reason it is worth doing is economic. The general model already encodes an enormous amount about language, structure and the world, and starting from that costs a tiny fraction of building it. A useful adaptation can be produced with a quantity of data and computation that an individual can afford.
The learning rate is usually kept small, because the point is to nudge rather than to rebuild. That single choice explains most of what fine-tuning is good and bad at.
Style and format transfer easily; facts do not
What fine-tuning reliably changes is behaviour: the register a model writes in, the structure of its output, whether it asks clarifying questions, how it handles a particular kind of request, what format it returns. A modest number of consistent examples shifts these substantially.
What it does poorly is install new factual knowledge. A fact appearing a handful of times in a fine-tuning set competes against everything the model absorbed during its original training, and small nudges do not overwrite dense prior statistics. The common outcome is a model that has learned the shape of the new material without reliably learning its content.
Worse, it can learn that it should sound confident about a domain it has only skimmed. If the fine-tuning data consists of authoritative answers, the model learns to produce authoritative-sounding answers, and the confidence transfers more readily than the correctness. This is a genuinely counterproductive result and it is common enough to be worth expecting.
Retrieval is usually the answer to the question people are asking
When somebody says they want to fine-tune a model on their documents, they generally mean they want it to answer questions about those documents. Those are different requests, and the second is better served by keeping the documents in a searchable store and supplying the relevant ones alongside each question.
The advantages are practical rather than theoretical. Documents can be updated without retraining anything, the answer can cite what it drew on, permissions can be enforced per document, and a wrong answer can be traced to a source. None of these are available once material has been dissolved into parameters.
Fine-tuning remains the right tool when the requirement concerns how the system behaves rather than what it knows, and the two are frequently combined. The failure mode worth avoiding is reaching for the expensive option to solve a retrieval problem.
Adapting a model can degrade it elsewhere
Training on a narrow distribution pulls the parameters towards it, and capabilities not represented in the new data can weaken. This is sometimes described as catastrophic forgetting, a term borrowed from earlier work on sequential learning, and it means an adapted model may be better at its target task and quietly worse at things it previously handled.
Safety behaviour is a specific and well-documented case. Adjustments made during a model’s final training stages can be substantially undone by subsequent fine-tuning, including fine-tuning on material that has nothing to do with the behaviour in question. This is why providers offering fine-tuning generally impose review of the data.
Techniques exist to limit the damage — modifying only a small set of added parameters rather than all of them, mixing general data back into the specific data, keeping the learning rate very low. They reduce the effect. Evaluating whether they reduced it enough requires testing the things you were not trying to change, which is the step most often skipped.
A word that has drifted
The term now covers procedures that differ considerably. Adjusting every parameter is one thing; freezing the original model and training a small additional set of parameters alongside it is another; training on preference comparisons rather than examples is another again. All three get called fine-tuning in ordinary conversation.
The distinctions matter for cost, for how much the model can change, and for whether the result is a whole new model or a small file applied on top of an existing one. That last difference determines how it can be stored, shared and served, which is often the practical consideration.
When someone reports that fine-tuning did or did not work for them, the useful follow-up is which of these they did, on how much data, and what they measured afterwards. The answers vary enough that general claims about the technique are close to meaningless without them.
Common questions
How much data does fine-tuning need?
Far less than training from scratch, and the useful range depends entirely on the goal. Adjusting output format or tone can work with a modest, highly consistent set of examples. Teaching genuinely new behaviour needs considerably more. Consistency in the examples matters more than quantity, because contradictory examples teach inconsistency.
Is fine-tuning the same as giving instructions in the prompt?
No, though they can address similar goals. Instructions influence one conversation and change nothing about the model; fine-tuning changes the parameters and applies to everything afterwards. Instructions are free and reversible, which is why they are usually the sensible thing to exhaust first.
Does fine-tuning expose the training data?
It can. Material that appears in fine-tuning data can be reproduced by the resulting model, sometimes verbatim, particularly if it is distinctive or repeated. Anyone fine-tuning on confidential or personal material should treat the resulting model as carrying that material rather than as having merely learned from it.
Consumer editor, AI Worth Knowing
Manish has written about how it works, in the world, limits & risks for most of the last decade and prefers a plain explanation to a clever one.





