Fine-tuning is when an existing AI model is trained on a smaller data set for performing a particular task. The model retains all the previous knowledge it had before fine-tuning. In fine-tuning, the focus of the model’s knowledge of the model gets narrowed down to a particular task, domain, or tone.
All six types of fine-tuning are discussed here along with the precise steps in them, cost details of 2026, and a question not asked by most of the guides: Is fine-tuning needed at all, or does RAG or prompt engineering provide a cheaper solution?
What Will I Learn?
Key takeaways
The following points summarize this guide:
- Fine-tuning involves the adjustment of some or all the pre-existing weights of the machine learning model using additional and task-specific data.
- There are six approaches including full fine-tuning, feature extraction, parameter-efficient fine-tuning (PEFT), LoRA, adapters, and prompt tuning.
- RLHF and DPO involve the fine-tuning of a model’s behavior as opposed to knowledge.
- OpenAI announced its discontinuation of self-serve fine-tuning on May 7, 2026 due to better instruction following in base models.
- RAG is one of the alternatives to fine-tuning for incorporating current or personal facts.
What Is Fine-Tuning?
Fine-Tuning is an approach to train an already trained model to adapt itself to perform a new task. Fine-tuning continues the training process on a smaller labeled dataset for a new task. The model comes with some weights which have already been trained in a general larger dataset.
Models are generally trained to recognize general patterns. A language model will learn grammar, factual knowledge, and logical reasoning from a general textual dataset. In the case of an image, the model will learn edges, textures, and shapes from a general image dataset. Fine-tuning refines this general model through another training process.
Three requirements are necessary for fine-tuning:
- It has to be done on a pre-trained model only.
- The second round of training takes place on a dataset that is smaller than the original dataset on which the first training was done.
- This time the target has to be smaller than the previous one.
Fine-Tuning vs. Pre-Training vs. Transfer Learning
The processes of pre-training and fine-tuning belong to one pipeline rather than competitive technologies.
Pre-training provides the initial model. The process is started by initializing the parameters randomly and training the model on a vast amount of data that is not specific to any one task – it could be millions of text tokens in the case of a language model or millions of images for a vision model.
The fine-tuning process begins when the pre-training stage is finished. It uses the learned weights from pre-training to initialize the model and continues training on a smaller, task-specific dataset. Fine-tuning is less expensive and quicker than pre-training.
Transfer learning is a more general approach that includes fine-tuning within it. Transfer learning is the practice of using knowledge gained from one task for helping with a new but similar task. Fine-tuning is just one form of transfer learning. Feature extraction, which we will see in the next section, is another.
How Fine-Tuning Works, Step by Step
Fine-tuning follows five steps in a typical workflow.
- Choose a pre-trained base model. The choice of the model has to be appropriate to the task to be performed. For instance, BERT and RoBERTa are good for text classification tasks. GPT-like models work well for text generation tasks.
- Freeze the early layers. The initial layers of a neural network learn general features. The early layers in an image model will learn to identify edges and textures while the early layers in a language model will learn grammar and simple word associations.
- Train the later layers on the new dataset. The following layers pick up specific patterns related to the task. During training, just these layers or even a few parameters get updated.
- Apply a small learning rate. A very high learning rate can cause catastrophic forgetting, i.e., forgetting what has already been learned. On the other hand, a low learning rate enables slow learning and retains all abilities of the base model.
- Evaluate the model and iterate. Test the trained model on the held-out dataset. Tune the hyperparameters such as the learning rate, batch size, and the number of epochs until it achieves the desired performance level.
The 6 Types of Fine-Tuning, Compared
There are six techniques that address all fine-tuning applications. Each technique varies based on the number of parameters that need to be updated and the computational requirements involved.
Full fine-tuning modifies all the parameters in the model. This is done in the same way as pre-training using the same training procedure but with a smaller training dataset. Full fine-tuning yields the greatest possible accuracy on tasks which are considerably different from the training data of the pre-trained model. In terms of memory consumption and training time, full fine-tuning is the most resource-demanding among the six techniques.
Feature extraction involves freezing the whole pre-trained model and training only a small number of additional layers on top of the pre-trained model. This means that the pre-trained model serves as a fixed feature extractor. Feature extraction is the least computationally demanding technique among the six.
Parameter-efficient fine-tuning (PEFT) is a broad term for techniques that modify a small portion of the model’s parameters as opposed to modifying the whole network. Researchers Lialin, Deshpande, and Rumshisky have noted that training of all the parameters of the model consumes 12 to 20 times more GPU memory than storing the model’s weights.
The LoRA and QLoRA models first freeze the weight of the original model and then learn two small matrices that capture the changes needed for the new task, thereby reducing the number of trainable parameters than full fine-tuning. QLoRA is an extension of LoRA that involves quantization, i.e., decreasing the numerical precision of the weight storage of the frozen base model, before applying LoRA to reduce memory consumption.
Adapter models, small and newly added layers are inserted into the already present layers of the pre-trained model. In training, only the adapter layer’s weight is updated while the original model weight remains frozen. It was observed by Houlsby et al. that adapter modules achieve comparable performance on BERT as that of full fine-tuning with only 3.6% of parameters being updated.
The prompt tuning technique freezes the whole model and learns a small set of input embeddings referred to as soft prompts. The model does not change in prompt tuning. Prompt tuning has the fewest trainable parameters among all the techniques discussed.
| Method | Parameters updated | Relative compute cost | Best for | Forgetting risk |
|---|---|---|---|---|
| Full fine-tuning | 100% | Highest | Tasks very different from pre-training | Highest |
| Feature extraction | New layers only | Lowest | Small datasets, similar tasks | Lowest |
| PEFT (general) | A small subset | Low | Limited compute budgets | Low |
| LoRA / QLoRA | Under 1% typically | Low | Large models, frequent task-switching | Low |
| Adapters | About 3.6% (BERT benchmark) | Low to moderate | Matching full fine-tuning cheaply | Low |
| Prompt tuning | Soft prompt only | Lowest | Fast task-switching, minimal changes | Lowest |
Full fine-tuning involves parameter updates compared to all other approaches shown in this table. Prompt tuning involves the least amount of parameter updates. The choice of the approach depends on the number of data samples and the difference between the tasks.
Fine-Tuning Large Language Models: Instruction Tuning, RLHF, and DPO
Language models such as GPT, Claude, Gemini, and Llama require an additional step in their training process after the six mentioned above in order to act like assistants and not just autocomplete tools.
Instruction tuning is a process whereby the model is trained on prompt/response pairs, whereby the response provides a demonstration of how to correctly respond to the prompt. A language model which hasn’t been tuned through instruction will produce a grammatically valid completion of the text, but it won’t always be able to perform the request made.
RLHF (Reinforcement Learning From Human Feedback) trains another reward model using human preference judgments and then applies reinforcement learning to tune the language model towards high-reward responses. RLHF tunes qualities that can’t be specified with labeled examples only, such as helpfulness, tone, and factual accuracy.
The direct preference optimization (DPO) technique was proposed by Rafailov et al. in 2023 and achieves the same alignment objective as RLHF without needing to train a reward model. This technique works with pairs of responses, where one is chosen while the other is not, and fine-tunes the model to increase the probability of the selected response. The published results indicate that DPO performs at least as well as RLHF in generating output of high quality for tasks such as sentiment control, summarization, and conversation.
Fine-Tuning vs. RAG vs. Prompt Engineering: Which Do You Need?
There are three methods that address three distinct problems. Mixing them up will cost valuable engineering resources and money.
Use RAG if you want the model to use information it is not aware of — latest news, confidential information from an organization, or a dynamic database. The RAG method searches for relevant pieces of text at the moment of making a request and includes them into the model’s context window. No modification of the model’s weights occurs.
Use fine-tuning if you need a model to develop a new capability, to have a specific tone, or an output format that cannot be guaranteed by the use of prompting.
Use prompt engineering if the task is a one-shot or if a good system prompt along with some examples provides satisfactory output. Prompt engineering does not need any training and infrastructure.
The three approaches are not exclusive. A deployed system normally makes use of RAG for knowledge and a fine-tuned or well-prompted model for behavior.
The reasoning provided by OpenAI about its decision to stop its self-serve fine-tuning program is an indication of this trend. On May 7, 2026, OpenAI informed developers that newer base models are accurate enough at following instructions such that prompt-based techniques are sufficient to accomplish what used to require fine-tuning. The process of stopping the program is phased into three periods: organizations not registered on or before May 7, 2026, could not initiate fine-tuning jobs anymore from that date; organizations that had no fine-tuned-models activity on or before July 2, 2026, could not initiate any new jobs from July 2, 2026; all other customers could not initiate fine-tuning jobs from January 6, 2027.
Real-World Use Cases
Fine-tuning is applicable in four different types.
Customer service and chatbots. Fine-tune models to answer questions according to their specific product, policy, and vocabulary rather than generic answers.
Healthcare and legal. Fine-tune models according to their specific vocabulary, which is either medical or legal vocabulary that is rarely found in normal datasets.
Code generation. Fine-tune models according to a particular dataset that includes internal libraries, specific programming languages, and their style guides.
Personalization and recommendations. Fine-tune models for recommendations based on purchase behavior and browsing habits specific to their platform.
Worked example: when fine-tuning was not the full answer
An intermediate-sized legal research company fine-tuned an open-source language model with 40,000 internal case summaries in order to achieve higher answer accuracy for statutory issues. Within two weeks, the fine-tuned model demonstrated better results on the internal accuracy benchmark of the company. After three months, accuracy of questions regarding new statutes fell since knowledge of the fine-tuned model did not change after training. Instead of fine-tuning the model each time there was some legislative change, the company introduced a retrieval layer atop its statute database. Query latency rose slightly. Time-dependent questions were answered accurately without additional training. The result correlates with the decision-making model presented above – fine-tuning was appropriate for the company’s legal reasoning style, while RAG was good for its ever-changing facts.
Benefits of Fine-Tuning
Fine-tuning provides four benefits that are objectively quantifiable when compared to training a new model.
- Lower Cost. Since fine-tuning uses existing weights that were previously trained, it is less computationally expensive than training a new model.
- Less training data. A fine-tuned model can reach high levels of performance on a narrow task using thousands of samples, whereas the pre-trained model needs billions of samples to perform well.
- Faster deployment. The fine-tuning process can be completed in hours or days, while the pre-training process takes weeks or months.
- Higher accuracy on a narrow task. A fine-tuned model will always outperform a pre-trained model on the specific task it was fine-tuned for.
Challenges and Risks of Fine-Tuning
Fine-tuning has four types of risks that are not relevant when working with a pre-trained model.
- Overfitting. Overfitting happens when the fine-tuned model is trained on a limited set of data that does not represent generalized information.
- Catastrophic forgetting. Fine-tuned aggressively, a model can erase all the general knowledge obtained during pre-training.
- Inherited bias. Bias, error, and security risk of a base model get inherited by a fine-tuned model. Inheriting means keeping the same properties even when they get amplified through fine-tuning.
- Static knowledge. Knowledge contained in a fine-tuned model becomes static after training; in order to incorporate new facts, you will need to perform a new round of fine-tuning.
How Much Does Fine-Tuning Cost?
The costs of fine-tuning depend on three variables: the approach, the model size, and whether the training takes place via a hosted API or rental infrastructure.
- Hosted API training for open-source models has ranged between $1 and around $25 per million training tokens in 2026, based on the size of the model and its provider.
- Self-hosted GPU rental training has cost between $ 1 and $ 4 per hour for 8-billion-parameter models, and $ 8 to $ 16 per hour for 70-billion-parameter models, when renting hardware.
- Fine-tuned inference costs usually exceed base-model inference on the same provider. It is still cheaper to use a fine-tuned small model instead of a larger base model when it can substitute it for the same job.
- Training cost scales with epochs. Training a dataset for three epochs costs approximately three times more than training for one epoch, because the per-token cost of training is charged per epoch.
Access to providers has also been modified in 2026. As of May 7, 2026, OpenAI has stopped taking new requests for fine-tuning and currently has only a few models available via its general-purpose fine-tuning platform. There is no fine-tuning API for the Claude models by Anthropic as of 2026; fine-tuning is carried out via enterprise agreements, and earlier via Claude 3 Haiku fine-tuning on Amazon Bedrock.
Best Practices for Fine-Tuning
There are five ways to improve fine-tuning results.
- Prepare clean, labeled data. It is necessary to delete duplicates, standardize the formatting, and make sure that the labeling is correct before starting the fine-tuning process.
- Match the base model to the task. An image model being used for a text task, for instance, will not be able to perform fine-tuning properly.
- Select parameters deliberately. Select which layers should be frozen and whether you should go for full fine-tuning, LoRA, or an adapter.
- Monitor for overfitting during training. Make sure to track how your model is performing on the validation set rather than on the training set.
- Test before full deployment. You should measure performance using metrics and analyze the fine-tuned model’s performance after deployment.
Who Works on Fine-Tuning?
The three main roles that usually do the task of fine-tuning are AI engineers, who build and maintain the infrastructure for the training process; data scientists, who prepare the data and analyze the models; and software developers.
Frequently Asked Questions
Q1. Is fine-tuning still relevant in 2026?
Ans. Yes, but only for those applications where a consistent level of skill, style, or formatting is necessary, given OpenAI’s shutdown of its self-serve fine-tuning capabilities in May 2026 in favor of its better-performing base models.
Q2. How much data do I need to fine-tune a model?
Ans. The amount of data used in fine-tuning datasets varies, but typically goes from a few hundred to tens of thousands of examples depending on the approach used and the dissimilarity of the task compared to that in which the base model was trained.
Q3. Can I fine-tune GPT, Claude, or Gemini myself?
Ans. As of 2026, OpenAI continues to offer fine-tuning for a certain number of models to existing customers, whereas Anthropic lacks a public fine-tuning API for Claude; refer to their current documentation to determine if you can build on either.
Q4. Fine-tuning vs. RAG: Which is cheaper?
Ans. The former will usually be cheaper to implement since it requires no training, while the latter entails a training fee but no retrieval fee per query.
Q5. Does fine-tuning make a model smarter, or just more specialized?
Ans. Fine-tuning won’t make the model more intelligent; rather, it will specialize the model’s existing capabilities towards solving the problem addressed by the fine-tuning dataset.