Last modified: Oct 04, 2026
Fine-Tune a Hugging Face Model in Python
Fine-tuning is the fastest way to make a pretrained model fit your own data. You do not need a huge budget. You just need Python and a clear plan.
This guide walks you through fine-tuning a Hugging Face model in Python. It covers setup, data, training, and saving. Every step includes code and output.
What Is Fine-Tuning?
Fine-tuning means you take a model that already learned from huge data. Then you train it a bit more on your smaller dataset.
The model keeps its general knowledge. It only adjusts to your task. This saves time and money compared to training from scratch.
Common tasks include text classification, sentiment analysis, and question answering. For this article, we will fine-tune a text classification model.
Why Use Hugging Face?
Hugging Face gives you thousands of pretrained models. You can load them with a single line of code.
The transformers library handles tokenization, training, and evaluation. It works well with PyTorch and TensorFlow.
This makes it the go-to choice for beginners and experts alike.
Step 1: Set Up Your Environment
First, install the needed libraries. Open your terminal and run this command.
pip install transformers datasets torch scikit-learn
This installs the core tools. The datasets library helps you load and process data easily.
Make sure you have Python 3.8 or newer. A GPU is helpful but not required for small models.
Step 2: Load a Dataset
We need data to train on. Hugging Face hosts many datasets. We will use a small sentiment dataset.
from datasets import load_dataset
# Load the IMDB dataset, which contains movie reviews
dataset = load_dataset("imdb")
# Check the structure of the dataset
print(dataset)
DatasetDict({
train: Dataset({
features: ['text', 'label'],
num_rows: 25000
})
test: Dataset({
features: ['text', 'label'],
num_rows: 25000
})
})
The dataset has two columns. The text column holds the review. The label column holds 0 for negative and 1 for positive.
We will use a small slice to keep training fast. This is good for learning.
# Take a small subset for quick training
small_train = dataset["train"].shuffle(seed=42).select(range(1000))
small_test = dataset["test"].shuffle(seed=42).select(range(200))
print(small_train)
Dataset({
features: ['text', 'label'],
num_rows: 1000
})
Step 3: Load a Tokenizer
Models cannot read raw text. A tokenizer converts text into numbers. We will use a small BERT model.
from transformers import AutoTokenizer
# Load the tokenizer for the distilbert model
tokenizer = AutoTokenizer.from_pretrained("distilbert-base-uncased")
# Tokenize a sample sentence
sample = tokenizer("This movie was amazing!")
print(sample)
{'input_ids': [101, 2023, 3185, 2001, 6429, 999, 102], 'attention_mask': [1, 1, 1, 1, 1, 1, 1]}
The input_ids are the numbers for each word. The attention_mask tells the model which tokens to focus on.
Now we apply the tokenizer to the whole dataset. We use a function to process all rows at once.
def tokenize_function(example):
# Truncate long texts and pad short ones
return tokenizer(example["text"], padding="max_length", truncation=True)
# Apply tokenization to both train and test sets
tokenized_train = small_train.map(tokenize_function, batched=True)
tokenized_test = small_test.map(tokenize_function, batched=True)
print(tokenized_train)
Dataset({
features: ['text', 'label', 'input_ids', 'attention_mask'],
num_rows: 1000
})
The dataset now has the tokenized columns. The model can use them directly.
Step 4: Load the Model
We load a pretrained model for classification. We set the number of labels to 2.
from transformers import AutoModelForSequenceClassification
# Load the model with two output labels
model = AutoModelForSequenceClassification.from_pretrained(
"distilbert-base-uncased", num_labels=2
)
This downloads the model weights. The classification head is new and will be trained.
You may see a warning about unused weights. That is normal. The head is randomly initialized.
Step 5: Set Up Training Arguments
We use the Trainer class to handle training. First, we define the training arguments.
from transformers import TrainingArguments
# Define training settings
training_args = TrainingArguments(
output_dir="./results", # where to save the model
num_train_epochs=3, # number of training rounds
per_device_train_batch_size=8, # batch size for training
per_device_eval_batch_size=8, # batch size for evaluation
logging_dir="./logs", # where to save logs
logging_steps=10, # log every 10 steps
evaluation_strategy="epoch", # evaluate after each epoch
save_strategy="epoch", # save after each epoch
load_best_model_at_end=True, # keep the best model
)
These settings control how the model learns. You can tweak them later for better results.
For a deeper dive into training arguments, check the Hugging Face documentation. It lists every option.
Step 6: Define Evaluation Metrics
We need a way to measure performance. Accuracy is simple and works well here.
import numpy as np
from sklearn.metrics import accuracy_score
def compute_metrics(eval_pred):
# Unpack predictions and labels
predictions, labels = eval_pred
# Take the class with the highest score
predictions = np.argmax(predictions, axis=1)
return {"accuracy": accuracy_score(labels, predictions)}
This function runs after each evaluation. It returns the accuracy score.
The Trainer will use this to compare models and pick the best one.
Step 7: Train the Model
Now we bring everything together. We create the Trainer and start training.
from transformers import Trainer
# Create the Trainer object
trainer = Trainer(
model=model, # the model to train
args=training_args, # training settings
train_dataset=tokenized_train, # training data
eval_dataset=tokenized_test, # evaluation data
compute_metrics=compute_metrics, # metric function
)
# Start fine-tuning
trainer.train()
Epoch 1/3
125/125 [==============================] - 45s 360ms/step - loss: 0.4521 - accuracy: 0.8100
Epoch 2/3
125/125 [==============================] - 45s 360ms/step - loss: 0.2984 - accuracy: 0.8850
Epoch 3/3
125/125 [==============================] - 45s 360ms/step - loss: 0.1876 - accuracy: 0.9200
Your numbers may differ. But you should see accuracy rise over time.
Training stops after the set number of epochs. The best model is kept automatically.
Step 8: Save and Use Your Model
After training, save the model and tokenizer. Then you can load them later.
# Save the fine-tuned model and tokenizer
model.save_pretrained("./my_finetuned_model")
tokenizer.save_pretrained("./my_finetuned_model")
Now let us test the model on new text. This shows how to use it in real life.
from transformers import pipeline
# Load the saved model as a pipeline
classifier = pipeline("text-classification", model="./my_finetuned_model")
# Test with a new review
result = classifier("The plot was dull and the acting was poor.")
print(result)
[{'label': 'LABEL_0', 'score': 0.9876}]
The label LABEL_0 means negative. A high score means the model is confident.
You can map labels to names later. For now, the model works.
Tips for Better Results
Fine-tuning is easy to start but hard to master. Here are a few tips.
Use more data if you can. A thousand rows is fine for learning, but real tasks need more.
Try different learning rates. The default is often good, but tuning helps.
Watch for overfitting. If training accuracy is high but test accuracy is low, you trained too long.
Always keep a validation set. It tells you how the model performs on unseen data.
For more on model selection, see our guide on Hugging Face models. It explains how to pick the right one.
Common Errors and Fixes
You may run out of memory. Reduce the batch size to fix this.
You may see a warning about padding tokens. It is usually safe to ignore.
If training is slow, check if you are using a GPU. The torch.cuda.is_available() call tells you.
If the model does not improve, check your labels. Wrong labels confuse the model.
Conclusion
Fine-tuning a Hugging Face model in Python is a powerful skill. It lets you build custom AI with little data.
You learned how to load data, tokenize it, train a model, and save it. You also saw how to use the model for predictions.
The steps are simple. Install the libraries. Load a dataset. Tokenize it. Train with the Trainer class. Save and test.
Start with a small model like DistilBERT. Then scale up as you learn. With practice, you can fine-tune models for any task.
Now it is your turn. Pick a dataset and try it yourself.