Last modified: Oct 04, 2026
Download Hugging Face Model in Python
Hugging Face Hub is a central place for thousands of AI models. You can download any model directly in Python. This guide shows you exactly how to do it.
We will cover the transformers library, the huggingface_hub library, and manual downloads. You will also see how to cache models and avoid common errors.
Why Download Models in Python?
Python is the main language for machine learning. Downloading a model in Python lets you use it immediately for inference or fine-tuning.
The Hugging Face Hub provides a simple API. You do not need to manually download files from the website. Everything happens inside your script.
This saves time and keeps your project reproducible. You can share your code and others can download the same model with one line.
Method 1: Using the Transformers Library
The transformers library is the easiest way. It automatically downloads the model and tokenizer.
First, install the library if you have not already.
pip install transformers
Then use the from_pretrained method. This method downloads the model from the Hub and loads it into memory.
from transformers import AutoModel, AutoTokenizer
# Choose a model name from Hugging Face Hub
model_name = "bert-base-uncased"
# Download and load the tokenizer
tokenizer = AutoTokenizer.from_pretrained(model_name)
# Download and load the model
model = AutoModel.from_pretrained(model_name)
print("Model and tokenizer downloaded successfully!")
When you run this code, you will see a progress bar. The files are saved to a local cache directory.
Downloading config.json: 100%|██████████| 570/570 [00:00<00:00, 2.1MB/s]
Downloading pytorch_model.bin: 100%|██████████| 440M/440M [00:15<00:00, 28.5MB/s]
Downloading vocab.txt: 100%|██████████| 232k/232k [00:00<00:00, 1.2MB/s]
Model and tokenizer downloaded successfully!
The default cache folder is ~/.cache/huggingface/hub. You can change it with the cache_dir argument.
Method 2: Using the Hugging Face Hub Library
The huggingface_hub library gives you more control. You can download specific files or the entire repository.
Install it with pip.
pip install huggingface_hub
Use the hf_hub_download function to download a single file.
from huggingface_hub import hf_hub_download
# Download only the model weights file
file_path = hf_hub_download(
repo_id="bert-base-uncased",
filename="pytorch_model.bin"
)
print(f"File downloaded to: {file_path}")
Output:
File downloaded to: /home/user/.cache/huggingface/hub/models--bert-base-uncased/snapshots/.../pytorch_model.bin
You can also download the whole repository using snapshot_download. This is useful when you need all files.
from huggingface_hub import snapshot_download
# Download all files for a model
local_dir = snapshot_download(
repo_id="bert-base-uncased",
local_dir="./my_bert_model"
)
print(f"All files downloaded to: {local_dir}")
This creates a folder called my_bert_model in your current directory. It contains config, weights, and tokenizer files.
Method 3: Manual Download with Requests
Sometimes you want full control. You can use the requests library to download files directly.
This method is not recommended for large models. It does not handle caching or progress bars well.
import requests
url = "https://huggingface.co/bert-base-uncased/resolve/main/config.json"
response = requests.get(url)
with open("config.json", "wb") as f:
f.write(response.content)
print("Config file downloaded manually.")
Output:
Config file downloaded manually.
This works for small files. For large model weights, use the official libraries instead.
Understanding the Cache Directory
Hugging Face stores downloaded models in a cache. This avoids downloading the same model twice.
The default location is ~/.cache/huggingface/hub on Linux and macOS. On Windows, it is C:\Users\username\.cache\huggingface\hub.
You can set the environment variable HF_HOME to change the cache root. Or pass cache_dir to the download function.
from transformers import AutoModel
# Use a custom cache directory
model = AutoModel.from_pretrained(
"bert-base-uncased",
cache_dir="./my_custom_cache"
)
This keeps your models organized. It also helps when you have limited disk space on the default drive.
Downloading Private or Gated Models
Some models are private or gated. You need an access token from Hugging Face.
First, log in using the CLI or Python. The login function saves your token.
from huggingface_hub import login
# Paste your token when prompted
login()
Then download the model normally. The library uses your token automatically.
from transformers import AutoModel
# Download a gated model like meta-llama/Llama-2-7b
model = AutoModel.from_pretrained("meta-llama/Llama-2-7b")
If you prefer not to log in interactively, pass the token directly.
from huggingface_hub import hf_hub_download
file_path = hf_hub_download(
repo_id="meta-llama/Llama-2-7b",
filename="config.json",
token="hf_your_token_here"
)
Always keep your token secret. Do not share it in public code repositories.
Common Errors and How to Fix Them
Error: OSError: Can't load config for 'model-name'
This means the model name is wrong or the model does not exist. Double-check the spelling on the Hugging Face website.
Error: ConnectionError: Couldn't reach server
Your internet may be down or the Hub is temporarily unavailable. Check your connection and try again.
Error: HTTPError: 401 Client Error
You are trying to access a gated model without a token. Log in or pass your token as shown above.
Error: Disk quota exceeded
Your cache drive is full. Delete old models or change the cache directory to a larger drive.
Best Practices for Downloading Models
Always use the transformers or huggingface_hub libraries. They handle retries, caching, and progress bars.
Pin your model version using a revision hash. This makes your code reproducible.
from transformers import AutoModel
# Use a specific commit hash for reproducibility
model = AutoModel.from_pretrained(
"bert-base-uncased",
revision="a1b2c3d4e5f6"
)
Use snapshot_download when you need the entire model folder. Use hf_hub_download for individual files.
Clean your cache regularly. Old models can take up hundreds of gigabytes.
Conclusion
Downloading a model from Hugging Face Hub in Python is simple. The transformers library is the fastest way for most users.
For more control, use the huggingface_hub library. It lets you download specific files or entire repositories.
Remember to handle private models with tokens. Also, watch your cache directory to avoid disk space issues.
Now you can start using thousands of pre-trained models in your Python projects. Happy coding!