Last modified: Oct 06, 2026
Numba vs Cython: Performance Compared
Python is easy to write but slow to run. For numerical code, two tools change that: Numba and Cython. Both compile Python-like code to machine code. But they work in very different ways. This article compares their performance, syntax, and best use cases.
What Is Numba?
Numba is a just-in-time (JIT) compiler. It translates Python functions to machine code at runtime. You add a decorator like @njit to a function. Numba then compiles it on the first call.
Numba works best with NumPy arrays and numerical loops. It understands NumPy operations. It can also run code in parallel with parallel=True.
Here is a simple Numba example:
from numba import njit
import numpy as np
@njit
def sum_array(arr):
total = 0.0
for i in range(arr.shape[0]):
total += arr[i]
return total
# Create a large array
data = np.random.rand(10_000_000)
# First call triggers compilation
result = sum_array(data)
print(result)
5000006.238
The first call includes compilation time. Later calls are very fast. Numba caches compiled code so you do not pay the cost again.
What Is Cython?
Cython is an ahead-of-time (AOT) compiler. It translates Python-like code to C or C++. You then compile that C code into a Python extension module.
Cython uses static type declarations with cdef. This gives the compiler more information. The result is code that runs at C speed.
Here is a Cython example in a .pyx file:
# sum_array.pyx
import numpy as np
cimport numpy as cnp
cnp.import_array()
def sum_array(cnp.ndarray[double, ndim=1] arr):
cdef double total = 0.0
cdef int i
cdef int n = arr.shape[0]
for i in range(n):
total += arr[i]
return total
You compile it with a setup script:
# setup.py
from setuptools import setup
from Cython.Build import cythonize
import numpy as np
setup(
ext_modules=cythonize("sum_array.pyx"),
include_dirs=[np.get_include()]
)
Then run python setup.py build_ext --inplace. The module is ready to import.
Performance Benchmarks
Benchmarks show a clear pattern. Cython can reach the highest raw speed. Numba gets very close with far less effort.
In a recent N-body benchmark, Cython ran in 10 ms with a 124x speedup over CPython. Numba ran in 22 ms with a 56x speedup[reference:0]. Cython was faster, but it required C knowledge and careful type declarations.
In a matrix-vector benchmark, Numba was slightly faster than Cython: 104 ms vs 142 ms[reference:1]. This shows that results depend on the problem.
For small arrays, Numba often wins. For large arrays, Cython can take the lead. One Stack Overflow comparison found Cython was 6 times slower for small dimensions but 4 times faster for large ones[reference:2].
The key takeaway: both tools deliver large speedups. The winner depends on your specific workload.
Ease of Use and Syntax
Numba is much easier to start with. You take existing Python code and add @njit. There is no separate compilation step. There is no new syntax to learn.
Cython requires more work. You write code in .pyx files. You declare types with cdef. You write a setup script. You compile the extension. Each step adds complexity.
One source puts it simply: Numba code is "very similar to Python code" and "a lot easier to maintain than Cython code"[reference:3]. Cython should only be used when Numba cannot solve the problem.
However, Cython gives you more control. You can call C libraries directly. You can manage memory manually. You can create C extension types with cdef class. Numba does not offer this level of control.
Compilation Time and Overhead
Numba compiles at runtime. The first call to a compiled function can be slow. This is the compilation overhead. For long-running programs, this cost is negligible. For short scripts, it can dominate.
Numba offers a cache option. Compiled code is saved to disk. Later runs skip compilation. This removes the overhead for repeated use.
Cython compiles ahead of time. There is no runtime overhead. The build step takes time, but the final module starts instantly. This is better for distributed systems and production deployments.
Numba's JIT compilation is "1.6x to 3.7x slower than C/C++ ahead-of-time compilation" in terms of compile time[reference:4]. But this only matters during the first run.
GPU and Parallel Support
Numba supports GPU acceleration through CUDA. You can write GPU kernels with @cuda.jit. This works only with NVIDIA GPUs.
Numba also supports automatic parallelization. Add parallel=True to @njit and use prange for loops. The compiler handles thread management.
Cython supports parallelism through OpenMP. You enable it with compiler flags. You use prange with cython.parallel. This requires more setup than Numba.
Cython can also release the GIL with with nogil:. This enables true thread-level parallelism. Numba handles the GIL automatically for compiled functions.
For GPU work, Numba is the easier path. For CPU parallelism, both are capable. Numba requires less code.
When to Use Numba
Use Numba when you want quick speedups with minimal code changes. It is ideal for:
• Numerical loops over NumPy arrays.
• Prototyping and research code.
• Code that must stay readable and maintainable.
• GPU acceleration on NVIDIA hardware.
Numba is the first choice for most scientific Python code. The Debian packaging guide for HyperSpy states: "If you need to improve the speed of a given part of the code your first choice should be Numba"[reference:5].
When to Use Cython
Use Cython when you need maximum control and performance. It is ideal for:
• Distributing compiled code to users without a compiler.
• Calling C or C++ libraries directly.
• Writing C extension types.
• Code that cannot be expressed in Numba's supported subset.
Cython should be considered "only if it is not possible to speed up the function using Numba"[reference:6]. It requires more effort but rewards you with C-level performance and flexibility.
Code Example: Side-by-Side
Here is the same function in both tools. It computes the Euclidean distance between two vectors.
Numba version:
from numba import njit
import numpy as np
@njit
def euclidean_distance(a, b):
total = 0.0
for i in range(a.shape[0]):
diff = a[i] - b[i]
total += diff * diff
return np.sqrt(total)
x = np.random.rand(1_000_000)
y = np.random.rand(1_000_000)
print(euclidean_distance(x, y))
Cython version:
# distance.pyx
import numpy as np
cimport numpy as cnp
from libc.math cimport sqrt
cnp.import_array()
def euclidean_distance(cnp.ndarray[double, ndim=1] a,
cnp.ndarray[double, ndim=1] b):
cdef double total = 0.0
cdef double diff
cdef int i
cdef int n = a.shape[0]
for i in range(n):
diff = a[i] - b[i]
total += diff * diff
return sqrt(total)
The Numba version needs one decorator. The Cython version needs type declarations, a cimport, and a build step. Both produce fast machine code.
Conclusion
Numba and Cython both solve Python's performance problem. They take different paths.
Numba is the easier and faster-to-write option. Add a decorator and get large speedups. It is perfect for numerical Python and research code.
Cython is the more powerful and flexible option. It reaches peak performance and integrates with C. But it demands more knowledge and effort.
For most developers, start with Numba. If you hit a wall, switch to Cython. Both tools are excellent. The best choice depends on your project's needs and your team's skills.