Benchmarks¶
kalman-py is compared with FilterPy in per-step mode (one predict + update call per measurement) and with pykalman in batch mode (the whole sequence in one call), on five frozen scenarios shared with the C, C++ and Rust implementations. Every library gets the same inputs, seeds and precision.
| ID | Scenario | n / m | Filters | Length |
|---|---|---|---|---|
| S1 | 1-D constant velocity | 2 / 1 | KF | 10,000 steps |
| S2 | 2-D constant velocity | 4 / 2 | KF | 10,000 steps |
| S3 | Range-bearing tracking | 4 / 2 | EKF, UKF | 500 steps × 200 seeds |
| S4 | Ill-conditioned problem, float32 | 4 / 2 | KF | 1,000,000 steps |
| S5 | 15-state INS error-state model | 15 / 6 | KF | 10,000 steps |
The scenario files and their format are described in tests/vectors/README.md in the
repository.
Preliminary
These numbers come from a development laptop (Apple M3 Pro), not a dedicated machine with a fixed CPU frequency, so timings can vary by 10–20% between runs.
Speed¶
In batch mode the JAX backend is 93× (S1), 67× (S2) and 9.5× (S5) faster than pykalman. In per-step mode kalman-py is 0.78–0.84× FilterPy's speed: both make the same number of small NumPy calls, and kalman-py also symmetrizes every covariance. Estimates match in every linear case.
Float32 stability (S4)¶
Warning
On this problem the default NumPy configuration (Joseph form) is the first to fail. Use
square_root=True when filtering in float32.
Full results¶
Measured on Apple M3 Pro (arm64), macOS 27.0.1, python-3.12.13, commit 9a2edc7, 2026-10-05.
| Scenario | Filter | Precision | Metric | kalman-py | filterpy | kalman-py-jax | kalman-py-numpy | kalman-py-sqrt | naive | pykalman | Ours vs best other |
|---|---|---|---|---|---|---|---|---|---|---|---|
| S1 | KF batch | float64 | max_abs_diff (state) | n/a | n/a | n/a | n/a | n/a | n/a | 5.99e-12 | n/a |
| S1 | KF batch | float64 | peak_memory (bytes) | n/a | n/a | n/a | 1,524,384 | n/a | n/a | 1,136,154 | n/a |
| S1 | KF batch | float64 | rmse (state) | 0.4233 | n/a | n/a | 0.4233 | n/a | n/a | 0.4233 | 1.00x |
| S1 | KF batch | float64 | time_per_step (ns) | 406.8 | n/a | n/a | 2,606 | n/a | n/a | 38,047 | 93.53x |
| S1 | KF per-step | float64 | max_abs_diff (state) | n/a | 1.78e-15 | n/a | n/a | n/a | n/a | n/a | n/a |
| S1 | KF per-step | float64 | peak_memory (bytes) | 164,260 | 163,912 | n/a | n/a | n/a | n/a | n/a | 1.00x |
| S1 | KF per-step | float64 | rmse (state) | 0.4233 | 0.4233 | n/a | n/a | n/a | n/a | n/a | 1.00x |
| S1 | KF per-step | float64 | time_per_step (ns) | 13,252 | 10,320 | n/a | n/a | n/a | n/a | n/a | 0.78x |
| S2 | KF batch | float64 | max_abs_diff (state) | n/a | n/a | n/a | n/a | n/a | n/a | 3.32e-12 | n/a |
| S2 | KF batch | float64 | peak_memory (bytes) | n/a | n/a | n/a | 4,165,368 | n/a | n/a | 3,866,229 | n/a |
| S2 | KF batch | float64 | rmse (state) | 0.5073 | n/a | n/a | 0.5073 | n/a | n/a | 0.5073 | 1.00x |
| S2 | KF batch | float64 | time_per_step (ns) | 577.5 | n/a | n/a | 2,786 | n/a | n/a | 38,978 | 67.49x |
| S2 | KF per-step | float64 | max_abs_diff (state) | n/a | 3.55e-15 | n/a | n/a | n/a | n/a | n/a | n/a |
| S2 | KF per-step | float64 | peak_memory (bytes) | 325,324 | 324,632 | n/a | n/a | n/a | n/a | n/a | 1.00x |
| S2 | KF per-step | float64 | rmse (state) | 0.5073 | 0.5073 | n/a | n/a | n/a | n/a | n/a | 1.00x |
| S2 | KF per-step | float64 | time_per_step (ns) | 13,442 | 10,556 | n/a | n/a | n/a | n/a | n/a | 0.79x |
| S3 | EKF | float64 | nees (-) | 4.002 | 4.002 | n/a | n/a | n/a | n/a | n/a | – |
| S3 | EKF | float64 | rmse (state) | 0.8896 | 0.8896 | n/a | n/a | n/a | n/a | n/a | 1.00x |
| S3 | UKF | float64 | nees (-) | 4.001 | 3.979 | n/a | n/a | n/a | n/a | n/a | – |
| S3 | UKF | float64 | rmse (state) | 0.8896 | 0.8896 | n/a | n/a | n/a | n/a | n/a | 1.00x |
| S4 | KF | float32 | steps_to_failure (steps) | 2,753 | n/a | 1,000,000 | n/a | 1,000,000 | 67,589 | n/a | 0.04x |
| S4 | KF | float32 | steps_to_indefinite (steps) | 2,753 | n/a | 1,000,000 | n/a | 1,000,000 | 1,000,000 | n/a | 0.0028x |
| S5 | KF batch | float64 | peak_memory (bytes) | n/a | n/a | n/a | 42,821,472 | n/a | n/a | 45,672,193 | n/a |
| S5 | KF batch | float64 | rmse (state) | 0.1865 | n/a | n/a | 0.1865 | n/a | n/a | 0.1865 | 1.00x |
| S5 | KF batch | float64 | time_per_step (ns) | 5,535 | n/a | n/a | 23,730 | n/a | n/a | 52,290 | 9.45x |
| S5 | KF per-step | float64 | peak_memory (bytes) | 1,222,492 | 1,220,296 | n/a | n/a | n/a | n/a | n/a | 1.00x |
| S5 | KF per-step | float64 | rmse (state) | 0.1865 | 0.1865 | n/a | n/a | n/a | n/a | n/a | 1.00x |
| S5 | KF per-step | float64 | time_per_step (ns) | 22,639 | 18,923 | n/a | n/a | n/a | n/a | n/a | 0.84x |
- KF per-step rows:
kalman-pyis the step-by-step API (NumPy). KF batch rows:kalman-pyis the JAX backend on CPU andkalman-py-numpythe NumPy backend; JAX has no peak_memory because tracemalloc can't see its buffers. max_abs_diff is the largest difference betweenkalman-pyand that library's estimates. - S3 UKF: kalman-py redraws sigma points from the predicted distribution before each update, while FilterPy reuses the propagated ones, so their estimates differ slightly; the EKFs agree to ~1e-13. A consistent filter has average NEES = 4.
- S4 (float32): FilterPy is n/a because it converts float32 inputs to float64 internally (its identity matrix is float64), so it can't run S4 as specified.
- S4:
kalman-pyis the default configuration (NumPy backend, Joseph form);kalman-py-sqrtusessquare_root=True;kalman-py-jaxis the JAX backend. 1,000,000 means the filter never failed.
Reproduce¶
ENV="$(git rev-parse --short HEAD),<cpu>,<os>,python-<version>,$(date +%F)"
uv run python benchmarks/run.py "$ENV"
uv run python benchmarks/accuracy.py "$ENV"
uv run python scripts/make_table.py results/results.csv kalman-py --readme README.md
uv run python scripts/plot_benchmarks.py