Python · CUDA/C++ · Gauss-Newton · RANSAC · Factor Graph · Manifold Optimization · Sparse Linear Algebra
cuNLS solves nonlinear least-squares problems on the GPU from Python. The pycunls
package works directly on CuPy arrays, ships built-in manifold states,
factors and robust losses, and lets you write custom factor and state kernels in Python with
NVIDIA Warp. It is built around batched factor evaluation,
sparse Jacobian assembly, and sparse linear solvers — designed for large-scale geometric
estimation workloads such as bundle adjustment, pose graph optimization, and ICP-style
alignment. For problems where many measurements are gross outliers (wrong matches), its GPU
RANSAC minimizers solve the same problems robustly and return the inlier set. The same
solver is available as a CUDA/C++ library with a C++ API for native applications.
cuNLS solves optimization problems of the form:
where
cuNLS refining two large estimation problems, one Gauss-Newton/LM iteration per frame:
- Left — Sphere pose-graph refinement. 200k points that should lie on a sphere are connected by relative ("between") constraints and start as a disturbed blob. cuNLS drives them back onto the sphere; color encodes per-point error, cooling from hot to calm as the solve converges.
-
Right — Kepler orbit fitting. A family of orbits is observed as noisy 3-D points;
cuNLS estimates the five Keplerian elements per orbit (custom NVIDIA Warp factor on an
$\mathbb{R}^5$ state) and a jittered tangle of loops organizes into a crisp nested rosette of tilted ellipses.
On real data from TartanGround, visualized with Rerun:
Visual-inertial odometry with RANSAC. A legged robot walks 82 m through a town; every
frame an ImuFactorBatch and priors stay always on while RansacLevenbergMarquardtMinimizer
classifies the feature matches, of which up to 75% are injected outliers. It drifts 0.9% of
the path; visual-only RANSAC drifts 2.9% and a Huber loss diverges.
A drone fleet in a supermarket. Quadrotors fly deliveries through a store fused from
depth images, among walking shoppers. The whole fleet is one batched MPC problem
(pycunls.mpc), re-solved at every 25 ms control step of the simulation; each drone keeps
clear of the shelves, the shoppers' personal space and the other drones' plans.
See python/examples to run them.
| Category | Details |
|---|---|
| APIs | Python (pycunls, nanobind bindings with first-class CuPy interop; constructors accept a cupy.ndarray or a raw device pointer) and C++ (cunls, shared or static library) |
| Manifold support | SO(2), SO(3), SE(2), SE(3), Sim(2), Sim(3), SL(4), Euclidean vectors |
| Solvers | Gauss-Newton, Levenberg-Marquardt with adaptive damping |
| Robust estimation (RANSAC) | RansacGaussNewtonMinimizer, RansacLevenbergMarquardtMinimizer: hundreds of hypotheses solved in parallel on the GPU from minimal samples, MSAC scoring, adaptive stopping, refinement on the inliers, per-factor inlier mask. Works with any factor and state type (built-in or custom) whose total free dimension is ≤ 64; any number of factors. Deterministic for a fixed seed; see RANSAC |
| Robust losses | Huber, Cauchy, Arctan, SoftL1, Tolerant, Tukey, Scaled |
| Built-in factors | Reprojection, PnP, between (SO(2)/SO(3)/SE(2)/SE(3)/Sim(2)/Sim(3)/SL(4)/vector), point-to-point, point-to-plane, symmetric point-to-plane, prior, constant-velocity/constant-acceleration motion priors (SO(2)/SO(3)/SE(2)/SE(3)) |
| Custom factors and states | CuPy / NVIDIA Warp kernels in Python, or user-defined CUDA kernels via SizedFactorBatch / SizedStateBatch in C++; the same types work with every minimizer, including RANSAC — see Custom factors and states |
| Numeric Jacobians | Finite-difference Jacobians for any factor batch (manifold-aware, reuses each state's Plus retraction), selectable globally (MinimizerOptions.jacobian_mode) or per factor group (the jacobian_mode_override of Problem.add_factor_batch) — write a factor with only a residual and let cuNLS differentiate it; see Numeric Jacobians |
| Linear solver | Block-sparse PCG (variable block-Jacobi preconditioner, default), NVIDIA cuDSS (optional, loaded via dlopen() at runtime — see C++ Installation), dense LDLT, dense Cholesky (cuSOLVER), dense QR (cuSOLVER) |
| Safety checks | Optional runtime validation (linear-solver diagnostics and more), off by default for low-latency solves — enable with MinimizerOptions::disable_safety_checks = false when debugging |
| Execution model | Fully asynchronous via CUDA streams |
pycunls exposes cuNLS to Python via nanobind,
with first-class CuPy interop and optional
NVIDIA Warp support for writing custom
factor kernels in Python. The wheel statically links the cuNLS core library and dynamically
links the CUDA runtime libraries, so it is specific to the CUDA version it was built with.
See Installation for details.
- NVIDIA GPU with compatible driver
- CUDA Toolkit (
nvcc,cudart,cuBLAS,cuSPARSE,cuSOLVER) - CMake >= 3.22
- C++17 compiler
- GNU Make
- Python >= 3.10
Requires Docker with the NVIDIA Container Toolkit.
./scripts/build_pycunls_in_docker.sh [local_output_dir]The output directory defaults to ./dist. The script builds the wheel inside
a container with the source mounted read-only. Intermediate build directories
live inside the container and are discarded; only the final .whl file is
written to the host output directory.
cd python
pip install scikit-build-core nanobind
pip wheel . --no-build-isolation --no-deps --wheel-dir ../distpip install ./dist/pycunls-*.whlFor an editable (in-place) install that reflects source changes without rebuilding:
cd python
pip install scikit-build-core nanobind
pip install -e ".[test]" --no-build-isolationThis installs pycunls along with all test dependencies (pytest,
cupy-cuda12x, warp-lang). Other optional dependency groups:
pip install -e ".[warp]" # warp-lang only
pip install -e ".[all]" # all optional extrasPrerequisites are the same as for the Python package, without Python.
./scripts/build_cunls.sh <build_dir> <Release|Coverage> [install_dir]Example — release build with install:
./scripts/build_cunls.sh build Release /tmp/cunls_install- Install the NVIDIA Container Toolkit.
- Run:
./scripts/build_cunls_in_docker.sh <Release|Coverage> [local_install_dir]The Docker build produces both shared and static variants, installed to separate
directories so their CMake package configs don't overwrite each other. The CMake build
directories are kept in the output directory (needed by test_cunls_in_docker.sh).
Install artifacts (default build_docker/, or the specified directory):
<install_dir>/
shared/ # set CMAKE_PREFIX_PATH here for libcunls.so
include/cunls/ # headers
lib/
libcunls.so # shared library
cmake/cunls/ # CMake package config (shared)
static/ # set CMAKE_PREFIX_PATH here for libcunls.a
include/cunls/ # headers
lib/
libcunls.a # static library (with bundled deps)
cmake/cunls/ # CMake package config (static)
cudss/ # cuDSS headers + libs (optional, dlopen()'d at runtime)
build_shared/ # CMake build directories (used by the C++ tests)
build_static/
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX=/tmp/cunls_install
cmake --build build -j
cmake --install buildBy default this builds a shared library. Pass -DBUILD_SHARED_LIBS=OFF to
build a static library instead.
CUDA compiler and architecture selections supplied by the caller are preserved. Without one,
/usr/local/cuda/bin/nvcc is used when it exists. For example, a native Jetson Orin build can avoid
PTX JIT compatibility requirements by compiling an SM 87 cubin:
cmake -S . -B build -DCMAKE_CUDA_ARCHITECTURES=87-realThe cuDSS archive (CUDSS_PLATFORM=auto) follows the CUDA Toolkit target: linux-x86_64,
linux-aarch64 for Jetson toolkits, or linux-sbsa for Arm server toolkits and CUDA 13 on Jetson.
Toolkits installed without a targets/ directory (e.g. distro packages under /usr) fall back to
linux-x86_64 on x86_64 and require an explicit platform on aarch64.
Set -DCUDSS_PLATFORM explicitly to override it.
Pass -DCUNLS_ENABLE_CUDSS=OFF to build without the cuDSS backend: the cuDSS archive is not
downloaded and SparseLinearSolverType::cuDSS throws std::runtime_error when a solver is created.
All other solvers work as usual, and the C++ tests skip the cuDSS cases.
Important
Capacity vs. active count. Every factor and state batch has two sizes. The constructor takes
the capacity: how many factors / states its device buffers hold, fixed for the batch's
lifetime. The active count starts at zero and is set with set_num_active_factors(n) /
set_num_active_states(n) (C++: SetNumActiveFactors / SetNumActiveStates), any n
up to the capacity. It is a host-only setter, so a real-time application allocates once for its
largest problem and changes the active counts every frame. A solve with nothing active throws. See
Capacity and active count.
Both quick starts solve the same 1-D prior problem: a scalar variable
minimal.py
import cupy as cp
import pycunls
stream = pycunls.CudaStream()
# Initial guess: x = 0. Target observation: o = 2.
state_gpu = cp.array([0.0], dtype=cp.float32)
obs_gpu = cp.array([2.0], dtype=cp.float32)
# Capacity: how many states / factors the buffers hold (fixed per batch).
# Batches start with 0 active entries; the active count is set separately and
# may change between solves up to the capacity. Here every slot is used.
state_batch = pycunls.VectorStateBatch1(state_gpu, 1)
state_batch.set_num_active_states(1)
prior = pycunls.PriorVectorFactorBatch1(obs_gpu, 1)
prior.set_num_active_factors(1)
state_ptrs = [state_batch.state_device_ptr(0)]
problem = pycunls.Problem()
problem.add_state_batch(state_batch)
problem.add_factor_batch(prior, state_ptrs)
minimizer = pycunls.LevenbergMarquardtMinimizer()
summary = minimizer.minimize(stream, problem) # updates state_gpu in place
print(f"Iterations: {summary.num_iterations}")
print(f"Initial cost: {summary.initial_cost}")
print(f"Final cost: {summary.final_cost}")
print(f"Solution x: {cp.asnumpy(state_gpu)}")Run:
python minimal.pySee the Quick Start for a step-by-step walkthrough.
main.cpp
#include <cuda_runtime.h>
#include <iostream>
#include <vector>
#include "cunls/cunls.h"
int main() {
cudaStream_t stream = nullptr;
cudaStreamCreate(&stream);
std::vector<float> h_state = {0.0f};
std::vector<float> h_obs = {2.0f};
cunls::dvector<float> d_state(h_state);
cunls::dvector<float> d_obs(h_obs);
// Capacity: how many states / factors the buffers hold (fixed per batch).
// Batches start with 0 active entries; the active count is set separately and
// may change between solves up to the capacity. Here every slot is used.
const size_t capacity = 1;
const size_t num_states = 1, num_factors = 1;
cunls::VectorStateBatch<1> state_batch(d_state.data(), capacity);
cunls::PriorVectorFactorBatch<1> prior(
reinterpret_cast<const cunls::Vector<1>*>(d_obs.data()), capacity);
state_batch.SetNumActiveStates(num_states); // active count
prior.SetNumActiveFactors(num_factors); // active count
std::vector<float*> state_ptrs = {state_batch.StateDevicePtr(0)};
cunls::Problem problem;
problem.AddStateBatch(&state_batch);
problem.AddFactorBatch(&prior, state_ptrs);
cunls::LevenbergMarquardtMinimizer minimizer;
auto summary = minimizer.Minimize(stream, problem);
std::cout << "Iterations: " << summary.num_iterations << "\n";
std::cout << "Initial cost: " << summary.initial_cost << "\n";
std::cout << "Final cost: " << summary.final_cost << "\n";
cudaStreamDestroy(stream);
return 0;
}To build it against an installed cuNLS, link libcunls and CUDA::cudart as the
C++ examples do.
When a fraction of the measurements are gross outliers, swap the minimizer — the problem stays the same. Mark each residual batch as sampled (may contain outliers, classified with an inlier threshold in residual units) or always-on (trusted priors):
Python
options = pycunls.RansacLevenbergMarquardtMinimizerOptions()
options.base_options.factor_batches = [
pycunls.RansacFactorBatchOptions(pycunls.RansacRole.sampled, 0.01)]
minimizer = pycunls.RansacLevenbergMarquardtMinimizer(options)
summary = minimizer.minimize(stream, problem) # estimate written back
mask = minimizer.inlier_mask(0) # numpy uint8, 1 = inlierC++
cunls::RansacLevenbergMarquardtMinimizerOptions options;
options.base_options.factor_batches = {{cunls::RansacRole::kSampled, /*inlier_threshold=*/0.01f}};
cunls::RansacLevenbergMarquardtMinimizer minimizer(options);
cunls::RansacSummary summary = minimizer.Minimize(stream, problem); // estimate written back
const uint8_t *inliers = minimizer.InlierMask(0); // device, 1 byte per factorOn PnP it recovers the pose at up to 90% outliers, where least squares (even with a Huber loss)
fails, in 1.4 ms for 1,000 correspondences and 15 ms for 1,000,000. See
RANSAC for the theory, options and tuning, and
python/examples/ransac_pnp.py / examples/ransac_pnp
for complete programs.
- Python examples (
pycunls) - C++ examples
Tests require a GPU.
pytest -v python/testsOr in Docker, against a wheel built by build_pycunls_in_docker.sh:
./scripts/test_pycunls_in_docker.sh ./distcmake -S . -B build/tests -DCMAKE_BUILD_TYPE=Release -DBUILD_TESTING=ON
cmake --build build/tests -j
ctest --test-dir build/tests --output-on-failureOr run the test binary directly:
./build/tests/bin/nls_testsOr in Docker, against the output of build_cunls_in_docker.sh:
./scripts/test_cunls_in_docker.sh ./build_dockerCoverage build:
./scripts/build_cunls.sh build/coverage Coveragepython -m pip install -r docs/sphinx/requirements.txt
python -m sphinx -b html docs/sphinx docs/sphinx/_buildOr build in Docker:
bash docs/build_in_docker.sh [output_dir]cuNLS follows the Google C++ Style Guide and uses a pre-commit hook for auto-formatting.
sudo apt install pre-commit
pre-commit installTo manually reformat:
sudo apt install clang-format
find . -iname '*.h' -o -iname '*.cpp' | xargs clang-format -icuNLS is licensed under the Apache License 2.0. Third-party license notices are in NOTICE and third_party/LICENSES/.




