Skip to content
nvidia-isaacPublic

About

Nvidia Isaac CUDA-Accelerated Nonlinear Least Square Solver

Resources

Contributing

Security policy

Stars

31 stars

Watchers

0 watching

Forks

Repository files navigation

cuNLS logo

GPU-Accelerated Nonlinear Least-Squares Solver

Python · CUDA/C++ · Gauss-Newton · RANSAC · Factor Graph · Manifold Optimization · Sparse Linear Algebra


cuNLS solves nonlinear least-squares problems on the GPU from Python. The pycunls package works directly on CuPy arrays, ships built-in manifold states, factors and robust losses, and lets you write custom factor and state kernels in Python with NVIDIA Warp. It is built around batched factor evaluation, sparse Jacobian assembly, and sparse linear solvers — designed for large-scale geometric estimation workloads such as bundle adjustment, pose graph optimization, and ICP-style alignment. For problems where many measurements are gross outliers (wrong matches), its GPU RANSAC minimizers solve the same problems robustly and return the inlier set. The same solver is available as a CUDA/C++ library with a C++ API for native applications.

cuNLS solves optimization problems of the form:

$$x^* = \arg\min_x \sum_i \rho_i!\left(\left|f_i(x)\right|^2_{\Sigma_i}\right)$$

where $x$ is the optimization variable (often on a manifold), $f_i(x)$ are residual functions, $\rho_i(\cdot)$ are optional robust loss functions, and $\left|v\right|^2_{\Sigma} = v^T \Sigma^{-1} v$ is the Mahalanobis norm.

Gallery

cuNLS refining two large estimation problems, one Gauss-Newton/LM iteration per frame:

Sphere pose-graph refinement   Kepler orbit fitting

  • Left — Sphere pose-graph refinement. 200k points that should lie on a sphere are connected by relative ("between") constraints and start as a disturbed blob. cuNLS drives them back onto the sphere; color encodes per-point error, cooling from hot to calm as the solve converges.
  • Right — Kepler orbit fitting. A family of orbits is observed as noisy 3-D points; cuNLS estimates the five Keplerian elements per orbit (custom NVIDIA Warp factor on an $\mathbb{R}^5$ state) and a jittered tangle of loops organizes into a crisp nested rosette of tilted ellipses.

On real data from TartanGround, visualized with Rerun:

Visual-inertial odometry with RANSAC

Visual-inertial odometry with RANSAC. A legged robot walks 82 m through a town; every frame an ImuFactorBatch and priors stay always on while RansacLevenbergMarquardtMinimizer classifies the feature matches, of which up to 75% are injected outliers. It drifts 0.9% of the path; visual-only RANSAC drifts 2.9% and a Huber loss diverges.

Drone fleet MPC in a supermarket

A drone fleet in a supermarket. Quadrotors fly deliveries through a store fused from depth images, among walking shoppers. The whole fleet is one batched MPC problem (pycunls.mpc), re-solved at every 25 ms control step of the simulation; each drone keeps clear of the shelves, the shoppers' personal space and the other drones' plans.

See python/examples to run them.

Features

Category Details
APIs Python (pycunls, nanobind bindings with first-class CuPy interop; constructors accept a cupy.ndarray or a raw device pointer) and C++ (cunls, shared or static library)
Manifold support SO(2), SO(3), SE(2), SE(3), Sim(2), Sim(3), SL(4), Euclidean vectors
Solvers Gauss-Newton, Levenberg-Marquardt with adaptive damping
Robust estimation (RANSAC) RansacGaussNewtonMinimizer, RansacLevenbergMarquardtMinimizer: hundreds of hypotheses solved in parallel on the GPU from minimal samples, MSAC scoring, adaptive stopping, refinement on the inliers, per-factor inlier mask. Works with any factor and state type (built-in or custom) whose total free dimension is ≤ 64; any number of factors. Deterministic for a fixed seed; see RANSAC
Robust losses Huber, Cauchy, Arctan, SoftL1, Tolerant, Tukey, Scaled
Built-in factors Reprojection, PnP, between (SO(2)/SO(3)/SE(2)/SE(3)/Sim(2)/Sim(3)/SL(4)/vector), point-to-point, point-to-plane, symmetric point-to-plane, prior, constant-velocity/constant-acceleration motion priors (SO(2)/SO(3)/SE(2)/SE(3))
Custom factors and states CuPy / NVIDIA Warp kernels in Python, or user-defined CUDA kernels via SizedFactorBatch / SizedStateBatch in C++; the same types work with every minimizer, including RANSAC — see Custom factors and states
Numeric Jacobians Finite-difference Jacobians for any factor batch (manifold-aware, reuses each state's Plus retraction), selectable globally (MinimizerOptions.jacobian_mode) or per factor group (the jacobian_mode_override of Problem.add_factor_batch) — write a factor with only a residual and let cuNLS differentiate it; see Numeric Jacobians
Linear solver Block-sparse PCG (variable block-Jacobi preconditioner, default), NVIDIA cuDSS (optional, loaded via dlopen() at runtime — see C++ Installation), dense LDLT, dense Cholesky (cuSOLVER), dense QR (cuSOLVER)
Safety checks Optional runtime validation (linear-solver diagnostics and more), off by default for low-latency solves — enable with MinimizerOptions::disable_safety_checks = false when debugging
Execution model Fully asynchronous via CUDA streams

Installation

Python package (pycunls)

pycunls exposes cuNLS to Python via nanobind, with first-class CuPy interop and optional NVIDIA Warp support for writing custom factor kernels in Python. The wheel statically links the cuNLS core library and dynamically links the CUDA runtime libraries, so it is specific to the CUDA version it was built with. See Installation for details.

Prerequisites

  • NVIDIA GPU with compatible driver
  • CUDA Toolkit (nvcc, cudart, cuBLAS, cuSPARSE, cuSOLVER)
  • CMake >= 3.22
  • C++17 compiler
  • GNU Make
  • Python >= 3.10

Build the wheel in Docker

Requires Docker with the NVIDIA Container Toolkit.

./scripts/build_pycunls_in_docker.sh [local_output_dir]

The output directory defaults to ./dist. The script builds the wheel inside a container with the source mounted read-only. Intermediate build directories live inside the container and are discarded; only the final .whl file is written to the host output directory.

Build the wheel locally

cd python
pip install scikit-build-core nanobind
pip wheel . --no-build-isolation --no-deps --wheel-dir ../dist

Install the wheel

pip install ./dist/pycunls-*.whl

Editable install for development

For an editable (in-place) install that reflects source changes without rebuilding:

cd python
pip install scikit-build-core nanobind
pip install -e ".[test]" --no-build-isolation

This installs pycunls along with all test dependencies (pytest, cupy-cuda12x, warp-lang). Other optional dependency groups:

pip install -e ".[warp]"   # warp-lang only
pip install -e ".[all]"    # all optional extras

C++ library

Prerequisites are the same as for the Python package, without Python.

Build locally

./scripts/build_cunls.sh <build_dir> <Release|Coverage> [install_dir]

Example — release build with install:

./scripts/build_cunls.sh build Release /tmp/cunls_install

Build with Docker

  1. Install the NVIDIA Container Toolkit.
  2. Run:
./scripts/build_cunls_in_docker.sh <Release|Coverage> [local_install_dir]

The Docker build produces both shared and static variants, installed to separate directories so their CMake package configs don't overwrite each other. The CMake build directories are kept in the output directory (needed by test_cunls_in_docker.sh).

Install artifacts (default build_docker/, or the specified directory):

<install_dir>/
  shared/               # set CMAKE_PREFIX_PATH here for libcunls.so
    include/cunls/      # headers
    lib/
      libcunls.so       # shared library
      cmake/cunls/      # CMake package config (shared)
  static/               # set CMAKE_PREFIX_PATH here for libcunls.a
    include/cunls/      # headers
    lib/
      libcunls.a        # static library (with bundled deps)
      cmake/cunls/      # CMake package config (static)
  cudss/                # cuDSS headers + libs (optional, dlopen()'d at runtime)
  build_shared/         # CMake build directories (used by the C++ tests)
  build_static/

Direct CMake build

cmake -S . -B build -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX=/tmp/cunls_install
cmake --build build -j
cmake --install build

By default this builds a shared library. Pass -DBUILD_SHARED_LIBS=OFF to build a static library instead.

CUDA compiler and architecture selections supplied by the caller are preserved. Without one, /usr/local/cuda/bin/nvcc is used when it exists. For example, a native Jetson Orin build can avoid PTX JIT compatibility requirements by compiling an SM 87 cubin:

cmake -S . -B build -DCMAKE_CUDA_ARCHITECTURES=87-real

The cuDSS archive (CUDSS_PLATFORM=auto) follows the CUDA Toolkit target: linux-x86_64, linux-aarch64 for Jetson toolkits, or linux-sbsa for Arm server toolkits and CUDA 13 on Jetson. Toolkits installed without a targets/ directory (e.g. distro packages under /usr) fall back to linux-x86_64 on x86_64 and require an explicit platform on aarch64. Set -DCUDSS_PLATFORM explicitly to override it.

Pass -DCUNLS_ENABLE_CUDSS=OFF to build without the cuDSS backend: the cuDSS archive is not downloaded and SparseLinearSolverType::cuDSS throws std::runtime_error when a solver is created. All other solvers work as usual, and the C++ tests skip the cuDSS cases.

Quick Start

Important

Capacity vs. active count. Every factor and state batch has two sizes. The constructor takes the capacity: how many factors / states its device buffers hold, fixed for the batch's lifetime. The active count starts at zero and is set with set_num_active_factors(n) / set_num_active_states(n) (C++: SetNumActiveFactors / SetNumActiveStates), any n up to the capacity. It is a host-only setter, so a real-time application allocates once for its largest problem and changes the active counts every frame. A solve with nothing active throws. See Capacity and active count.

Both quick starts solve the same 1-D prior problem: a scalar variable $x$ pulled toward a target $o = 2$ by a prior factor with residual $r = x - o$, so the optimum is $x^* = 2$.

Python

minimal.py

import cupy as cp
import pycunls

stream = pycunls.CudaStream()

# Initial guess: x = 0.  Target observation: o = 2.
state_gpu = cp.array([0.0], dtype=cp.float32)
obs_gpu   = cp.array([2.0], dtype=cp.float32)

# Capacity: how many states / factors the buffers hold (fixed per batch).
# Batches start with 0 active entries; the active count is set separately and
# may change between solves up to the capacity. Here every slot is used.
state_batch = pycunls.VectorStateBatch1(state_gpu, 1)
state_batch.set_num_active_states(1)

prior = pycunls.PriorVectorFactorBatch1(obs_gpu, 1)
prior.set_num_active_factors(1)

state_ptrs = [state_batch.state_device_ptr(0)]

problem = pycunls.Problem()
problem.add_state_batch(state_batch)
problem.add_factor_batch(prior, state_ptrs)

minimizer = pycunls.LevenbergMarquardtMinimizer()
summary   = minimizer.minimize(stream, problem)  # updates state_gpu in place

print(f"Iterations:   {summary.num_iterations}")
print(f"Initial cost: {summary.initial_cost}")
print(f"Final cost:   {summary.final_cost}")
print(f"Solution x:   {cp.asnumpy(state_gpu)}")

Run:

python minimal.py

See the Quick Start for a step-by-step walkthrough.

C++

main.cpp

#include <cuda_runtime.h>
#include <iostream>
#include <vector>
#include "cunls/cunls.h"

int main() {
  cudaStream_t stream = nullptr;
  cudaStreamCreate(&stream);

  std::vector<float> h_state = {0.0f};
  std::vector<float> h_obs   = {2.0f};

  cunls::dvector<float> d_state(h_state);
  cunls::dvector<float> d_obs(h_obs);

  // Capacity: how many states / factors the buffers hold (fixed per batch).
  // Batches start with 0 active entries; the active count is set separately and
  // may change between solves up to the capacity. Here every slot is used.
  const size_t capacity = 1;
  const size_t num_states = 1, num_factors = 1;
  cunls::VectorStateBatch<1> state_batch(d_state.data(), capacity);
  cunls::PriorVectorFactorBatch<1> prior(
      reinterpret_cast<const cunls::Vector<1>*>(d_obs.data()), capacity);
  state_batch.SetNumActiveStates(num_states);  // active count
  prior.SetNumActiveFactors(num_factors);       // active count

  std::vector<float*> state_ptrs = {state_batch.StateDevicePtr(0)};

  cunls::Problem problem;
  problem.AddStateBatch(&state_batch);
  problem.AddFactorBatch(&prior, state_ptrs);

  cunls::LevenbergMarquardtMinimizer minimizer;
  auto summary = minimizer.Minimize(stream, problem);

  std::cout << "Iterations: "   << summary.num_iterations << "\n";
  std::cout << "Initial cost: " << summary.initial_cost   << "\n";
  std::cout << "Final cost: "   << summary.final_cost     << "\n";

  cudaStreamDestroy(stream);
  return 0;
}

To build it against an installed cuNLS, link libcunls and CUDA::cudart as the C++ examples do.

Robust Estimation with RANSAC

When a fraction of the measurements are gross outliers, swap the minimizer — the problem stays the same. Mark each residual batch as sampled (may contain outliers, classified with an inlier threshold in residual units) or always-on (trusted priors):

Python

options = pycunls.RansacLevenbergMarquardtMinimizerOptions()
options.base_options.factor_batches = [
    pycunls.RansacFactorBatchOptions(pycunls.RansacRole.sampled, 0.01)]
minimizer = pycunls.RansacLevenbergMarquardtMinimizer(options)
summary = minimizer.minimize(stream, problem)  # estimate written back
mask = minimizer.inlier_mask(0)                # numpy uint8, 1 = inlier

C++

cunls::RansacLevenbergMarquardtMinimizerOptions options;
options.base_options.factor_batches = {{cunls::RansacRole::kSampled, /*inlier_threshold=*/0.01f}};

cunls::RansacLevenbergMarquardtMinimizer minimizer(options);
cunls::RansacSummary summary = minimizer.Minimize(stream, problem);  // estimate written back
const uint8_t *inliers = minimizer.InlierMask(0);                    // device, 1 byte per factor

On PnP it recovers the pose at up to 90% outliers, where least squares (even with a Huber loss) fails, in 1.4 ms for 1,000 correspondences and 15 ms for 1,000,000. See RANSAC for the theory, options and tuning, and python/examples/ransac_pnp.py / examples/ransac_pnp for complete programs.

Examples

Testing

Tests require a GPU.

Python tests

pytest -v python/tests

Or in Docker, against a wheel built by build_pycunls_in_docker.sh:

./scripts/test_pycunls_in_docker.sh ./dist

C++ tests

cmake -S . -B build/tests -DCMAKE_BUILD_TYPE=Release -DBUILD_TESTING=ON
cmake --build build/tests -j
ctest --test-dir build/tests --output-on-failure

Or run the test binary directly:

./build/tests/bin/nls_tests

Or in Docker, against the output of build_cunls_in_docker.sh:

./scripts/test_cunls_in_docker.sh ./build_docker

Coverage build:

./scripts/build_cunls.sh build/coverage Coverage

Building Documentation

python -m pip install -r docs/sphinx/requirements.txt
python -m sphinx -b html docs/sphinx docs/sphinx/_build

Or build in Docker:

bash docs/build_in_docker.sh [output_dir]

Code Style

cuNLS follows the Google C++ Style Guide and uses a pre-commit hook for auto-formatting.

sudo apt install pre-commit
pre-commit install

To manually reformat:

sudo apt install clang-format
find . -iname '*.h' -o -iname '*.cpp' | xargs clang-format -i

License

cuNLS is licensed under the Apache License 2.0. Third-party license notices are in NOTICE and third_party/LICENSES/.

About

Nvidia Isaac CUDA-Accelerated Nonlinear Least Square Solver

Resources

Contributing

Security policy

Stars

31 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages