Skip to content

Build against cuOpt 26.10 nightly: multi-GPU PDLP, new options, iteration/node reporting - #20

Open
0x17 wants to merge 17 commits into
mainfrom
multigpu-pdlp-nightly
Open

0x17 wants to merge 17 commits into
mainfrom
multigpu-pdlp-nightly

Conversation

@0x17

@0x17 0x17 commented Oct 2, 2026 •

Copy link
Copy Markdown
Member
  • Switch both workflows to the latest cuOpt nightly wheels (>=26.10.0a0) from the RAPIDS nightly index
    • needs --prerelease=allow and --index-strategy unsafe-best-match so RAPIDS packages come from the nightly index and CUDA libraries from pypi.nvidia.com
  • Adapt link and bundles to cuOpt 26.10's modular component wheels (both workflows and build-link.sh)
    • libcuopt is now a thin metapackage; libcuopt.so is a linker script, not the engine. Build against libcuopt_mathopt/include plus libcuopt_client/include, link -lcuopt_mathopt and the separate client library, and bundle libcuopt_mathopt/lib64/libcuopt_mathopt.so + libcuopt_client/lib64/libcuopt_client.so (routing is not needed by the GAMS link).
    • Copy libgomp, libtbb, libtbbmalloc and libcudart from libcuopt_mathopt_cuXX.libs, rather than libcuopt_cuXX.libs. The client also uses bundled libgomp/libcares: include both libcuopt_client_cuXX.libs and libcuopt_client.libs, since the unsuffixed client wheel may own the installed libcuopt_client.so. libomp is gone.
    • Copy cuOpt's cuDSS threading layer from libcuopt_mathopt/lib64/libcudss_mtlayer_cuopt.so; do not bundle libcudss_mtlayer_gomp.so. Read build version/hash from the mathopt wheel in build-link.sh instead of using placeholder defaults.
    • libcuopt_client.so needs system libssl.so.3/libcrypto.so.3/libz.so.1 (not bundled).
    • Verified local CUDA 12 and CUDA 13 builds and their ELF dependencies; PR CI builds and packages both CUDA variants on x86_64 and ARM64 (no multi-GPU solve in CI).
  • Support multi-GPU PDLP (Add C API support for multiGPU PDLP  NVIDIA/cuopt#1958)
    • optcuopt.def: num_gpus now accepts -1..72, new multigpu_pdlp_partitioner option
    • bug fix: gmscuopt.c creates the problem with cuOptCreateRangedProblem instead of cuOptCreateProblem, since cuOpt's multi-GPU path only reads constraint lower/upper bounds and saw 0 constraints (validation error with presolve 0)
    • skip initial primal/dual solutions for multi-GPU PDLP, cuOpt rejects them
    • README section on usage and limitations
  • Update optcuopt.def to cuOpt 26.10
    • remove 4 options dropped upstream (mip_hyper_heuristic_presolve_time_ratio, mip_hyper_heuristic_presolve_max_time, mip_hyper_diving_min_node_depth, mip_hyper_submip_node_limit_base)
    • new method 4 (primal simplex) and barrier_dual_initial_point 2 (SeDuMi-style), fix changed default of mip_hyper_heuristic_related_vars_time_limit
    • add 18 new options (e.g. concurrent_nnz_cutoff, primal_simplex_pricing, mip_rens, mip_mutation, barrier regularization, Curtis-Reid scaling); sequence_solve left out (Python re-solve cache only)
  • Report iterations and nodes to GAMS via the new solution attributes (cuOptGetSolutionIntAttribute)
    • LP iterations, MIP nodes and simplex iterations; node-limit detection now uses the actual node count
    • guarded by #ifdef, still compiles against 26.08 headers
  • Pass GAMS nodlim to cuOpt node_limit (was ignored before)
  • Bug fix: drop the PDLP reduced-cost workaround in gmscuopt.c, cuOpt now returns correct reduced costs (Correctly return reduced costs for PDLP (stable3) NVIDIA/cuopt#1797)
  • Verified locally against nightly 26.10.0a222 (CUDA 13, single RTX A1000):
    • regression tests 10/10, gamslib test 73/73 match CPLEX
    • multi-GPU code path (method 1, num_gpus -1, runs even on 1 GPU): 33/35 gamslib LPs match CPLEX within 1e-3; egypt hits the iteration limit (also with single-GPU PDLP), indus89 doesn't converge within 120 s (single-GPU PDLP: optimal in 25 s)
    • iterations/nodes and nodlim checked on trnsport and cube
    • not tested on a machine with several GPUs
  • Not done: cuOptSetLogCallback for a live log, since it only receives lines from the calling thread (a MIP log loses ~75% of its lines)

@0x17 0x17 self-assigned this Oct 2, 2026
@0x17 0x17 changed the title Build against cuOpt nightly and support multi-GPU PDLP Build against cuOpt nightly, add support multi-GPU PDLP, preparations for 26.10 Oct 2, 2026
@0x17 0x17 changed the title Build against cuOpt nightly, add support multi-GPU PDLP, preparations for 26.10 Build against cuOpt 26.10 nightly: multi-GPU PDLP, new options, iteration/node reporting Oct 2, 2026
Comment thread gmscuopt.c
}

status = cuOptCreateProblem(
// Ranged form, since cuOpt's multi-GPU PDLP (without presolve) ignores the row types + RHS

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Huh? please let us know about these types of issues :)
@Bubullzz

@0x17 0x17 Oct 2, 2026 •

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, you're right. Didn't have the time to properly structure this finding. The issue is now here NVIDIA/cuopt#2042.

@0x17
0x17 marked this pull request as draft October 4, 2026 07:55
@0x17
0x17 marked this pull request as ready for review October 4, 2026 07:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants