Fork of cudarc (MIT OR Apache-2.0) — upstream v0.19.7 plus one build.rs commit that probes the host for CUDA link libraries and falls back to the dlopen path when they are absent. Consumed by Blazen via [patch.crates-io].
  • Rust 99.9%
  • C 0.1%
Find a file
Zachary Handley f0601ac153 build: make the CUDA link mode a host probe instead of a hard link edge
cudarc's `dynamic-linking` mode unconditionally emits
`rustc-link-lib=dylib={cuda,nvrtc,curand,cublas,cublasLt}`, and
candle-core 0.10.x hardcodes that feature on its cudarc edge. On a host
with no CUDA toolkit the link then fails, and the `nvcc` version probe
panics outright unless `fallback-latest` happens to be enabled.

Decide the link mode up front instead:

  * `cuda_link_libs_present()` reuses cudarc's own `link_searches()`
    directories and looks for the `cublas` linker-name (libcublas.so /
    libcublas.dylib / cublas.lib) that `ld` would resolve against.
  * When those libs ARE present the upstream path runs verbatim, so
    CUDA hosts and GPU CI lanes are byte-identical to upstream.
  * When they are absent -- or `BLAZEN_NO_CUDA=1` is set -- emit
    `cfg(feature = "dynamic-loading")` and skip both the link-libs and
    the `nvcc` probe. cudarc's source keys the link mode solely on
    `dynamic-loading` (it never reads `dynamic-linking`), so this
    switches the crate to its dlopen path even though the cargo feature
    selected was `dynamic-linking`. The libs are then dlopen'd at
    runtime, which never happens on a host without CUDA.

build.rs only; no change to src/. Applied on top of upstream v0.19.7
(3e5d38b5fe), the exact rev that produced
the crates.io cudarc 0.19.7 artifact.
2026-08-06 19:49:26 -07:00
.gemini Adding gemini config 2025-07-28 16:14:59 -04:00
.github Modify FUNDING.yml to include placeholders 2026-05-06 15:21:35 -04:00
bindings_generator Load symbol on use instead of at initialization (#573) 2026-05-11 11:31:21 -04:00
examples Adds peer memcpy (#520) 2026-01-21 18:24:11 -05:00
src Adding safe Group api to nccl (#578) 2026-05-15 11:57:51 -04:00
.gitignore Adds an attribute method for device and a .rustfmt.toml to the root so formatting is respected when inside a Cargo workspace (#154) 2023-06-21 09:19:12 -04:00
.rustfmt.toml Adds an attribute method for device and a .rustfmt.toml to the root so formatting is respected when inside a Cargo workspace (#154) 2023-06-21 09:19:12 -04:00
build.rs build: make the CUDA link mode a host probe instead of a hard link edge 2026-08-06 19:49:26 -07:00
Cargo.toml v0.19.7 2026-05-15 11:58:32 -04:00
LICENSE-APACHE Updating stuff for cargo 2022-09-16 19:06:35 -04:00
LICENSE-MIT Updating stuff for cargo 2022-09-16 19:06:35 -04:00
README.md Revise version support details in README 2026-05-06 15:20:42 -04:00

cudarc: minimal and safe api over the cuda toolkit

crates.io docs.rs

Checkout cudarc on crates.io and docs.rs.

Contributions welcome!

Safe CUDA wrappers for:

library dynamic load dynamic link static link
CUDA driver N/A
NVRTC
cuRAND
cuBLAS
cuBLASLt
NCCL
cuDNN
cuSPARSE
cuSOLVER N/A
cuFILE
CUPTI
nvtx N/A
cuFFT

CUDA Versions supported (choose with -F cuda-<version>, like cuda-13010):

  • 11.4-11.8
  • 12.0-12.9
  • 13.0

CUDNN versions supported (choose with -F cudnn-<version>, like cudnn-09021):

  • 8.9.7
  • 9.10.2
  • 9.21.1

NCCL versions supported (choose with -F nccl-<version>, like nccl-02023):

  • 2.22-2.30

Configuring CUDA version

Select cuda version with one of:

  • -F cuda-version-from-build-system: At build time will get the cuda toolkit version using nvcc
    • -F fallback-latest: can be used to control behavior if this fails. default is not enabled, which will cause the build script to panic. if -F fallback-latest is enabled, we will use the highest bindings we have.
  • -F cuda-<major>0<minor>0 to build for a specific version of cuda

Configuring linking

By default we use -F dynamic-loading, which will not require any libraries to be present at build time.

You can also enable -F dynamic-linking or -F static-linking for your use case.

Getting started

It's easy to create a new device and transfer data to the gpu:

// Get a stream for GPU 0
let ctx = cudarc::driver::CudaContext::new(0)?;
let stream = ctx.default_stream();

// copy a rust slice to the device
let inp = stream.clone_htod(&[1.0f32; 100])?;

// or allocate directly
let mut out = stream.alloc_zeros::<f32>(100)?;

You can also use the nvrtc api to compile kernels at runtime:

let ptx = cudarc::nvrtc::compile_ptx("
extern \"C\" __global__ void sin_kernel(float *out, const float *inp, const size_t numel) {
    unsigned int i = blockIdx.x * blockDim.x + threadIdx.x;
    if (i < numel) {
        out[i] = sin(inp[i]);
    }
}")?;

// Dynamically load it into the device
let module = ctx.load_module(ptx)?;
let sin_kernel = module.load_function("sin_kernel")?;

cudarc provides a very simple interface to launch kernels using a builder pattern to specify kernel arguments:

let mut builder = stream.launch_builder(&sin_kernel);
builder.arg(&mut out);
builder.arg(&inp);
builder.arg(&100usize);
unsafe { builder.launch(LaunchConfig::for_num_elems(100)) }?;

And of course it's easy to copy things back to host after you're done:

let out_host: Vec<f32> = stream.clone_dtoh(&out)?;
assert_eq!(out_host, [1.0; 100].map(f32::sin));

Design

Goals are:

  1. As safe as possible (there will still be a lot of unsafe due to ffi & async)
  2. As ergonomic as possible
  3. Allow mixing of high level safe apis, with low level sys apis

To that end there are three levels to each wrapper (by default the safe api is exported):

use cudarc::driver::{safe, result, sys};
use cudarc::nvrtc::{safe, result, sys};
use cudarc::cublas::{safe, result, sys};
use cudarc::cublaslt::{safe, result, sys};
use cudarc::curand::{safe, result, sys};
use cudarc::nccl::{safe, result, sys};

where:

  1. sys is the raw ffi apis generated with bindgen
  2. result is a very small wrapper around sys to return Result from each function
  3. safe is a wrapper around result/sys to provide safe abstractions

Heavily recommend sticking with safe APIs

License

Dual-licensed to be compatible with the Rust project.

Licensed under the Apache License, Version 2.0 http://www.apache.org/licenses/LICENSE-2.0 or the MIT license http://opensource.org/licenses/MIT, at your option. This file may not be copied, modified, or distributed except according to those terms.