Skip to content

Migrate samples - #2959

Open
danielfrg wants to merge 46 commits into
NVIDIA:mainfrom
danielfrg:migrate-samples
Open

danielfrg wants to merge 46 commits into
NVIDIA:mainfrom
danielfrg:migrate-samples

Conversation

@danielfrg

Copy link
Copy Markdown
Contributor

Description

close #2866

Moving the cuda-python samples from the cuda-samples repo.

Migrating also the sample runners and added a small pytest wrapper to run the test.

Let me know if we want to do any changes to when these are run.

Checklist

  • New or existing tests cover these changes.
  • The documentation is up to date with these changes.

danielfrg and others added 30 commits June 25, 2026 14:55
Resolves conflicts caused by concurrent changes on main:

- Relicense (NVIDIA#2293) updated SPDX headers on cuda_bindings/examples/*.py
  and cuda_bindings/cuda/bindings/_example_helpers/*. All of those files
  were removed on this branch, so main's edits are discarded via git rm.
- Bindings-side test_examples.py, cuda_core/examples/pytorch_example.py
  and cuda_core/tests/example_tests/test_basic_examples.py were removed
  on this branch; main's edits to them are discarded.
- scripts/run_tests.sh was removed on main (NVIDIA#2271); accept the deletion.
- ci/tools/run-tests: keep this branch's optional sample-deps install
  (nvtx, pillow) but pick up main's rename CUDA_VER_MINOR ->
  TEST_CUDA_MAJOR_MINOR.
- Four new texture / GL-interop examples added to cuda_core/examples/
  on main (gl_interop_fluid.py, gl_interop_fluid_numba_cuda_mlir.py,
  gl_interop_mipmap_lod.py, texture_sample.py) are migrated into their
  own directories under samples/ with README.md and requirements.txt.
  Added a DISPLAY guard to glInteropFluid and glInteropMipmapLod so they
  waive with exit code 2 on headless runners, matching glInteropPlasma.
- ruff.toml: ignore RUF059 under samples/** (matches the pre-existing
  ignore for examples/**).
- .spdx-ignore: drop the obsolete cuda_bindings/examples/* line.

Verified end-to-end:
- pixi install -e samples: solves cleanly.
- pytest cuda_core/tests/example_tests/test_samples.py: 46 passed,
  9 skipped (all legitimate: no DISPLAY, no P2P, sub-Hopper GPU, or
  intentionally omitted deps), 0 failed. 55 samples collected.
- ruff check + format: clean.
- toolshed/check_spdx.py: clean on edited files.
Relative imports fail when example_tests has no __init__.py; import run_samples as a top-level module instead.
- systemInfo: guard get_process_name() call behind CUDA_BINDINGS_NVML_IS_COMPATIBLE
  to avoid NameError when NVML bindings are incompatible with the driver.

- vectorAdd: replace cp.random.rand() with NumPy-generated data transferred via
  cp.asarray() to avoid a dependency on libcurand which is not present in the
  cuda-13x wheel environment.
CuPy's FFT backend requires libcufft (so.11 or so.12), which is not
bundled in the cuda-13x wheel environment used in CI.  Probe
cupy.cuda.cufft at startup and exit with the runner waiver code
(CUDA_PYTHON_SAMPLE_WAIVER_EXIT_CODE=77) so the test is reported as
SKIPPED rather than FAILED.
…tSignalAnalysis

CuPy's FFT backend requires libcufft (shipped by nvidia-cufft-cu12 on PyPI).
Adding it to the PEP 723 inline script block and requirements.txt lets the
sample runner automatically waive the test when cuFFT is not installed,
rather than crashing with an ImportError at runtime.
…umpyVsCupy

cp.dot() requires libcublasLt (shipped by nvidia-cublas-cu12 on PyPI).
Adding it to the PEP 723 inline script block and requirements.txt lets the
sample runner automatically waive the test when cuBLAS is not installed.
…gram

cp.random requires libcurand which is not installed in the cuda-13x
wheel CI environment. The random data is only used as histogram input,
so generating it on CPU via numpy and uploading with cp.asarray() is
equivalent and removes the curand dependency.
… is unsupported

ManagedMemoryResource with preferred_location requires concurrent managed
memory access support, which is not available on all devices (e.g. RTX 4090
in the CI environment). Catch CUDA_ERROR_NOT_SUPPORTED at runtime and exit
with the runner waiver code so the test is reported as SKIPPED rather than
FAILED. This mirrors the existing Windows waiver already in the sample.
…cess is unavailable

Per CUDA docs (cuMemPoolCreate): managed memory pools require all devices to
have non-zero concurrentManagedAccess; otherwise CUDA_ERROR_NOT_SUPPORTED is
returned. This attribute is 0 on Windows, WSL, and some Linux VM environments.

Check device.properties.concurrent_managed_access before attempting to create
the ManagedMemoryResource and exit with the runner waiver code when it is False,
mirroring the existing Windows waiver already in the sample.
…cess is unavailable

Managed memory pools (cuMemPoolCreate) require all devices to have non-zero
concurrentManagedAccess; otherwise CUDA_ERROR_NOT_SUPPORTED is returned.
This attribute is 0 on Windows, WSL, and some Linux VM environments.

Move the concurrent_managed_access check to just after device info is printed
and simplify the message, keeping it alongside the existing Windows waiver.
NVIDIA WSL User Guide states that only legacy CUDA IPC APIs are supported
from driver R510; the newer IPC memory pool API (cuMemPoolCreate with POSIX
FD handles) is not supported on WSL2. Detect WSL by checking for 'microsoft'
in /proc/sys/kernel/osrelease and return False from check_ipc_support() so
the sample exits cleanly with a waiver instead of crashing with
CUDA_ERROR_INVALID_VALUE.
…rrent managed access is unavailable

Both samples use ManagedMemoryResource which requires concurrent managed
memory access (concurrentManagedAccess != 0). On WSL and Windows this
attribute is 0, causing CUDA_ERROR_NOT_SUPPORTED at pool creation.

Add concurrent_managed_access check after device init, consistent with
the existing Windows waiver and the fix applied to blurImageUnifiedMemory.
…UDA context

nvmlSystemGetProcessName returns NVML_ERROR_NOT_FOUND when the process
has not yet established a CUDA context (i.e. no cudaMalloc or equivalent
has been called). This is expected NVML behavior per the API docs.

Guard the get_process_name() call with a try/except NotFoundError so
the sample gracefully prints N/A instead of crashing.
Catch OSError (FileNotFoundError is a subclass) when pyglet tries to
load opengl32.dll on headless Windows CI runners (TCC/MCDM mode or
missing display driver). Exit with EXIT_WAIVED so the test is skipped
rather than failing.

See: NVIDIA#2767
Catch OSError (FileNotFoundError is a subclass) when pyglet tries to
load opengl32.dll on headless Windows CI runners (TCC/MCDM mode or
missing display driver). Exit with EXIT_WAIVED so the test is skipped
rather than failing.

Applies the same fix to glInteropFluid, glInteropPlasma, and
glInteropFluidNumbaCudaMlir (same root cause as glInteropMipmapLod).

See: NVIDIA#2767
lijinf2 and others added 12 commits September 4, 2026 19:53
The file already defined EXIT_WAIVED at module level; the extra
definition added near the imports created a duplicate that caused
test_core_sample_waivers_use_negotiated_exit_code to fail (expects
exactly one assignment).
PinnedMemoryResource internally calls cuMemPoolCreate; on Windows with
WDDM T4 (memory_pools_supported=False) this raises CUDA_ERROR_NOT_SUPPORTED.
Check device.properties.memory_pools_supported early and exit with
EXIT_WAIVED so the test is skipped rather than failing.
…unsupported

PinnedMemoryResource internally calls cuMemGetMemPool; on Windows with
TCC mode T4 (memory_pools_supported=False) this raises CUDA_ERROR_NOT_SUPPORTED.
Check device.properties.memory_pools_supported early and exit with
EXIT_WAIVED so the test is skipped rather than failing.
On Python 3.13, ctypes.windll.opengl32 wraps the underlying
FileNotFoundError as AttributeError: opengl32 rather than propagating
the OSError directly. Extend the except clause to (OSError, AttributeError)
so the waiver fires on both Python 3.12 and 3.13 Windows CI runners.
…Harness

When pytest-rerunfailures retries test_main on the same instance,
self.buffer from the previous run may still be set while its backing
DeviceMemoryResource has already been closed by the finally block.
process.start() (spawn mode) pickles self.child_main which requires
serializing self.__dict__, hitting the closed MR and raising:

  RuntimeError: DeviceMemoryResource has been closed
  when serializing dict item 'buffer'

Fix by popping 'buffer' from self.__dict__ at the start of each
test_main run so a stale closed-MR buffer never reaches the pickler.
…nsorMap

The kernel uses <cuda/barrier> from libcudacxx (CCCL). On Windows CI
with pip-installed CUDA, cuda.pathfinder cannot locate the cccl include
directory because nvidia-cuda-cccl-cu12 is not installed, causing NVRTC
to fail with 'cannot open source file cuda/barrier'.

Add nvidia-cuda-cccl-cu12 to the PEP 723 inline dependencies and
requirements.txt so the sample runner installs (or waives) it, matching
the same pattern used for nvidia-cufft-cu12 in fftSignalAnalysis.
cuFile stats level is process-global state (a C library global
variable inside libcufile). Running test_set_stats_level concurrently
with test_stats_start_stop (which also calls set_stats_level(1)) can
cause a race condition where get_stats_level() returns the value set
by the other test, failing the assertion.

All other cuFile stats tests (test_stats_start_stop, test_get_stats_l1/l2/l3)
were already marked thread_unsafe; this commit fixes the oversight for
test_set_stats_level.
SMResource.split(dry_run=True) fails with cuDevSmResourceSplit when
result=NULL and smCount is non-zero. This causes the auto SM-split
probe in greenContext to fail on CUDA 13.1+.

Waive the sample on affected drivers until the cuda.core bug is fixed.
cuda_graphs.py and memory_ops.py were migrated from cuda_core/examples/
to samples/cuda_core/ in this PR. Update the 0.3.1 release notes to
use the :sample-file: role pointing to the new locations so lychee
link checks pass.
* origin/main: (128 commits)
  chore: refresh pixi lockfiles (NVIDIA#2956)
  [no-ci] Improve the Pixi lockfile freshness check  (NVIDIA#2942)
  docs(cuda.core): add 1.2.1 to the docs version switcher (NVIDIA#2944)
  chore: refresh pixi lockfiles (NVIDIA#2943)
  ci: temporarily pause GB300 nightly run (NVIDIA#2941)
  [doc-only] Prepare the 13.4.3 release (NVIDIA#2936)
  fix(bindings): cover missing helper attrs and fix CU_JIT_WALL_TIME pointer (NVIDIA#2931)
  cuda.core: fix pool setup and builder teardown under stream capture (NVIDIA#2838)
  Fix CUDA pointer and coredump attribute types (NVIDIA#2929)
  Use default setuptools-scm tag parsing (NVIDIA#2922)
  chore: refresh pixi lockfiles (NVIDIA#2927)
  chore: refresh pixi lockfiles (NVIDIA#2921)
  feat(cuda.core): support nvJitLink incremental linking (NVIDIA#2867)
  tests: enumerate CUDA devices for GPU-only system checks (NVIDIA#2916)
  Merge 13.4.2 changes into main (NVIDIA#2913)
  fix(cuda.core): preserve adjacency set semantics (NVIDIA#2893)
  Fix primary context cleanup during Python shutdown (NVIDIA#2744)
  Simplify and speed up param packer (NVIDIA#2904)
  Fix VMM handle leaks in virtual memory paths (NVIDIA#2235)
  [chore] Remove obsolete sccache settings (NVIDIA#2860)
  ...

# Conflicts:
#	cuda_bindings/pixi.lock
#	cuda_core/pixi.lock
@copy-pr-bot

copy-pr-bot Bot commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@github-actions github-actions Bot added CI/CD CI/CD infrastructure cuda.bindings Everything related to the cuda.bindings module cuda.core Everything related to the cuda.core module cuda.pathfinder Everything related to the cuda.pathfinder module labels Sep 28, 2026
@danielfrg danielfrg self-assigned this Sep 28, 2026
@danielfrg danielfrg added P0 High priority - Must do! example Improvements or additions to code examples labels Sep 28, 2026
Revert the cp.asarray(np.random...) CPU-generate-then-transfer workaround
back to cp.random.rand()/cp.random.randint() in both samples, per review
feedback that the CPU fallback deviated from the original CuPy-based
content and taught users an unnecessary host-to-device transfer.

cp.random dlopens libcurand at runtime, which the cuda_core CI test
environment does not currently install. Rather than widening the shared
test-cu12/test-cu13 dependency groups, declare nvidia-curand explicitly
in each sample's PEP 723 script header and requirements.txt so the
shared sample runner's dependency gate (missing_dependencies) waives
the sample cleanly when curand is unavailable instead of the process
crashing on dlopen.
@lijinf2

lijinf2 commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

/ok to test 6b5e79d

@github-actions

Copy link
Copy Markdown
Contributor

Replace nvidia-cufft-cu12/nvidia-cublas-cu12 with nvidia-cufft/nvidia-cublas
in fftSignalAnalysis and numpyVsCupy, per review feedback not to mix
CUDA 12 and CUDA 13 packages. Follow-up on switching to cupy-cuda13x[ctk]
tracked in NVIDIA#2961.
@lijinf2

lijinf2 commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

/ok to test 562e4c6

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CI/CD CI/CD infrastructure cuda.bindings Everything related to the cuda.bindings module cuda.core Everything related to the cuda.core module cuda.pathfinder Everything related to the cuda.pathfinder module example Improvements or additions to code examples P0 High priority - Must do!

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Drive PR #2266 (Migrating cuda-python sample codes from cuda-samples repository and Testing) across the finish line

2 participants