NVIDIA NVSHMEM is a parallel programming interface for NVIDIA GPUs based on OpenSHMEM. NVSHMEM can significantly reduce multi-process communication and coordination overheads by allowing programmers to perform one-sided communication from within CUDA kernels and on CUDA streams.
A single-header C++ library for simplifying the use of CUDA Runtime Compilation (NVRTC).
HPC Container Maker
Tilus is a tile-level kernel programming language with explicit control over shared memory and registers.
A fast GPU memory copy library based on NVIDIA GPUDirect RDMA technology
NVIDIA cuOpt examples for decision optimization
Providing reproducibility in deep learning frameworks
Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
Tooling for optimized, validated, and reproducible GPU-accelerated AI runtime in Kubernetes
NVSentinel detects and remediates GPU faults on Kubernetes nodes
A tool for testing and validating container requirements against versioned manifests
A toolkit showing GPU's all-round capability in video processing
Multi-language agent runtime and library for execution scope management, lifecycle events, and middleware on tool and LLM calls.
A nvImageCodec library of GPU- and CPU- accelerated codecs featuring a unified interface
NVIDIA Dataset Utilities (NVDU)
NeMo-Speech.cpp is a lightweight C++ inference runtime for Speech models
Docker image for Swift all-in-one demo deployment
NVIDIA Design System and UI Agent Harness for AI/ML Factories, Robotics, and Autonomous Vehicles
Extended and advanced applications to the Rivermax SDK (Networking SDK for Media and Data Streaming).
NGC Container Replicator