The hardware behind autonomous machines


Nvidia Puts Rust on the GPU: CUDA Rust Arrives in Two Early-Stage Flavors

Two open-source projects, cuda-oxide and cutile-rs, let developers write GPU kernels in Rust and compile them natively to PTX, with Nvidia saying plainly that neither is production-ready yet.

Nvidia Puts Rust on the GPU: CUDA Rust Arrives in Two Early-Stage Flavors

Nvidia has published two ways to write GPU kernels in Rust, and it is unusually candid that neither one is finished. In a post on its developer blog, Nvidia introduces CUDA Rust as a pair of open-source projects, cuda-oxide and cutile-rs, that turn Rust code directly into instructions its graphics processors can run. The company says it will keep building the effort into 2027 and beyond.

Some definitions first, because I once tried to explain this at a wedding and lost the table by the second acronym. CUDA, which everyone pronounces koo-duh, is Nvidia's programming platform for its GPUs, the chips that do the heavy arithmetic behind AI models. A kernel is the small function that thousands of GPU threads execute at once, like handing the same recipe card to every cook in a very large kitchen. Rust is a language systems engineers love because its compiler refuses to build code that could corrupt memory, so a whole family of bugs is caught before the program runs.

The argument in the post is that the software wrapped around AI models is increasingly written in Rust already. Nvidia's own Nova Linux driver is in Rust, and its Dynamo inference framework sits on a Rust core. Until now the kernel was the holdout: you could launch one from Rust, but you had to write it in another language. The post puts it plainly: "NVIDIA CUDA Rust closes that gap."

Both tracks mirror programming models CUDA already offers. SIMT, short for single instruction, multiple threads, is the classic approach where you describe what one thread does and launch thousands of copies. That is cuda-oxide's job: it plugs into the Rust compiler and translates kernels down to PTX, Nvidia's assembly-like instruction format for GPUs. The Tile track, handled by cutile-rs, works one level up. You say what should happen to one tile of data and let Nvidia's compiler decide how threads and memory get arranged. Nvidia tells developers to reach for Tile first and drop to SIMT only when they need that control.

What I find clever is what the compiler now refuses to do. Nvidia shows that passing a kernel's output buffer as one of its own inputs does not compile on either track, which shuts the door on a classic GPU bug where two threads write the same memory address in an unpredictable order. As the post notes, "Those bugs rarely reproduce on demand, and they pass tests before failing in production."

Then comes the caveat, in the company's own words: "Both projects are early-stage and neither is production-ready." The post says cuda-oxide is in early alpha and needs a pinned nightly Rust toolchain plus a separate LLVM install. cutile-rs is further along. It runs on stable Rust 1.89 or newer with CUDA 13.3, is published on crates.io, and Nvidia says it is already used outside the company in Hugging Face's Grout inference engine and in mistral.rs. Both require Linux and a GPU with compute capability 8.0 or later.

As for the business bet, every language that can target CUDA natively gives developers one more reason to stay on Nvidia hardware. Nvidia credits the community projects that got there first, naming rust-cuda, rust-gpu and cudarc, and says it has been working with the rust-cuda maintainers. Its researcher Melih Elibol is presenting the work at RustConf 2026 in Montreal. The post sums up the difference this time in one line: "What is new is the engineering we are putting behind it, and a clear sense of where it is going." For once the corporate sentence and the code repositories seem to agree.

Leave a Reply

Discover more from Autonomy Magazine

Subscribe now to keep reading and get access to the full archive.

Continue reading