The hardware behind autonomous machines


Nvidia CUDA 13.4 Lands on Windows on Arm and Opens the Door to Rubin

The toolkit release adds a Windows on Arm target, a preview of the Rubin GPU architecture, and finer-grained ways to share one GPU among many jobs.

Nvidia CUDA 13.4 Lands on Windows on Arm and Opens the Door to Rubin

Nvidia has shipped CUDA Toolkit 13.4, and for the first time the software that turns its graphics chips into general-purpose number crunchers runs on Windows machines built around Arm processors. CUDA, which you pronounce KOO-duh, is the programming layer developers use to write code that executes on an Nvidia GPU instead of on the main processor. Until now, if your Arm laptop ran Windows, that layer simply was not there.

On the company's developer blog, the wording is plain: "CUDA applications have long been supported on Arm platforms through Linux; this release extends that capability to the Windows on Arm platform." Arm, for anyone who has only met it on a phone spec sheet, is the family of low-power processor designs inside nearly every smartphone and a growing share of thin laptops.

Rubin gets its first public outing here too. The toolkit adds functional support, labelled a preview, for Nvidia's next-generation GPU architecture, listed as compute capability 107, so developers can start porting code before Rubin support reaches general availability in a later release. Think of it as receiving the floor plan of a building before the concrete is poured, so the furniture fits on move-in day. I once tried that with an actual apartment and measured the wrong wall, which is why I now trust compilers over tape measures.

Data center operators will care more about a less glamorous change. Multi-Process Service, or MPS, is the mechanism that lets several programs share one GPU at the same time, a bit like a landlord splitting a large flat into separate units. Version 3 adds a scriptable command line, named server instances, configuration files in the TOML format, and memory limits wired into cgroups, the Linux kernel feature that fences off how much of a machine each container may use. In the blog's words, "These capabilities enable precise GPU partitioning, where compute performance, memory boundaries, and execution priority are defined programmatically."

GPUs are the most expensive item on the rack, and a chip sitting half idle because one job could not share it politely is money on fire. Tools that carve a GPU into well-fenced slices let a cloud provider or an enterprise schedule the same hardware to more tenants. That is the bet Nvidia is making with this release: keep the chips busy, and keep the people who program them on CUDA.

There is a quieter change tucked in as well. "CUDA SDK installers no longer bundle the NVIDIA driver." You now install the driver and the toolkit separately through your package manager, which is one more line in the setup guide and one fewer surprise when the two drift out of step.

Further down the stack, a new transport layer called CUDA Compute Fabric Transport lets communication libraries move data across NVLink, Nvidia's high-speed chip-to-chip interconnect, by addressing named endpoints rather than mapping every remote memory allocation into a process's address space. Nvidia says it is meant for the developers who write those libraries, not for most application programmers. In the CCCL 3.4 libraries, the company reports a rewritten device-wide scan reaching up to 92 percent of memory bandwidth on a Blackwell GPU in its own benchmarks, up from around 50 percent in the previous implementation. That is Nvidia's own number on Nvidia's own hardware, so treat it as a claim to test rather than a law of physics.

Downloads are live now, and the Rubin support is explicitly a preview, so anything you build against it today may need adjusting when the production release arrives. Consider that the fine print on the floor plan.

Leave a Reply

Discover more from Autonomy Magazine

Subscribe now to keep reading and get access to the full archive.

Continue reading