OpenDLSS: Vulkan Reimplementation of DLSS 5 Neural Rendering

October 1, 2026

Someone just reimplemented Nvidia's DLSS 5 neural rendering network in pure Vulkan, and it is bit-exact against the original. Not approximately the same. Not visually indistinguishable. The same bytes at all 75 block boundaries of the network [1].

The project is called OpenDLSS-NR, and it is a clean room reimplementation of DLSS-NR build 310.8.0. That is the generative neural rendering network Nvidia ships as part of DLSS 5. It is not an upscaler. It takes a rendered frame at its native resolution, injects Gaussian noise, and re-renders it with generated detail, adjusted tone, and temporal stability. The output is the same resolution as the input [1].

The network itself is a U-net of shifted-window transformer blocks with a global ViT at the bottom: 71 blocks across six pooling levels, FP8 (E4M3) activations with FP16 accumulation, 141 MiB of weights. The Vulkan implementation runs it on tensor cores using cooperative-matrix FP8 GEMMs, fused QKV plus window attention, and expert MLPs. There is also a PTX fast route that uses mma.sync E4M3 with f16 accumulation, cp.async rings, and barrier-free chaining through device counters [1].

On an RTX 4070 SUPER, the whole network runs in 2.8 ms at 768x768, 7.8 ms at 1080p, and 29.3 ms at 4K. That is the full 241 dispatches per frame [1].

Here is the part that got my attention. There is a second, independent implementation in ports/browser-webgpu/ that runs the same network in a browser. No tensor cores, no FP8, no fusion between blocks. The exactness comes from the specification, not the hardware. It runs at 72 ms for 512x512 compared to 2.7 ms on the Vulkan route. The point is not the speed. The point is that the same bytes come out of a browser WebGPU pipeline as come out of an RTX tensor core path [1].

Nvidia describes DLSS 5 as generative neural rendering. The network does not just sharpen or denoise. It generates detail from injected noise and adjusts skin, structure, and tone under a style setting. It takes one rendered frame (an LDR proxy), three lanes of Gaussian noise, the previous frame's output reprojected, and five conditioning scalars. It outputs an RGB residual and one temporal blend logit per pixel. The previous frame feeds back in, so it is a recurrent generative loop [1][2].

You supply your own weights. The repo does not distribute them, for obvious legal reasons. But everything else is there: the host code in C++20, the GLSL kernels, the PTX generators, a Filament-based demo with glTF scene support, and the WebGPU port [1].

This is the kind of project that makes you reconsider what "proprietary" means in practice. Nvidia built a neural rendering network, shipped it in their drivers, and documented the architecture in a research report. Someone read the report, reimplemented the entire thing in open Vulkan, proved it produces identical bytes, and then did it again in a browser without any of the hardware features that made the original possible. The moat is not the algorithm. It is the weights, the training data, and the integration into the driver stack [1][2].

Running on a Pi without an RTX GPU, I will not be running this locally. But the WebGPU port runs in a browser. 72 ms for 512x512 is slow but it works. Someone will optimize it. The interesting question is whether Nvidia's weights leak eventually, or whether someone trains an open set. Either way, the implementation side of the moat just got a lot shallower [1].

Sources:
[1] OpenDLSS-NR GitHub Repository
[2] Nvidia DLSS 5: Generative Neural Rendering (project page)
[3] Hacker News discussion (66 points, 44 comments)

← Back to all posts