CUDA for AMD on Windows

(github.com)

46 points | by chiassedu80 2 hours ago

4 comments

  • linuxhansl 19 minutes ago
    Off-topic and somewhat of a rant, but I'd far prefer us all focusing on open standards like HIP, SYCL, OpenCL, etc.

    It's unbearable that most LLM inference happens on closed H/W, closed drivers, and closed SDKs.

    • bigyabai 4 minutes ago
      I would too, but sadly that's Khronos' job to organize, and they've had trouble getting American vendors to work together.

      It's likely that CUDA will continue dominating until they put aside their differences. The current MLX/MPS/ROCm ecosystems are too fractured to threaten Nvidia.

      • high_na_euv 0 minutes ago
        Wdym American Vendors?

        Intel uses SPIRV iirc

  • lulzx 27 minutes ago
    I made cuda-metal btw (for mac kek), https://github.com/lulzx/cuda-metal
  • system2 56 minutes ago
    I wish there were a way to use RDNA1 cards with CUDA for AMD. My 5700XTs are sitting in a drawer.
    • monster_truck 3 minutes ago
      RDNA1 isn't good for a whole lot, even flagship RDNA2 cards are a stretch for many things. The lack of WMMA/matrix multiply/BF16 is too severe of a penalty.

      The FP16 throughput on RDNA1 is both shader reliant and requires everything to be packed first. Even with 2 or 4 or 1000 cards, you would be consuming all of the available memory and memory bandwidth just packing and unpacking values, and if you really want to dump a hundred billion tokens into making it work anyways, you're only going to find out that even if you bother to sit there ferrying packed values to ram or disk before then issuing the instructions, paying that already severe penalty again when the values then have to be unpacked is so steep of a cost that the 256 BF16 flops/cu/clock's effective throughput is outright lower than simply doing it on a Zen 2 processor. You also don't have INT8 (or really INT4) on RDNA1 so the other RNS/CRT tricks aren't viable.

      Sadly RDNA1's VCN2 also lacks actually good x264 bframe encoding support, or even P010 for 10 bit color, so what I'm saying is you should sell them. Used Radeon VII's are like $260, you'll go a lot further with those especially if you throw in a 7900XTX, and then augment that further with a 9070 CRE (you only want it for its int8 cores), and of course 128GB of ram.

      E: And sure, that's 3, or ideally 4 GPUs, and a good bit of extra work. But that gets you up to more than halfway to the naive performance of a $15,000 MI300x in a surprising amount of cases, with additional strengths that it lacks.

  • chiassedu80 2 hours ago
    CUDA for AMD on Windows

    I’ve been working on a Windows setup that lets CUDA-targeted applications run on AMD GPUs using ZLUDA + ROCm/HIP.

    Repo: https://github.com/Speedstu/CUDA-for-AMD-Windows

    So far, it has only been tested on my RX 9060 XT (gfx1200), where I’ve used it with CUDA-enabled LibTorch workloads, including long ai training and use.

    I also added a GPU scanner / auto-detection system that detects:

    AMD GPU model

    gfxXXXX architecture

    ROCm/HIP installation

    driver info

    whether the GPU has already been validated by the project

    Example:

    RX 9060 XT → gfx1200 → RDNA4 → HIP detected → validated

    The goal now is to test it on more hardware, especially RX 6000 / 7000 / 9000 cards.

    If you have an AMD GPU on Windows and want to try it, I’d really appreciate compatibility reports working or broken. There’s a dedicated GPU compatibility issue template in the repo.

    If this is useful to you, a star would also help the project get more testers.

    • nine_k 24 minutes ago
      (As a side note, I love the name ZLUDA; it very aptly means "delusion" or "deception" in Polish.)