I got Whisper running locally on my Radeon 890M and XDNA 2 NPU — could they work together?
I've been experimenting with local speech-to-text on my Windows laptop. After getting Whisper running on both the integrated GPU and NPU separately, I'm wondering whether there's a practical way to combine their processing capacity. I thought I'd share what I've tested so far and see if anyone has tried something similar. Hardware and setup Laptop: Lenovo ThinkPad P14s Gen 6 AMD CPU: AMD Ryzen AI 9 HX PRO 370 GPU: Integrated Radeon 890M NPU: AMD XDNA 2 OS: Windows 11 Model: Whisper Large-v3 Turbo Backend: Whisper.cpp (Vulkan and Vitis AI configurations) How this started I purchased a lifetime license for Vowen Pro because I wanted a Windows alternative to MacWhisper for transcribing downloaded lectures, interviews, and videos locally. Its built-in transcription worked, but processing a 103-minute English SEO lecture took approximately 74 minutes. I wanted to see whether I could make better use of the AMD hardware already in my laptop. After some experimentation, I got two separate Whisper.cpp configurations working. Benchmark results Using the same five-minute section of an English recording: Configuration Processing time Approx. real-time speed XDNA 2 NPU (Vitis AI) 80.29 seconds 3.74× Radeon 890M (Vulkan) 42.76 seconds 7.02× These are individual local tests, not averages from repeated benchmark runs. The NPU configuration uses the Vitis AI encoder with CPU participation elsewhere in the inference pipeline. The Vulkan build uses the Radeon 890M as its GPU compute backend. During Vulkan inference, Windows Task Manager showed GPU Compute utilization reaching 100%. Integrating it with Vowen I then started the Vulkan build's Whisper.cpp HTTP server on 127.0.0.1:8088 , using /v1/audio/transcriptions as the inference endpoint. Vowen Pro has a custom OpenAI-compatible transcription server option, so I tried connecting it to my local server. It worked. Vowen could import MP3 and MP4 files, prepare the audio, send it to the local Whisper server, and display the resulting transcript. I subsequently tested a full 103-minute lecture, followed by a 15:45 video that finished in roughly two minutes. I didn't use a precise timer for the latter, so consider that an approximate end-to-end result rather than a formal benchmark. I also configured Windows Task Scheduler to launch the server silently at login. After rebooting, the server came back online without opening a console window, and Vowen successfully completed another transcription. I now have a usable local transcription workflow without manually launching the inference server every time. My question: can the GPU and NPU work in parallel? Since the Radeon 890M and XDNA 2 can both run Whisper separately, I'm curious whether using them simultaneously could improve throughput for long recordings. One idea would be to: Split a long recording into overlapping segments. Send some segments to a Vulkan GPU worker and others to a Vitis AI NPU worker. Process both queues concurrently. Merge the transcripts and reconcile timestamps at segment boundaries. I'm aware this introduces potential problems: shared memory bandwidth, CPU-side decoding, thermal and power limits, uneven workloads, and transcription consistency between segments. I also realize that splitting work between devices doesn't necessarily mean the total processing time will improve. Has anyone attempted this kind of heterogeneous GPU + NPU workload on AMD Ryzen AI hardware? Are there existing projects or approaches worth looking into? And would the shared system resources likely make this slower than simply running the Vulkan backend alone? I'd be interested in hearing about actual benchmarks, implementation experiences, or technical limitations I may have overlooked. The attached screenshots show my Vowen workflow and Radeon 890M utilization during Vulkan inference. The NPU benchmarks were conducted separately. I haven't implemented simultaneous GPU+NPU processing. No affiliation with Vowen, AMD, or the Whisper-related projects mentioned. This is just a personal experiment with hardware and software I already own. I've been experimenting with local speech-to-text on my Windows laptop, and I wanted to share some results and ask a question about using AMD's GPU and NPU together. My hardware: Lenovo ThinkPad P14s Gen 6 AMD Ryzen AI 9 HX PRO 370 Integrated Radeon 890M GPU XDNA 2 NPU Windows 11 What I set up I bought a lifetime license for Vowen Pro because I wanted something similar to MacWhisper for transcribing long lectures, interviews, and videos locally. The built-in transcription worked, but it was relatively slow. A 103-minute English SEO lecture took roughly 74 minutes to transcribe. I started experimenting with Whisper.cpp and AMD's acceleration backends. I got two separate configurations working with the same Whisper Large-v3 Turbo model: AMD XDNA 2 NPU via Vitis AI: 5 minutes of audio processed in 80.29 seconds. Radeon 890M via Vulkan: The same 5 minutes processed in 42.76 seconds. That's approximately 3.7× real-time for the NPU configuration and 7.0× for Vulkan. The NPU configuration still uses CPU resources for parts of the inference pipeline, while the Vulkan version puts a substantial amount of work on the integrated GPU. Getting it into a usable workflow I then launched the Vulkan build's Whisper.cpp HTTP server locally on port 8088, with its transcription endpoint mapped to /v1/audio/transcriptions . To my surprise, Vowen Pro's custom OpenAI-compatible server option connected successfully. It also handled audio extraction from MP4 files automatically, so I could import videos directly into the GUI. I tested the setup with a full 103-minute lecture and, after a Windows reboot, with another 15:45 video. The latter finished in roughly two minutes, although I didn't measure that run with a precise timer. I also configured Windows Task Scheduler to launch the server silently at login. After rebooting, the server was listening on localhost without opening a console window, and Vowen successfully completed another transcription. So I now have a convenient, GPU-accelerated local transcription workflow without manually launching a terminal each time. Here's what I'm curious about: Since I have both an integrated Radeon GPU and an XDNA 2 NPU, could I make better use of them simultaneously? I'm not necessarily talking about having them execute the same model operations together. I was thinking about splitting a long recording into segments, sending some to a Vulkan GPU worker and others to a Vitis AI NPU worker, and then merging the transcripts and timestamps. I realize there could be complications involving shared memory bandwidth, CPU decoding, power limits, segmentation boundaries, and scheduling overhead. Has anyone experimented with this kind of heterogeneous GPU + NPU transcription pipeline on AMD hardware? Would parallel processing actually improve total throughput, or would the shared resource constraints outweigh the benefits? I'd be interested in any existing implementations, benchmarks, or suggestions. Not affiliated with Vowen or any of these projects. Just experimenting with the hardware and software I already own.