Last updated: 2026-09-02
ROCm backend#
ZINC has a supported Linux ROCm backend for AMD GPUs. It reuses the CUDA-side model loader and scheduler behind the same C ABI, implemented with HIP, hipRTC, and hipBLAS. Vulkan and ROCm are independently supported AMD paths; select ROCm explicitly at build time.
Validated stack#
The current reference configuration is:
- Linux kernel
7.2.2-070202-generic - ROCm userspace
7.2.4 - Radeon AI PRO R9700 (
gfx1201, wave32) - Zig
0.15.2or newer - Qwen 3.5 9B, Qwen 3.6 35B-A3B, and Qwen 3.8 27B
- Gemma 4 26B-A4B MoE and Gemma 4 31B
Kernel 7.2.2 is the validated reference, not a build-time requirement. The
running amdgpu driver, ROCr runtime, and ROCm userspace must agree well enough
for rocminfo and a HIP device query to expose the target GPU.
Build and verify#
Install the ROCm HIP development files that provide amdhip64, hiprtc, and
hipblas, then build with the toolkit root:
ROCM_PATH=/opt/rocm zig build -Dbackend=rocm -Doptimize=ReleaseFast
ROCR_VISIBLE_DEVICES=0 ./zig-out/bin/zinc --check
./zig-out/bin/zinc --version
If the ROCm libraries are not registered with the system dynamic linker, add
/opt/rocm/lib to LD_LIBRARY_PATH for the build and run commands.
Run a GGUF directly:
ROCR_VISIBLE_DEVICES=0 ./zig-out/bin/zinc \
-m /path/to/Qwen3.8-27B-Q4_K_M.gguf \
--chat --prompt "Explain why fair LLM benchmarks need a warmup." -n 96
Qwen 3.8 GGUFs that include the model's appended NextN block use it
automatically for CLI generation. ZINC drafts two tokens, verifies the seed and
drafts together with the full 64-layer model, and restores the recurrent state
at the accepted boundary when a draft is rejected. ZINC_MTP=0 returns to
ordinary greedy decode. ZINC_MTP_DRAFTS=2 is the measured R9700 default;
1 and 3 remain available for GPUs or workloads with a different acceptance
profile. This path is currently CLI-only, so the reusable-server benchmark
table below does not include its speedup.
Start the OpenAI-compatible server:
ROCR_VISIBLE_DEVICES=0 ./zig-out/bin/zinc \
-m /path/to/Qwen3.8-27B-Q4_K_M.gguf \
--parallel 1 --context 2048 --port 9090
Gemma 4 26B-A4B currently serves through its single-sequence MoE path, so ZINC
uses one request slot for that model even if a larger --parallel value is
given. The other validated models keep the requested slot count.
The ROCm-tuned prefill and decode fusions are default-on. These switches remain available as correctness/performance A/B opt-outs:
ZINC_BATCHED_PREFILL=0ZINC_PREFILL_Q8=0ZINC_PREFILL_Q8_REUSE=0ZINC_PREFILL_WMMA_TILES=0ZINC_PREFILL_WMMA_T80=0ZINC_PREFILL_WMMA_TAIL=0ZINC_PREFILL_WMMA_Q4_DIRECT=0ZINC_QWEN_MOE_BATCHED=0ZINC_ATTN_V2=0ZINC_SSM_PREPARED=0ZINC_SSM_COL_WARP=0ZINC_SSM_COL_WARP_FAST=0ZINC_ROCM_FUSED_FFN=0ZINC_ROCM_FUSED_SSM_F32=0ZINC_ROCM_FUSED_Q4_PAIRS=0ZINC_ROCM_FUSED_ATTN_FRONTEND=0ZINC_ROCM_DECODE_PAIR_REDUCE=0ZINC_ROCM_Q4_PAIR_REDUCE=0ZINC_ROCM_DECODE_Q8_FFN=0ZINC_ROCM_DECODE_Q8_Q6=0ZINC_ROCM_DECODE_Q8_LM=0ZINC_ROCM_DECODE_Q8_Q4=0ZINC_ROCM_DECODE_Q8_Q6_PROJ=0ZINC_ROCM_DECODE_Q8_Q5=0ZINC_ROCM_DECODE_Q8_Q4_PAIR=0ZINC_ROCM_RMS_Q8=0ZINC_ROCM_ARGMAX_V2=0ZINC_ROCM_DECODE_SSM_COL_WARP=0ZINC_ROCM_DECODE_SSM_FAST=0ZINC_BATCH_B1_MATVEC=0ZINC_MTP=0(Qwen 3.8 CLI: disable model-native NextN drafting)ZINC_MTP_Q8=0(Qwen 3.8 CLI: use the float multi-token verifier)
The Qwen 3.8 decode path uses packed Q8 activations for the hot Q4_K, Q5_K,
and Q6_K projections, paired projection kernels where the shapes permit it, a
prepared wave32 column scan for the recurrent SSM state, an ILP-unrolled
single-query attention kernel, and a two-stage GPU argmax. RMS normalization
and Q8 activation packing share one wide per-token kernel in prefill and decode,
avoiding a second activation read and launch. Shape-specific WMMA tiles cover
short prompts and long-prompt tails without padding them to a substantially
larger tile. All are ROCm defaults only; setting the corresponding variable to
0 restores the preceding path for comparison.
Qwen 3.6 MoE prefill uses the native token-batched Q4_K/Q5_K/Q6_K expert kernels. CUDA's grouped tensor-core expert path is not dispatched by the ROCm build; its compatibility symbols are intentionally inert.
Fair llama.cpp comparison#
Use the reusable-server performance suite for claims. It runs one engine at a time, uses the same GGUF and scenario matrix, discards a warmup, and reports the median of measured runs:
The RDNA runner applies the same stable host policy before either engine: PCIe ASPM is set to performance and the discrete GPU's memory DPM policy is locked high. This prevents an idle server from changing clocks between phases.
source .env
bun tools/performance_suite.mjs \
--target rdna \
--phase all \
--models qwen38-27b-q4k-m \
--runs 5 \
--warmup 1 \
--rdna-backend rocm \
--rdna-build \
--rdna-start-llama \
--rdna-llama-server /path/to/llama-server \
--rdna-llama-device Vulkan0 \
--no-site-write \
--output /tmp/zinc-rocm-qwen38.json
On the validated R9700 stack, the 2026-09-02 Qwen 3.8 27B four-scenario server
matrix produced these five-run medians. The comparison used llama.cpp commit
9400c8946, which was the tip of upstream master when the run started; it was
not an older historical pin. Overall is the comparable
prompt-plus-generation phase-wall-time score; higher is better. Each server
receives the same textual scenario prompt and applies its native chat template;
prefill and decode rates below are the server-reported phase rates.
| Scenario | ZINC prefill | llama.cpp prefill | ZINC decode | llama.cpp decode | Overall |
|---|---|---|---|---|---|
| Quick Chat | 391.76 tok/s | 150.51 tok/s | 32.27 tok/s | 29.83 tok/s | 122.55% |
| Coding Review | 612.68 tok/s | 253.09 tok/s | 31.97 tok/s | 29.83 tok/s | 116.59% |
| Incident Context | 680.43 tok/s | 308.56 tok/s | 31.88 tok/s | 29.83 tok/s | 121.97% |
| Long Coding Draft | 426.03 tok/s | 186.73 tok/s | 31.97 tok/s | 29.84 tok/s | 112.11% |
ZINC won prefill, decode, and end-to-end phase time in all four scenarios and reached 116.88% of llama.cpp on the summed full-matrix phase-wall-time score (21.111 s versus 24.675 s), or 14.45% less phase wall time. Across the four scenarios, the prefill lead ranged from 120.52% to 160.29% and the decode lead from 6.86% to 8.16%. All four captured ZINC previews passed the suite's output quality checks.
Troubleshooting#
hipErrorNoBinaryForGpu: verify hipRTC receives the detectedgfx*target and that the installed ROCm release supports it.- No device or the wrong device: set
ROCR_VISIBLE_DEVICESand pass-d 0after visibility filtering. - Missing shared library: register
/opt/rocm/libwith the dynamic linker or setLD_LIBRARY_PATH. - Performance drift: stop other GPU processes, discard at least one warmup, and compare persistent server processes rather than a cold CLI launch.