TeksEdge reports that another builder documented a progression running Qwen3.8-27B on the same pair of Radeon AI PRO R9700 GPUs, with llama.cpp default at 19.2 tok/s, llama.cpp tuned for gfx1201 with tensor split at 29.7 tok/s, and vLLM Radiance with MXFP4 at 184.9 tok/s. The latest Radiance benchmarks are now reporting 183.1 tok/s weighted single-stream, 188.8 tok/s code, 218.8 tok/s file editing, 243.7 tok/s math, 249.4 tok/s JSON, and up to 523.5 aggregate tok/s at concurrency 8.
A separate deployment on 2× Radeon AI PRO R9700 with 64GB total dedicated VRAM, a Ryzen 7500F, 64GB DDR5, and PCIe 5.0 x8 per GPU used vLLM Radiance with TP2 and MTP. Running Qwen3.8-27B in Quark AWQ MXFP4, that setup reported 111.4 tok/s median decode, 4,224–4,410 tok/s prefill, 81ms TTFT, and 131K server context, while native FP8 achieved 87.6 tps.