Researchers introduce PTXBench, a benchmark for evaluating and adapting large language models to generate architecture-specific PTX code for GPU kernel optimization, focusing on correctness, instruction execution, and speedup on H100 and B200 hardware. The study tests GEMM and attention workloads, revealing uneven PTX capabilities. Models especially struggle with complex attention backward passes, and successfully executing targeted instructions often fails to yield performance on par with leading GPU libraries. Authors fine-tune Qwen3.6-27B with supervised, repair-conditioned training, improving some tasks but exposing limited generalization. They emphasize data coverage, balance, and teacher quality, positioning PTXBench as an auditable platform for tracking progress on evolving architectures.
This update represents a notable development in the Ai sector. Organizations and founders tracking this space should evaluate potential strategic and technical implications on their operations.