DataKernelBench tests whether LLMs can optimize GPU database queries
A new benchmark evaluates whether language models can generate and improve GPU code for database-style queries, with the paper reporting speedups of up to 2.11× on one H100 GPU and 2.54× across four H100 GPUs.