ChemDIRT benchmark finds chemistry LLM performance changes with prompts and molecular representations
A new arXiv benchmark evaluates chemistry-focused large language models across varied instructions, molecular representations and eight task categories, reporting substantial prompt sensitivity, representation dependence and uneven performance.