J-Miner paper reports executable decision rules extracted from language-model classifiers
Researchers describe J-Miner, a method that converts internal signals from fine-tuned language-model classifiers into inspectable rules, reporting up to 98.3% reproduction of source decisions and nearly the same average task accuracy in much smaller student models.