Paper reports a fine-tuning method for long-context AI with sparse attention
An arXiv preprint describes a method that trains transformer language models to work with sparse attention and reports frequent gains over models trained with exact attention.