Researchers target attention heads linked to object hallucinations in LLaVA
A preprint reports that targeting 32 attention heads in LLaVA-1.5-7B reduced hallucinated objects in captions on 400 held-out COCO images.
Dañu koy yeesal bis bu nekk1800 jaar-jaar yuñ firnde
IA buñu saytu bu baax ci lu jëm ci genne ay fasoŋu porodiwi, coppite ci politik, gestu ci kaaraange, ak toxu usine yi, ap ekipu njang buy def te Yàlla tax moo ko leeral ci Àngle bu leer.
Bépp jaar-jaar dafay lëkkale ak firnde yi gëna am doole: balluwaay yu njëkk yi suñu ko amee, luko moy rapoor yuñ joxe ci anam wu leer.
Li xewoon, lu tax mu am solo, ak li ñu wara seetaan - te amul jargon.
Su siñaal bi sew, dunu siiwal dara ludul padding feed bi.
Jaar-jaari IA yuñ saytu fi ñu bawoo, li gëna bees daal di njëkka am, ngir nit ñi bëgga xam IA te baña topp lu bari.
A preprint reports that targeting 32 attention heads in LLaVA-1.5-7B reduced hallucinated objects in captions on 400 held-out COCO images.
A new arXiv preprint presents JIT-Agent, a model designed to generate, repair and improve the software harnesses that guide AI agents. The authors report sizable benchmark gains across several language-model families, but the work remains an unreviewed preprint without independent validation in the source.
An arXiv preprint introduces $R^3$, a post-training method that uses free-form language reasoning to guide robotic manipulation policies. The authors report gains on two controlled benchmarks, while leaving real-world performance and the size of those gains unspecified.
A new arXiv preprint reports that several physical-AI benchmarks measure overlapping information, potentially changing how researchers rank models and choose evaluation suites.
A new preprint reports that medical vision-language models can appear reliable on familiar data while failing cross-dataset transfer, multimodal alignment and shortcut tests.
A new arXiv preprint proposes ProViP, a training-free method that progressively removes redundant visual tokens and focuses pruning on the attention heads most useful for selecting critical visual information. In one reported LLaVA-1.5-7B experiment, it retained 95.9% of the original performance while delivering a…
An audit of 15 frozen hematology, pathology and general-vision foundation models reports steep drops in cross-dataset accuracy and confidence calibration when white-blood-cell images come from different acquisition conditions.
A new preprint introduces LiDAR-SAM2, which transfers video segmentation from SAM2 to temporally consistent 4D LiDAR labeling using multi-view projection and spatio-temporal aggregation.
A new preprint describes SHIFT-LLM, a post-pruning correction method that uses lightweight linear adapters to approximate the computations removed from large language models. The authors report accuracy gains of up to 15.7 points on Llama-3.1-8B-Instruct across seven zero-shot benchmarks.
A new paper introduces YOLOEZ, an open-source graphical tool that combines image labeling, YOLO model training and defect inference in one no-code workflow for structural inspection.
A new arXiv preprint reports a compact monocular-vision pipeline that identifies sidewalk paths and plans routes on CPU hardware. In the authors’ tests, its SegFormer-B0 model reached a hand-annotated intersection-over-union score of 0.946 at 11.7 milliseconds per frame, while image-space midpoint planning produced…
An arXiv paper introduces PointRL, a reinforcement-learning method that uses hidden annotation evidence to train vision-language models to point more reliably at targets.
Benn nettali bu am njariñ ayu-bis bu nekk
Wutal xibaar IA buñ firndeel ci ayu-bis bi, done yu baax, jumtukaay yu am njariñ, tànneefi jàng, ak liggéey IA yu bees.
Nga jël ab liggéeykat IA wala nga genne ab produit IA bu am njariñ? Tegal ko ci kanamu nit ñi ñëw fi ngir jàng ak jëf.
Publie ab liggey IA Yonnee ab jumtukaayu IA