Researchers proposed SetwiseEvalKit, a three-level, nine-dimension document set evaluation benchmark covering both short-text and long-text scenarios, cont...
论文研究HuggingFace Daily Papers(社区热门论文)
Today AI Intelligence Brief
Researchers proposed SetwiseEvalKit, a three-level, nine-dimension document set evaluation benchmark
covering both short-text and long-text scenarios, containing approximately 28K high-quality evaluation criteria. Evaluation of 12 re-rankers showed that the best method had a coverage of less than 45%, and cross-document coordination dimensions were generally weak. Based on this, the training-free method Rubric4Setwise was proposed, which converts criteria into set selection signals, achieving optimal downstream generation performance with fewer documents and search rounds, and is the only method that maintains SOTA in both scenarios.
Original Article Excerpt
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs :https://info.arxiv.org/labs/index.html.
The readable text on this page was extracted from the public source and organized with attribution, publication time and the original link. Copyright remains with the original author and publisher.
Researchers proposed SetwiseEvalKit, a three-level, nine-dimension document set evaluation benchmark covering both short-text and long-text scenarios, cont...