Skip to main navigation Skip to search Skip to main content

Finding the paper behind the data: Automatic identification of research articles related to data publications

  • Marton Ribary
  • , Barbara McGillivray
  • , Kaveh Aryan
  • , Viola Harperath
  • , Mandy Wigdorowitz

Research output: Chapter in Book/Report/Conference proceedingConference contribution

2 Downloads (Pure)

Abstract

Data papers are scholarly publications that describe datasets in detail, including their structure, collection methods, and potential for reuse, typically without presenting new analyses. As data sharing becomes increasingly central to research workflows, linking data papers to relevant research papers is essential for improving transparency, reproducibility, and scholarly credit. However, these links are rarely made explicit in metadata and are often difficult to identify manually at scale. In this study, we present a comprehensive approach to automating the linking process using natural language processing (NLP) techniques. We evaluate both set-based and vector-based methods, including Jaccard similarity, TF-IDF, SBERT, and reranking with large language models. Our experiments on a curated benchmark dataset reveal that no single method consistently outperforms others across all metrics, in line with the multifaceted nature of the task. Set-based methods using frequent words (N=50) achieve the highest top-10% accuracy, closely followed by TF-IDF, which also leads in MRR and top-1% and top-5% accuracy. SBERT-based reranking with LLMs yields the best results in top-N accuracy. This dispersion suggests that different approaches capture complementary aspects of similarity (lexical, semantic, and contextual), showing the value of hybrid strategies for robust matching between data papers and research articles. For several methods, we find no statistically significant difference between using abstracts and full texts, suggesting that abstracts may be sufficient for effective matching. Our findings demonstrate the feasibility of scalable, automated linking between data papers and research articles, enabling more accurate bibliometric analyses, improved tracking of data reuse, and fairer credit assignment for data sharing. This contributes to a more transparent, interconnected, and accessible research ecosystem.
Original languageEnglish
Title of host publicationProceedings of the third Workshop on Information Extraction from Scientific Publications (WASP 2025)
EditorsTirthankar Ghosal, Alberto Accomazzi, Kelly Lockhart, Felix Grezes
Place of PublicationKerrville, TX
PublisherAssociation for Computational Linguistics
Chapter5
Pages34-43
Number of pages10
ISBN (Electronic)979-8-89176-310-4
Publication statusPublished - 23 Dec 2025

Publication series

NameA workshop series associated with the International Joint Conference on Natural Language Processing & Asia-Pacific Chapter of the Association for Computational Linguistics (IJCNLP-AACL)
PublisherAssociation for Computational Linguistics

Keywords

  • scholarly discovery
  • natural language processing
  • Jaccard similarity
  • TF-IDF
  • Sentence Transformers
  • LLM Reranking

Cite this