Mazoni, Alysson; Borges, Luís; Macedo, Estevão Fernandes; Tuesta, Esteban Fernández; Mena‐Chalco, Jesús Pascual
Funded by CAPES, FAPESP, UNICAMP
Disambiguating research entities remains a long-standing methodological problem in scientometric analyses, as inconsistencies and ambiguous metadata limit the interoperability of major bibliographic databases. While global systems such as OpenAlex provide extensive coverage, they often lack the granularity and accuracy provided by national-level research databases. This study proposes a large-scale methodology to enhance author and institutional disambiguation by integrating local (Lattes and CAPES) and global (OpenAlex) databases. The method combines shared Digital Object Identifiers with an adapted Levenshtein distance algorithm to handle variations in author and institutional names across multilingual research databases, achieving over 97% accuracy for authors and 64% for institutions. The proposed framework provides a scalable and replicable approach for entity disambiguation in tabular research databases. Beyond the Brazilian context, this integration strategy offers a globally applicable approach for harmonizing national research information systems with open scientometric infrastructures.
