Mugnaini, Rogério; Rodrigues, Wellington Barbosa; Damaceno, Rafael J. P.
Os resultados indicam que a OpenAlex amplia a recuperação da produção científica em relação às bases analisadas, sobretudo no caso brasileiro. A comparação entre Brasil e Espanha evidencia diferenças associadas ao peso das revistas nacionais e à inserção em bases internacionais e regionais. O estudo reforça o potencial da OpenAlex para análises bibliométricas, mas também aponta a necessidade de interpretar seus dados de forma complementar e contextualizada.
O estudo analisa a integração entre a Plataforma Lattes e a OpenAlex para identificação de pesquisadores brasileiros, avaliando a cobertura do ORCID autodeclarado e a inferência por meio da combinação DOI–OpenAlex com heurística nominal. A partir de 307 bolsistas PQ-SR, observa-se baixa cobertura de ORCID no Lattes (45,9%), mas elevada recuperação via heurística (94,5% com ao menos um ORCID). Embora a estratégia amplie a identificação, surgem ambiguidades, evidenciando limites metodológicos e a necessidade de cautela em análises cientométricas integradas.
Choji, Thamyres Tetsue; Mazoni, Alysson; Moral-Muñoz, José A.; Cobo, Manuel J.; Costas, Rodrigo
Este estudo compara a presença de autores latino-americanos nas bases OpenAlex, Ranking Aberto de Leiden, Scopus e SciELO, analisando diferenças de gênero e tempo de carreira. Os resultados indicam que OpenAlex apresenta perfil mais inclusivo e maior representação de autores em início de carreira. LRCCOE e Scopus mostram maior seletividade e maior presença masculina, enquanto SciELO apresenta maior proporção de mulheres e menor diferença de gênero em estágios seniores. As desigualdades variam de acordo com a base utilizada.
Choji, Thamyres Tetsue; Mazoni, Alysson; Moral-Muñoz, José A.; Costas, Rodrigo; Cobo, Manuel J.
Nowadays bibliographic data sources are available, and their coverage is shaped by their scope and indexing criteria. Their comparison can reveal what they underrepresent or exclude, particularly in regions such as Latin America, which presents specific characteristics in knowledge production. This study compares four data sources: OpenAlex, Leiden Ranking Core Collection Open Edition, Scopus, and SciELO. We analyze document coverage and author visibility by gender, career length, and language of publication. OpenAlex was used as the baseline, and comparisons were made using DOI and persistent author identifiers. Results show that, after OpenAlex, Scopus presents the largest coverage of documents (47.7%) and authors (52.1%), but with larger gender differences than SciELO. SciELO shows higher representation of women, particularly among exclusive authors, career lengths, and more authors publishing in Spanish (30.4%) and Portuguese (46%). These findings suggest that data sources may shape gender representation and multilingual publication practices in Latin America.
English functions as the principal shared language of international scholarly communication, but its dominance also creates inequalities in the production, dissemination, visibility, and assessment of research. Developed as a complementary output of the CoARA Working Group on Multilingualism and Language Biases in Research Assessment, this report examines the historical development of English dominance and the contemporary mechanisms through which it is reproduced. It combines evidence from the literature with descriptive bibliometric analyses based on Web of Science, Scopus, OpenAlex, and CWTS Journal Indicators. The findings show that English is especially dominant in journals with high field normalised citation impact and in the selective segments of scholarly communication most frequently used in research evaluation. However, the extent of this dominance varies considerably across disciplines and countries, and publication in national languages remains important in the social sciences, humanities, and fields directed toward local professional and societal audiences. The report also demonstrates how database selection, metadata quality, and analytical choices affect the visibility of multilingual scholarship. Web of Science and Scopus predominantly represent English language journal literature, whereas OpenAlex reveals a much broader body of multilingual research, although incomplete language, authorship, and institutional metadata remain important limitations. The report further examines the financial, linguistic, and time burdens associated with English language publication and considers how publishing infrastructures, bibliometric indicators, rankings, and assessment practices influence which research becomes visible and valued. This report contributes to the landscape analysis conducted by the Coalition for Advancing Research Assessment (CoARA) Working Group (WG) on Multilingualism and language biases in research assessment and it complements the Implementation proposal for language-aware assessments (Annex II.3).
Local research has gained increasing attention in recent years, with a growing body of literature emphasizing its significance—especially for peripheral communities underrepresented in mainstream scientific discourse. However, despite this growing interest, there remains no clear consensus on how to define local research. Conceptual clarity is essential for conducting consistent and replicable analyses that capture the specific characteristics, contexts, and societal relevance of local research. This need is particularly acute in the Global South, where structural inequalities in visibility and recognition often marginalize regionally focused scholarship. In Africa, for instance, a substantial share of research is published in local or regional journals that are poorly represented in global indexing databases. As bibliometric analyses typically rely on these databases, this exclusion diminishes their perceived relevance. To address this gap, the present study incorporates data from regional and open-access sources to examine citation patterns associated with African local research. It uses multiple conceptualizations of local research to classify publications as either local or non-local depending on the definition used. These classifications are correlated with three citation-based indicators: average citations per publication, time to first citation, and the proportion of uncited publications. Additionally, the study explores the characteristics of citing publications, distinguishing between local and non-local sources as well as domestic and international citations. The findings reveal that citation patterns differ significantly across locality definitions, underscoring the importance of adopting nuanced, context-sensitive approaches to assess and support local research.
Disambiguating research entities remains a long-standing methodological problem in scientometric analyses, as inconsistencies and ambiguous metadata limit the interoperability of major bibliographic databases. While global systems such as OpenAlex provide extensive coverage, they often lack the granularity and accuracy provided by national-level research databases. This study proposes a large-scale methodology to enhance author and institutional disambiguation by integrating local (Lattes and CAPES) and global (OpenAlex) databases. The method combines shared Digital Object Identifiers with an adapted Levenshtein distance algorithm to handle variations in author and institutional names across multilingual research databases, achieving over 97% accuracy for authors and 64% for institutions. The proposed framework provides a scalable and replicable approach for entity disambiguation in tabular research databases. Beyond the Brazilian context, this integration strategy offers a globally applicable approach for harmonizing national research information systems with open scientometric infrastructures.
Numerous initiatives are currently underway to disambiguate databases worldwide. In this paper, we propose a methodology for disambiguating research entities using big data techniques, adopting an approach that goes from local to global databases. Our objective is to enhance the quality of data in the OpenAlex database by leveraging information from Brazilian databases, particularly data from the Lattes Platform and the Brazilian Federal Agency for Support and Evaluation of Graduate Education. We compare similar names of authors and institutions, employing Digital Object Identifiers to link entities, along with an adaptation of the Levenshtein distance algorithm. The proposed method is straightforward to implement in tabular databases and facilitates disambiguation, thereby contributing to open science practices and providing an effective solution for research information systems. The findings indicate the potential for integrating local and global databases to address issues related to ambiguous names and incomplete metadata.