multiobs

Our Recent Productions

Datasets, working papers, and peer-reviewed articles produced by MultiObs researchers.

Showing 1–6 of 6

  1. Working Paper2026Open Access
    Collective action for Open Research Information: Introducing ORION-DBs and a new FORCE11 Working Group(opens in a new tab)

    Mazoni, Alysson; Kramer, Bianca; Neylon, Cameron; Costas, Rodrigo; van Eck, Nees Jan

    With ORION-DBs, a growing set of actors are hosting large research information datasets in shared spaces, notably in Google BigQuery. To coordinate amongst current and future contributors and work towards community standards for discovery, processing, preservation and documentation, we are starting a FORCE11 working group. The working group will also seek to build a community around future alternatives to proprietary cloud systems for data sharing at scale and identify future paths to community-led training and development resources. Our goal is to build a community with a focus on supporting the users of Open Research Information resources to make more sophisticated and complex use of these data sources. We will do this by consolidating and expanding the resources themselves through supporting providers, gathering user stories and developing training materials. Finally we will scope pathways to technological independence from proprietary systems consistent with the needs of users. This webinar will introduce the ORION-DB initiative, present a number of use cases and introduce the FORCE11 Working Group. If you are a data provider, metadata user or infrastructure developer and want to learn more about ORION-DB and the planned activities of the working group, you are invited to attend the webinar. The version of this slide deck on Google was originally: https://docs.google.com/presentation/d/1ua88tr2gP2HTurQPkM4-A7eWjFiDLUI_ht3MJdHxOxc/edit?usp=sharing

  2. Conference Paper2026Open Access
    Gender representation of Latin American authors across bibliographic data sources: OpenAlex, Leiden Ranking Core Collection Open Edition, Scopus and SciELO(opens in a new tab)

    Choji, Thamyres Tetsue; Mazoni, Alysson; Moral-Muñoz, José A.; Costas, Rodrigo; Cobo, Manuel J.

    Nowadays bibliographic data sources are available, and their coverage is shaped by their scope and indexing criteria. Their comparison can reveal what they underrepresent or exclude, particularly in regions such as Latin America, which presents specific characteristics in knowledge production. This study compares four data sources: OpenAlex, Leiden Ranking Core Collection Open Edition, Scopus, and SciELO. We analyze document coverage and author visibility by gender, career length, and language of publication. OpenAlex was used as the baseline, and comparisons were made using DOI and persistent author identifiers. Results show that, after OpenAlex, Scopus presents the largest coverage of documents (47.7%) and authors (52.1%), but with larger gender differences than SciELO. SciELO shows higher representation of women, particularly among exclusive authors, career lengths, and more authors publishing in Spanish (30.4%) and Portuguese (46%). These findings suggest that data sources may shape gender representation and multilingual publication practices in Latin America.

  3. Conference Paper2026Open Access
    The effect of open access on numbers of downloads and their geographic diversity(opens in a new tab)

    van Bellen, Simon; Larivière, Vincent

    Despite their usefulness in analyzing the uses of research papers within and beyond academia, the availability of download data remains scarce. Using data from the Érudit journal platform, this paper analyzes the effect of open access availability on numbers of downloads and their geographic diversity. Results show that OA positively affects the numbers of downloads by a factor of 3.0 during the embargo period, and that downloads may maintain this advantage during the following years. Eighty percent of downloads originated from non-subscribing visitors, with the end of the embargo period being strongly associated with an expansion of their geographic diversity. On the whole, our results show that journals can substantially broaden international readership by adopting immediate OA, and that metadata quality is at least as influential as OA for downloads.

  4. Working Paper2026Open Access
    English dominance in scholarly communication: Historical roots and bibliometric evidence(opens in a new tab)

    Brasil, André

    English functions as the principal shared language of international scholarly communication, but its dominance also creates inequalities in the production, dissemination, visibility, and assessment of research. Developed as a complementary output of the CoARA Working Group on Multilingualism and Language Biases in Research Assessment, this report examines the historical development of English dominance and the contemporary mechanisms through which it is reproduced. It combines evidence from the literature with descriptive bibliometric analyses based on Web of Science, Scopus, OpenAlex, and CWTS Journal Indicators. The findings show that English is especially dominant in journals with high field normalised citation impact and in the selective segments of scholarly communication most frequently used in research evaluation. However, the extent of this dominance varies considerably across disciplines and countries, and publication in national languages remains important in the social sciences, humanities, and fields directed toward local professional and societal audiences. The report also demonstrates how database selection, metadata quality, and analytical choices affect the visibility of multilingual scholarship. Web of Science and Scopus predominantly represent English language journal literature, whereas OpenAlex reveals a much broader body of multilingual research, although incomplete language, authorship, and institutional metadata remain important limitations. The report further examines the financial, linguistic, and time burdens associated with English language publication and considers how publishing infrastructures, bibliometric indicators, rankings, and assessment practices influence which research becomes visible and valued. This report contributes to the landscape analysis conducted by the Coalition for Advancing Research Assessment (CoARA) Working Group (WG) on Multilingualism and language biases in research assessment and it complements the Implementation proposal for language-aware assessments (Annex II.3).

  5. Article2026Open Access
    Interoperability between local and global databases in scientometrics: lattes, CAPES, and OpenAlex(opens in a new tab)

    Mazoni, Alysson; Borges, Luís; Macedo, Estevão Fernandes; Tuesta, Esteban Fernández; Mena‐Chalco, Jesús Pascual

    Funded by CAPES, FAPESP, UNICAMP

    Disambiguating research entities remains a long-standing methodological problem in scientometric analyses, as inconsistencies and ambiguous metadata limit the interoperability of major bibliographic databases. While global systems such as OpenAlex provide extensive coverage, they often lack the granularity and accuracy provided by national-level research databases. This study proposes a large-scale methodology to enhance author and institutional disambiguation by integrating local (Lattes and CAPES) and global (OpenAlex) databases. The method combines shared Digital Object Identifiers with an adapted Levenshtein distance algorithm to handle variations in author and institutional names across multilingual research databases, achieving over 97% accuracy for authors and 64% for institutions. The proposed framework provides a scalable and replicable approach for entity disambiguation in tabular research databases. Beyond the Brazilian context, this integration strategy offers a globally applicable approach for harmonizing national research information systems with open scientometric infrastructures.

  6. Working Paper2025Open Access
    Exploring Interoperability Between Local and Global Databases in Scientometrics: Lattes, Capes, and OpenAlex(opens in a new tab)

    Mazoni, Alysson Fernandes; Borges, Luís Fabiano Farias; Macedo, Estevao Fernandes; Tuesta, Esteban Fernandez

    Numerous initiatives are currently underway to disambiguate databases worldwide. In this paper, we propose a methodology for disambiguating research entities using big data techniques, adopting an approach that goes from local to global databases. Our objective is to enhance the quality of data in the OpenAlex database by leveraging information from Brazilian databases, particularly data from the Lattes Platform and the Brazilian Federal Agency for Support and Evaluation of Graduate Education. We compare similar names of authors and institutions, employing Digital Object Identifiers to link entities, along with an adaptation of the Levenshtein distance algorithm. The proposed method is straightforward to implement in tabular databases and facilitates disambiguation, thereby contributing to open science practices and providing an effective solution for research information systems. The findings indicate the potential for integrating local and global databases to address issues related to ambiguous names and incomplete metadata.