De Freitas, Sergio Antonio Andrade; da Silva, André Corrêa; Silva, Milena de Faria; Batista, Renan Carneiro; Segundo, Washington
Datasets, working papers, and peer-reviewed articles produced by MultiObs researchers.
Showing 1–3 of 3
De Freitas, Sergio Antonio Andrade; da Silva, André Corrêa; Silva, Milena de Faria; Batista, Renan Carneiro; Segundo, Washington
Bascur, Juan Pablo; Costas, Rodrigo; Verberne, Suzan
Purpose Traditional science maps cluster documents into topics but are inherently biased toward clustering certain topics over others. This study investigates the extent to which topic bias can be influenced through the selection of data sources used to construct document networks. Design/methodology/approach We evaluate the clustering effectiveness of several topic categories using document networks constructed from two traditional science mapping sources (citations and text similarity) and six non-traditional data sources (policy documents, patent families, document authors, Facebook users, Twitter users, and Twitter conversations). Each source is evaluated both independently and in combination with a text similarity network. Findings Different data sources favor different kinds of topics. Facebook users favor health issues, patent families favor biotechnology topics, policy documents favor government and social issues, Twitter conversations favor food topics, Twitter users favor nursing topics, and document authors favor geographical entities. These findings demonstrate that topic bias can be systematically influenced through data source selection. Research limitations The study focuses on biomedical publications and is limited to the topic categories and data sources examined. Additional domains, data sources, and source combination methods may exhibit different patterns of topic bias. Practical implications The ability to influence topic bias through data source selection opens up the possibility of creating science maps tailored to different information needs. The reported source-specific biases can support the design of science maps optimized for particular users, tasks, or domains. Originality/value This study provides one of the first large-scale investigations of how alternative data sources affect topic emergence in science maps. It introduces an expanded methodology for evaluating topic-level clustering effectiveness and systematically characterizes the topical biases associated with different data sources, providing a foundation for future science map customization.
Maruyama, William Takahiro; Digiampietri, Luciano A.
The increasing use of digital data in electoral prediction has motivated a growing body of computational research, yet the field remains methodologically diverse and lacks consolidated comparative frameworks. This article presents a systematic review of computational approaches for electoral outcome prediction using digital data between 2020 and 2025. Following rigorous systematic methodology, searches were conducted across three scientific databases, resulting in 80 primary studies analyzed after applying explicit quality criteria. The review proposes a taxonomy classifying studies by data integration and predictive complexity, enabling systematic identification of methodological patterns. Results reveal geographic concentration in few countries, with Twitter as the dominant platform and sentiment analysis as the most frequent technique. Vote percentage prediction and winner identification represent the primary objectives, evaluated mainly through regression and classification metrics. The field demonstrates numerical expansion with modest geographic diversification, yet persistent challenges remain regarding sample representativeness, cross-context generalization, and absence of standardized validation protocols. Findings indicate the need for broader geographic coverage, reduced platform dependency, and establishment of uniform evaluation criteria to advance methodological maturity in computational electoral prediction.