multiobs

Data Infrastructure

Federated Data Infrastructure

Open datasets for the study and monitoring of Science, Technology and Innovation, published on Google BigQuery.

Overview

What the infrastructure is

The federated Big Data Research Infrastructure is the first of the three components of MultiObs. It brings together the largest available set of interconnected databases and analytical tools to support a wide varierty of data-driven questions on the dynamics of Science, Technology and Innovation.

Bibliometric, economic, funding and intellectual property datasets are collected and structured for academic use, enabling cross-analysis to drive future global institutional research. The infrastructure supports the four thematic branches of the project and the Living Lab, and it is open to researchers beyond the MultiObs team.

Access

How to access the data

Our datasets live in the public Google BigQuery project multiobs. They can be queried with SQL directly from the BigQuery console, with no need to download or host anything yourself.

1

Get a Google Cloud project

Use an existing GCP project or create one.

2

Open the BigQuery sandbox

Use BigQuery interface or any database management tool.

Open the sandbox
3

Query the multiobs project

Run SQL against our public datasets.

Dataset documentation, update times, row counts and sizes are maintained on ORION-DBs, which is rebuilt daily. MultiObs is not affiliated with Google.

Community

Part of ORION-DBs

MultiObs is one of the contributing collections of ORION-DBs, a community of independent groups that host open scholarly data on Google BigQuery. Rather than each group separately maintaining copies of the same core sources, the groups share the load — coordinating storage, preprocessing and documentation so that key open research information resources stay combinable at scale.

Each project covers different sources or versions, so users benefit from the combination without any single group having to maintain everything. For the motivation behind the effort, see the collective's mission statement.

Collections in the network

People

Infrastructure Team

The people who build and maintain the pipelines, datasets and BigQuery infrastructure behind MultiObs.

Mariana Moretti

Mariana Moretti

Data Architecture & Analytics Engineering

Mariana's role involves the technical aspects of MultiObs data infrastructure such as Data Architecture: Designing databases, modeling data structures and defining schemas; and Analytical Engineering: Developing SQL-based workflows to organize and structure complex datasets. Working alongside technicians and researchers, she contributes to the conversion of scientific requirements into robust data structures and workflows. She also conducts training sessions on data infrastructure and practical dataset usage in line with MultiObs' knowledge-sharing initiatives.

Marcelo de Barros

Marcelo de Barros

Cloud Infrastructure & Pipelines

Marcelo works at the collection and infrastructure end of MultiObs, building and maintaining the pipelines that ingest large-scale sources, from bibliographic databases to Wikipedia and Bluesky. He administers the Google Cloud environment where these datasets are stored, processed and published, including BigQuery and Cloud Storage, and evaluates new sources, testing collection and integration strategies. He designs and documents the pipelines and processes he takes part in, producing the technical reference that supports reproducibility and onboarding across the team. He also developed and maintains the MultiObs website.

Jorge Borges

Jorge Borges

Data Collection & Ingestion

Jorge works on the collection and ingestion of a significant portion of the datasets available through MultiObs. Using Python, APIs, web scraping, and other data collection methods, he brings external data sources into the project and publishes them to Google BigQuery. His work also includes initial validation, deduplication, and first-stage data processing, helping ensure that the datasets are ready for further treatment and analysis by the research team.

MultiObs is a São Paulo Chair of Excellence (SPEC) project funded by the São Paulo Research Foundation (FAPESP) and hosted at the State University of Campinas (UNICAMP).