SOURCES & METHODS
Behind the numbers
The scale of the network. The reach of its research. Here is what we count, how we count it, and what each number can—and cannot—tell us.
These explanations accompany the figures displayed on the homepage, plus the count of citations made directly to dataset DOIs, which is kept separate from scholarly citations. Usage figures cover Harvard Dataverse only and start in 2020. They are not live counters. Network totals describe multiple installations; the citation analysis covers a defined set of Harvard Dataverse datasets.
HARVARD DATAVERSE · RESEARCH REACH
Scholarly Citations
3,946,325Citations received by publications linked to the datasets in the analysis. Each linked publication DOI contributes its citation count once, even when multiple datasets link to it.
From shared data to scholarly reach
Connect datasets to publications
Start with the publication links associated with the 111,651 Harvard Dataverse datasets in the analysis. Resolve publication identifiers and remove repeated links.
Choose one count per publication
For each linked publication DOI with a successful lookup, use the largest available provider citation count. If counts tie, prefer the most recently updated record. Do not add provider counts together.
Sum across distinct publication DOIs
Add the selected counts once per linked publication DOI across the analysis—not once per dataset. A publication shared by several datasets contributes only once.
What “deduplicated” means here
The cited publication is deduplicated. The citing papers are not deduplicated across different publications. A paper that cites two linked publications can contribute to both counts. This total does not establish that every citing paper reused the underlying data.
Sources and coverage
Publication links and stored provider results from Crossref, OpenAlex, and Semantic Scholar support the calculation. Missing links, unresolved identifiers, and differences in provider coverage affect the result.
THE GLOBAL COMMUNITY
Dataverse installations
150+An installation is an independently operated Dataverse repository. It is not a collection, a dataset, or a server within an installation.
The registry extract used for this site’s map contains 150 installation entries. The headline retains the project’s “150+” wording; the map is not a continuously updated census.
Where installations are based
Dataverse Network
Explore the interactive map ↗Browse all 150 repositories (in alphabetical order)
- Abacus — Canada
- ACSS Dataverse — Lebanon
- ADA Dataverse — Australia
- ADP - Slovenian Social Science Data Archives — Slovenia
- Arca Dados — Brazil
- ARP — Hungary
- ASU Library Research Data Repository — USA
- AUSSDA Dataverse — Austria
- BioData.pt Data Management Portal (DMPortal) — Portugal
- bonndata — Germany
- Borealis — Canada
- Botswana Harvard Data — Botswana
- BRIN Dataverse — Indonesia
- BSC Dataverse — Spain
- CedapDados — Brazil
- Centro Brasileiro de Pesquisas Físicas - CBPF — Brazil
- CESA | Repositorio de datos académicos — Colombia
- CIDACS — Brazil
- CIFOR — Indonesia
- CIMMYT Research Data — Mexico
- CIRAD Dataverse — France
- CORA. Research Data Repository (RDR) — Spain
- CoronaWhy Dataverse — Netherlands
- CROSSDA — Croatia
- CSDA Dataverse — Czechia
- CUHK Research Data Repository — Hong Kong
- DADOS IPB — Portugal
- Dane Badawcze UW — Poland
- DANS Data Station Archaeology — Netherlands
- DANS Data Station Life Sciences — Netherlands
- DANS Data Station Physical and Technical Sciences — Netherlands
- DANS Data Station Social Sciences and Humanities — Netherlands
- dare — Germany
- Dartmouth Dataverse — USA
- DaRUS — Germany
- Data Suds — France
- data.sciencespo — France
- DATADOI — Estonia
- DataPB — Brazil
- dataportal.ing.pan.pl — Poland
- DataRepositoriUM — Portugal
- DataSpace@HKUST — Hong Kong
- Dataverse e-cienciaDatos — Spain
- DataverseLV — Latvia
- DataverseNL — Netherlands
- DataverseNO — Norway
- DataverseUA — Ukraine
- DATICE — Iceland
- Datos para Resiliencia — Chile
- DeiC Dataverse — Denmark
- Domus Dados — Brazil
- DR-NTU (Data) — Singapore
- DUnAs — Portugal
- Edmond — Germany
- FGV Dataverse — Brazil
- Florida International University Research Data Portal — USA
- Fudan University — China
- George Mason University Dataverse — USA
- Göttingen Research Online — Germany
- Gustave Eiffel University Dataverse — France
- Harvard Dataverse — USA
- Health Study Hub — Germany
- HeiDATA — Germany
- IBICT — Brazil
- ICRISAT — India
- ICWSM — Germany
- IDSC Dataverse — Germany
- IFDC Dataverse — USA
- IISH Dataverse — Netherlands
- Indata — Ecuador
- Institute of Russian Literature Dataverse — Russia
- International Potato Center — Peru
- ioerDATA — Germany
- IPGP Research Collection — France
- ISSDA Dataverse — Ireland
- Italian Institute of Technology (IIT) — Italy
- Johns Hopkins Research Data Repository — USA
- JPL Open Repository — USA
- Jülich DATA — Germany
- KEEN Data Management Platform — Germany
- KU Leuven RDR — Belgium
- Libra Data — USA
- Lithuanian Data Archive for Social Sciences and Humanities (LiDA) — Lithuania
- LORE - LIST Open Repository — Luxembourg
- Maine Dataverse Network — USA
- MBLWHOI Library Dataverse — USA
- MELDATA — Lebanon
- Ministerio de las Culturas, las Artes y los Saberes — Colombia
- NIE Data Repository — Singapore
- NIOZ Dataverse — Netherlands
- NYCU Dataverse — Taiwan (ROC)
- ODISSEI Portal — Netherlands
- Open Data @ UCLouvain — Belgium
- Open Forest Data — Poland
- osnaData — Germany
- PAPYRUS — Colombia
- Peking University — China
- POLEN DataHub — Portugal
- PolyU Research Data Repository — Hong Kong
- Pontificia Universidad Católica del Perú — Peru
- QDR Main Collection — USA
- Recherche Data Gouv — France
- Redape - Repositório de Dados de Pesquisa da Embrapa — Brazil
- RepOD — Poland
- Repositório de Dados de Pesquisa da UFABC — Brazil
- Repositório de Dados de Pesquisa do ILEEL — Brazil
- Repositorio de Datos Abiertos de Investigación (Redata) — Uruguay
- Repositorio de Datos Académicos RDA-UNR — Argentina
- Repositorio de datos de investigación de la Universidad de Chile — Chile
- Repositorio de Datos de Investigación de la Universidad Nacional de La Plata — Argentina
- Repositorio de datos de Investigación UdeA — Colombia
- Repositorio de Datos de Investigación Universidad del Rosario — Colombia
- Repositorio de Datos de Investigación USACH — Chile
- Repositorio de datos de la Universidad de Concepción — Chile
- Repositorio de Datos de la Universidad del Pacífico (DatasetsUP) — Peru
- Repositorio de Datos Pontificia Universidad Javeriana — Colombia
- Repositorio de Datos Universidad Distrital Francisco José de Caldas — Colombia
- Reposítorio SoilData — Brazil
- Repositorio TECdatos — Costa Rica
- Repositórios Piloto da Rede Nacional de Ensino e Pesquisa — Brazil
- RODBUCK UKEN — Poland
- RODBUK — Poland
- RODBUK AGH — Poland
- RODBUK PK — Poland
- RODBUK UEK — Poland
- RODBUK UJ — Poland
- RSU Dataverse — Latvia
- Sano — Poland
- SciELO Data — Brazil
- SODHA — Belgium
- Tecnológico de Monterrey Data Hub — Mexico
- Texas Data Repository Dataverse — USA
- The Henryk Niewodniczański Institute of Nuclear Physics Polish Academy of Sciences — Poland
- TRR170-DB — Germany
- TUDOdata — Germany
- UC Berkeley Library Dataverse — USA
- UCLA Dataverse — USA
- UD Dataverse — USA
- ULiège Open Data Repository — Belgium
- UNB Libraries Dataverse — Canada
- UNC Dataverse — USA
- Università Ca’ Foscari Venezia Datarepository — Italy
- Università degli Studi di Milano — Italy
- University of Physical Culture in Krakow — Poland
- University of Wroclaw — Poland
- USC Dataverse — USA
- VTTI — USA
- World Agroforestry - Research Data Repository — Kenya
- Wyoming Data Repository — USA
- Yale Dataverse — USA
Locations: Dataverse community registry. Geography: Natural Earth.
Counting rule
Count repository entries in the community-maintained registry. Several installations in one country are separate entries; collections within a repository do not increase the installation count.
Coverage limit
A registry entry does not by itself confirm current availability or whether a repository’s metrics service responds. The installation count and the number of repositories contributing dataset totals are different measures.
ACROSS THE NETWORK
Total Datasets in the Network
589,645+The reported aggregate of published dataset counts from responding installations. This is a repository-record total, not a verified count of unique datasets across all installations.
From repositories to a network total
How the total is assembled
Request the dataset metric from each installation and sum the returned counts. An unavailable response is missing information, not a zero. A collection’s count must not be added again when it is already included in its installation’s total.
What is not yet independently reproducible
The original site record attributes this total to 123 responding installations. The installation-by-installation responses and query filters were not retained with the site, so we cannot show an audited breakdown or confirm whether harvested records were excluded. The total should not be described as a complete, cross-repository-deduplicated census.
HARVARD DATAVERSE · DATASET DOIs
Citations of the data itself
16,947The sum of DataCite’s recorded citation counts for dataset DOIs in the Harvard Dataverse analysis. This measure is separate from citations received by linked publications.
Dataset coverage in the analysis
Identify dataset DOIs
Use the 111,651 dataset records included in the analysis, with one entry per dataset identifier.
Read DataCite counts
Match dataset DOIs to DataCite records and read the citationCount field. 111,651 records matched; 0 had no matching record.
Sum the recorded counts
Add the recorded citation counts across dataset identifiers. Missing records contribute no observed citations, but remain a coverage gap rather than evidence of no use.
What counts as a citation?
DataCite aggregates DOI relationships, including citation, reference, and supplement relationships. Its rules avoid counting equivalent links twice for the same DOI pair. The sum is not a count of unique citing papers across datasets.
DataCite’s citation definitions and counting rulesKeep the measures separate
Dataset Citations identify relationships to dataset DOIs. Scholarly Citations measure the citation reach of linked publications. Their scopes can overlap; adding them would not produce a deduplicated count of citing papers.
HARVARD DATAVERSE REPOSITORY
Published datasets
118,106The Harvard Dataverse dataset metric recorded for this preview. It is a repository-wide holdings measure, not the size of the citation-analysis cohort.
Two different scopes
Read the dataset metric
The repository’s /api/info/metrics/datasets endpoint returns a count. The metric describes released datasets, not collections or files; unpublished and deaccessioned versions are excluded.
The source endpoint can change independently of this preview’s displayed value.
Check the query scope
Dataset metrics support local, harvested, or combined records through the dataLocation filter. The original source link has no explicit filter. We do not present the recorded total as a DOI-deduplicated count across repositories.
HARVARD DATAVERSE · USAGE SINCE 2020
Views and downloads
16,292,190Unique dataset views by people since mid-2020, when Harvard Dataverse began reporting usage under the Make Data Count standard. “Unique” means one count per session, per dataset, per month, so this is a count of visits, not of distinct people.
Usage of Harvard Dataverse since 2020
| Measure | All traffic | Unique | Unique, by people |
|---|---|---|---|
| Dataset views | 126,533,697 | 25,960,025 | 16,292,190 |
| Downloads | 89,429,116 | 3,626,111 | 2,895,422 |
Record usage events
Every dataset page view and file download on Harvard Dataverse is logged as an event under the COUNTER Code of Practice for Research Data, the standard behind Make Data Count.
Collapse repeats into unique visits
Repeated views or downloads of the same dataset within one session and month count once. The result is a count of visits to datasets, not a count of distinct people.
Separate people from machines
COUNTER classifies known crawlers, harvesters and API clients as machine traffic. The headline figures use the human share; the table shows both.
Since 2020, Harvard only
Make Data Count reporting started on Harvard Dataverse in mid-2020, so these totals cover only the years since. Only 14 of the network’s installations publish usage this way, so no comparable network total exists; the sum of what those installations report (32,472,322 unique views) is a floor, not a census.
Sources
The repository’s /api/info/metrics/makeDataCount/ endpoints for total, unique and “regular” (human) views and downloads. The all-time file-download counter, which predates 2020 and is not deduplicated, is reported separately by the metrics API.