Harvard Dataverse ↗
Dataverse Project

SOURCES & METHODS

Behind the numbers

The scale of the network. The reach of its research. Here is what we count, how we count it, and what each number can—and cannot—tell us.

These explanations accompany the figures displayed on the homepage, plus the count of citations made directly to dataset DOIs, which is kept separate from scholarly citations. Usage figures cover Harvard Dataverse only and start in 2020. They are not live counters. Network totals describe multiple installations; the citation analysis covers a defined set of Harvard Dataverse datasets.

HARVARD DATAVERSE · RESEARCH REACH

Scholarly Citations

3,946,325

Citations received by publications linked to the datasets in the analysis. Each linked publication DOI contributes its citation count once, even when multiple datasets link to it.

From shared data to scholarly reach

Shared datasetsThe foundation
Linked publicationsThe research
Scholarly CitationsThe reach
38,890datasets connected to publications with recorded citations
See the wider reach of data-associated research through the publications connected to Harvard Dataverse datasets.
  1. Connect datasets to publications

    Start with the publication links associated with the 111,651 Harvard Dataverse datasets in the analysis. Resolve publication identifiers and remove repeated links.

  2. Choose one count per publication

    For each linked publication DOI with a successful lookup, use the largest available provider citation count. If counts tie, prefer the most recently updated record. Do not add provider counts together.

  3. Sum across distinct publication DOIs

    Add the selected counts once per linked publication DOI across the analysis—not once per dataset. A publication shared by several datasets contributes only once.

What “deduplicated” means here

The cited publication is deduplicated. The citing papers are not deduplicated across different publications. A paper that cites two linked publications can contribute to both counts. This total does not establish that every citing paper reused the underlying data.

Sources and coverage

Publication links and stored provider results from Crossref, OpenAlex, and Semantic Scholar support the calculation. Missing links, unresolved identifiers, and differences in provider coverage affect the result.

THE GLOBAL COMMUNITY

Dataverse installations

150+

An installation is an independently operated Dataverse repository. It is not a collection, a dataset, or a server within an installation.

The registry extract used for this site’s map contains 150 installation entries. The headline retains the project’s “150+” wording; the map is not a continuously updated census.

Where installations are based

United States21
Germany15
Brazil14
Poland14
Netherlands9
Colombia7
Other countries and territories70
Entries grouped by the registry’s country or territory field. Every entry in this site’s map contributes once; the bars sum to 150.
Go to the alphabetical repository list
Dataverse installations around the worldLocations of 150 installations in the Dataverse community registry, across six continents. Select a dot to open its repository in a new tab. Use Tab to reach individual repositories where dots overlap.Abacus — CanadaACSS Dataverse — LebanonADA Dataverse — AustraliaADP - Slovenian Social Science Data Archives — SloveniaArca Dados — BrazilARP — HungaryASU Library Research Data Repository — USAAUSSDA Dataverse — AustriaBioData.pt Data Management Portal (DMPortal) — Portugalbonndata — GermanyBorealis — CanadaBotswana Harvard Data — BotswanaBRIN Dataverse — IndonesiaBSC Dataverse — SpainCedapDados — BrazilCentro Brasileiro de Pesquisas Físicas - CBPF — BrazilCESA | Repositorio de datos académicos — ColombiaCIDACS — BrazilCIFOR — IndonesiaCIMMYT Research Data — MexicoCIRAD Dataverse — FranceCORA. Research Data Repository (RDR) — SpainCoronaWhy Dataverse — NetherlandsCROSSDA — CroatiaCSDA Dataverse — CzechiaCUHK Research Data Repository — Hong KongDADOS IPB — PortugalDane Badawcze UW — PolandDANS Data Station Archaeology — NetherlandsDANS Data Station Life Sciences — NetherlandsDANS Data Station Physical and Technical Sciences — NetherlandsDANS Data Station Social Sciences and Humanities — Netherlandsdare — GermanyDartmouth Dataverse — USADaRUS — GermanyData Suds — Francedata.sciencespo — FranceDATADOI — EstoniaDataPB — Brazildataportal.ing.pan.pl — PolandDataRepositoriUM — PortugalDataSpace@HKUST — Hong KongDataverse e-cienciaDatos — SpainDataverseLV — LatviaDataverseNL — NetherlandsDataverseNO — NorwayDataverseUA — UkraineDATICE — IcelandDatos para Resiliencia — ChileDeiC Dataverse — DenmarkDomus Dados — BrazilDR-NTU (Data) — SingaporeDUnAs — PortugalEdmond — GermanyFGV Dataverse — BrazilFlorida International University Research Data Portal — USAFudan University — ChinaGeorge Mason University Dataverse — USAGustave Eiffel University Dataverse — FranceGöttingen Research Online — GermanyHarvard Dataverse — USAHealth Study Hub — GermanyHeiDATA — GermanyIBICT — BrazilICRISAT — IndiaICWSM — GermanyIDSC Dataverse — GermanyIFDC Dataverse — USAIISH Dataverse — NetherlandsIndata — EcuadorInstitute of Russian Literature Dataverse — RussiaInternational Potato Center — PeruioerDATA — GermanyIPGP Research Collection — FranceISSDA Dataverse — IrelandItalian Institute of Technology (IIT) — ItalyJohns Hopkins Research Data Repository — USAJPL Open Repository — USAJülich DATA — GermanyKEEN Data Management Platform — GermanyKU Leuven RDR — BelgiumLibra Data — USALithuanian Data Archive for Social Sciences and Humanities (LiDA) — LithuaniaLORE - LIST Open Repository — LuxembourgMaine Dataverse Network — USAMBLWHOI Library Dataverse — USAMELDATA — LebanonMinisterio de las Culturas, las Artes y los Saberes — ColombiaNIE Data Repository — SingaporeNIOZ Dataverse — NetherlandsNYCU Dataverse — Taiwan (ROC)ODISSEI Portal — NetherlandsOpen Data @ UCLouvain — BelgiumOpen Forest Data — PolandosnaData — GermanyPAPYRUS — ColombiaPeking University — ChinaPOLEN DataHub — PortugalPolyU Research Data Repository — Hong KongPontificia Universidad Católica del Perú — PeruQDR Main Collection — USARecherche Data Gouv — FranceRedape - Repositório de Dados de Pesquisa da Embrapa — BrazilRepOD — PolandRepositorio de Datos Abiertos de Investigación (Redata) — UruguayRepositorio de Datos Académicos RDA-UNR — ArgentinaRepositorio de datos de investigación de la Universidad de Chile — ChileRepositorio de Datos de Investigación de la Universidad Nacional de La Plata — ArgentinaRepositorio de datos de Investigación UdeA — ColombiaRepositorio de Datos de Investigación Universidad del Rosario — ColombiaRepositorio de Datos de Investigación USACH — ChileRepositorio de datos de la Universidad de Concepción — ChileRepositorio de Datos de la Universidad del Pacífico (DatasetsUP) — PeruRepositorio de Datos Pontificia Universidad Javeriana — ColombiaRepositorio de Datos Universidad Distrital Francisco José de Caldas — ColombiaRepositorio TECdatos — Costa RicaRepositório de Dados de Pesquisa da UFABC — BrazilRepositório de Dados de Pesquisa do ILEEL — BrazilRepositórios Piloto da Rede Nacional de Ensino e Pesquisa — BrazilReposítorio SoilData — BrazilRODBUCK UKEN — PolandRODBUK — PolandRODBUK AGH — PolandRODBUK PK — PolandRODBUK UEK — PolandRODBUK UJ — PolandRSU Dataverse — LatviaSano — PolandSciELO Data — BrazilSODHA — BelgiumTecnológico de Monterrey Data Hub — MexicoTexas Data Repository Dataverse — USAThe Henryk Niewodniczański Institute of Nuclear Physics Polish Academy of Sciences — PolandTRR170-DB — GermanyTUDOdata — GermanyUC Berkeley Library Dataverse — USAUCLA Dataverse — USAUD Dataverse — USAULiège Open Data Repository — BelgiumUNB Libraries Dataverse — CanadaUNC Dataverse — USAUniversity of Physical Culture in Krakow — PolandUniversity of Wroclaw — PolandUniversità Ca’ Foscari Venezia Datarepository — ItalyUniversità degli Studi di Milano — ItalyUSC Dataverse — USAVTTI — USAWorld Agroforestry - Research Data Repository — KenyaWyoming Data Repository — USAYale Dataverse — USA
Browse all 150 repositories (in alphabetical order)

Locations: Dataverse community registry. Geography: Natural Earth.

Counting rule

Count repository entries in the community-maintained registry. Several installations in one country are separate entries; collections within a repository do not increase the installation count.

Coverage limit

A registry entry does not by itself confirm current availability or whether a repository’s metrics service responds. The installation count and the number of repositories contributing dataset totals are different measures.

ACROSS THE NETWORK

Total Datasets in the Network

589,645+

The reported aggregate of published dataset counts from responding installations. This is a repository-record total, not a verified count of unique datasets across all installations.

From repositories to a network total

Community registry150+ installations
Responses recorded for this total118 installations
Reported aggregate589,645 datasets
“+” indicates incomplete network coverage, not an estimate of the missing repositories’ holdings.

How the total is assembled

Request the dataset metric from each installation and sum the returned counts. An unavailable response is missing information, not a zero. A collection’s count must not be added again when it is already included in its installation’s total.

What is not yet independently reproducible

The original site record attributes this total to 123 responding installations. The installation-by-installation responses and query filters were not retained with the site, so we cannot show an audited breakdown or confirm whether harvested records were excluded. The total should not be described as a complete, cross-repository-deduplicated census.

HARVARD DATAVERSE · DATASET DOIs

Citations of the data itself

16,947

The sum of DataCite’s recorded citation counts for dataset DOIs in the Harvard Dataverse analysis. This measure is separate from citations received by linked publications.

Dataset coverage in the analysis

Datasets with a recorded Dataset Citation12,410
Datasets with no recorded Dataset Citation99,241
12,410 datasets account for 16,947 Dataset Citations. No recorded citation is not proof that a dataset has never been cited or used.
  1. Identify dataset DOIs

    Use the 111,651 dataset records included in the analysis, with one entry per dataset identifier.

  2. Read DataCite counts

    Match dataset DOIs to DataCite records and read the citationCount field. 111,651 records matched; 0 had no matching record.

  3. Sum the recorded counts

    Add the recorded citation counts across dataset identifiers. Missing records contribute no observed citations, but remain a coverage gap rather than evidence of no use.

What counts as a citation?

DataCite aggregates DOI relationships, including citation, reference, and supplement relationships. Its rules avoid counting equivalent links twice for the same DOI pair. The sum is not a count of unique citing papers across datasets.

DataCite’s citation definitions and counting rules

Keep the measures separate

Dataset Citations identify relationships to dataset DOIs. Scholarly Citations measure the citation reach of linked publications. Their scopes can overlap; adding them would not produce a deduplicated count of citing papers.

HARVARD DATAVERSE REPOSITORY

Published datasets

118,106

The Harvard Dataverse dataset metric recorded for this preview. It is a repository-wide holdings measure, not the size of the citation-analysis cohort.

Two different scopes

Published datasets shown on the homepage118,106
Dataset records in the citation analysis111,651
Different source scopes, not a growth comparison or a citation coverage percentage. Do not substitute one denominator for the other.

Read the dataset metric

The repository’s /api/info/metrics/datasets endpoint returns a count. The metric describes released datasets, not collections or files; unpublished and deaccessioned versions are excluded.

The source endpoint can change independently of this preview’s displayed value.

Check the query scope

Dataset metrics support local, harvested, or combined records through the dataLocation filter. The original source link has no explicit filter. We do not present the recorded total as a DOI-deduplicated count across repositories.

HARVARD DATAVERSE · USAGE SINCE 2020

Views and downloads

16,292,190

Unique dataset views by people since mid-2020, when Harvard Dataverse began reporting usage under the Make Data Count standard. “Unique” means one count per session, per dataset, per month, so this is a count of visits, not of distinct people.

Usage of Harvard Dataverse since 2020

MeasureAll trafficUniqueUnique, by people
Dataset views126,533,69725,960,02516,292,190
Downloads89,429,1163,626,1112,895,422
Counts refresh daily from the repository’s Make Data Count metrics. “All traffic” counts every event; “unique” counts one per session, dataset and month; “by people” excludes machine traffic such as crawlers and API harvesters.
  1. Record usage events

    Every dataset page view and file download on Harvard Dataverse is logged as an event under the COUNTER Code of Practice for Research Data, the standard behind Make Data Count.

  2. Collapse repeats into unique visits

    Repeated views or downloads of the same dataset within one session and month count once. The result is a count of visits to datasets, not a count of distinct people.

  3. Separate people from machines

    COUNTER classifies known crawlers, harvesters and API clients as machine traffic. The headline figures use the human share; the table shows both.

Since 2020, Harvard only

Make Data Count reporting started on Harvard Dataverse in mid-2020, so these totals cover only the years since. Only 14 of the network’s installations publish usage this way, so no comparable network total exists; the sum of what those installations report (32,472,322 unique views) is a floor, not a census.

Sources

The repository’s /api/info/metrics/makeDataCount/ endpoints for total, unique and “regular” (human) views and downloads. The all-time file-download counter, which predates 2020 and is not deduplicated, is reported separately by the metrics API.