Human Biographical Record (HBR)
9 Pages Posted: 15 Dec 2020 Last revised: 17 Feb 2021
Date Written: December 2, 2020
We construct a new dataset of more than seven million notable individuals across recorded human history, the Human Biographical Record (HBR). With Wikidata as the backbone, HBR adds further information from various digital sources, including Wikipedia in all 292 languages. Machine learning and text analysis combine the sources and extract information on date and place of birth and death, gender, occupation, education, and family background. This paper discusses HBR's construction and its completeness, coverage, accuracy, and also its strength and weakness relative to prior datasets. HBR is the first part of a larger project, the human record project that we briefly introduce.
Keywords: Bid data, machine learning, economic history
JEL Classification: N00,
Suggested Citation: Suggested Citation