Uploaded March 2019 | Updated September 2026, 10 hours ago
The University of Cambridge Herbarium has over a million specimens (including nearly 1000 personally collected by Charles Darwin). Many of these have alternative names and classifications. Public resources such as theplantlist.org have brought together a consensus list of names and their alternatives, but many specimens are stored under their old names. Your task is to use machine learning methods to optimise the indexing of the Herbarium specimens, provide simpler and more intuitive retrieval, and visualise the relationships between parts of the collection, using data from their databases, theplantlist.org and other scientific resources.
The University of Cambridge Herbarium has over a million specimens (including nearly 1000 personally collected by Charles Darwin). Many of these have alternative names and classifications. Public resources such as theplantlist.org have brought together a consensus list of names and their alternatives, but many specimens are stored under their old names. Your task is to use machine learning methods to optimise the indexing of the Herbarium specimens, provide simpler and more intuitive retrieval, and visualise the relationships between parts of the collection, using data from their databases, theplantlist.org and other scientific resources.










