Logo do repositório
 
Publicação

Datasheets for Digital Cultural Heritage Datasets

dc.contributor.authorAlkemade, Henk
dc.contributor.authorClaeyssens, Steven
dc.contributor.authorColavizza, Giovanni
dc.contributor.authorFreire, Nuno
dc.contributor.authorLehmann, Jörg
dc.contributor.authorNeudecker, Clemens
dc.contributor.authorOsti, Giulia
dc.contributor.authorvan Strien, Daniel
dc.contributor.institutionFaculdade de Ciências Sociais e Humanas (FCSH)
dc.contributor.pblUbiquity Press
dc.date.accessioned2024-04-09T00:56:34Z
dc.date.available2024-04-09T00:56:34Z
dc.date.issued2023
dc.descriptionCNECT/LUX/2021/OP/0070 [2022-2024
dc.description.abstractSparked by issues of quality and lack of proper documentation for datasets, the machine learning community has begun developing standardised processes for establishing datasheets for machine learning datasets, with the intent to provide context and information on provenance, purposes, composition, the collection process, recommended uses or societal biases reflected in training datasets. This approach fits well with practices and procedures established in GLAM institutions, such as establishing collections' descriptions. However, digital cultural heritage datasets are marked by specific characteristics. They are often the product of multiple layers of selection; they may have been created for different purposes than establishing a statistical sample according to a specific research question; they change over time and are heterogeneous. Punctuated by a series of recommendations to create datasheets for digital cultural heritage, the paper addresses the scope and characteristics of digital cultural heritage datasets; possible metrics and measures; lessons from concepts similar to datasheets and/or established workflows in the cultural heritage sector. This paper includes a proposal for a datasheet template that has been adapted for use in cultural heritage institutions, and which proposes to incorporate information on the motivation and selection criteria, digitisation pipeline, data provenance, the use of linked open data, and version information.en
dc.description.versionpublishersversion
dc.description.versionpublished
dc.format.extent11
dc.format.extent569634
dc.identifier.doi10.5334/johd.124
dc.identifier.issn2059-481X
dc.identifier.otherPURE: 87840842
dc.identifier.otherPURE UUID: d598cfff-0f33-489d-8457-d9eff3f6e8e8
dc.identifier.otherScopus: 85177736331
dc.identifier.urihttp://hdl.handle.net/10362/165967
dc.identifier.urlhttps://www.scopus.com/pages/publications/85177736331
dc.identifier.urlhttps://openhumanitiesdata.metajnl.com/articles/10.5334/johd.124
dc.language.isoeng
dc.peerreviewedyes
dc.subjectDatasets
dc.subjectDatasheets
dc.subjectDigital cultural heritage
dc.subjectGLAM institutions
dc.subjectMachine learning
dc.subjectModel cards
dc.subjectInformation Systems
dc.subjectGeneral Arts and Humanities
dc.subjectLibrary and Information Sciences
dc.titleDatasheets for Digital Cultural Heritage Datasetsen
dc.typejournal article
degois.publication.firstPage1
degois.publication.lastPage11
degois.publication.titleJournal of Open Humanities Data
degois.publication.volume9
dspace.entity.typePublication
rcaap.rightsopenAccess

Ficheiros

Principais
A mostrar 1 - 1 de 1
A carregar...
Miniatura
Nome:
65531b13159a7.pdf
Tamanho:
556.28 KB
Formato:
Adobe Portable Document Format