Logo do repositório
 
Publicação

Intelligent Support for Low Literacy Adults

dc.contributor.authorReis, Maria Leonor
dc.contributor.authorBarbosa, Sílvia
dc.contributor.authorMoutinho, Michell
dc.contributor.authorMonteiro, Ricardo
dc.contributor.authorAmaro, Raquel
dc.contributor.authorCorreia, Susana Mesquita de Deus
dc.contributor.institutionCentro de Linguística da UNL (CLUNL)
dc.contributor.institutionDepartamento de Linguística (DL)
dc.contributor.pblInternational Association of Online Engineering
dc.date.accessioned2025-01-13T21:24:12Z
dc.date.available2025-01-13T21:24:12Z
dc.date.issued2024
dc.descriptionUIDB/03213/2020 UIDP/03213/2020
dc.description.abstractThis paper presents the Portuguese dataset of the iRead4Skills project (Dataset 1: corpora by complexity level for FR, PT, and SP – v.2.0), a representative sample of written European Portuguese for automatic complexity assessment that addresses a gap in existing resources for Portuguese. The corpus was created within the framework of the iRead4Skills project, which encompasses Portuguese, French, and Spanish. The project aims to develop an intelligent system to evaluate text complexity while recommending appropriate reading materials to native adult learners with low literacy skills. The corpus compilation involved a manual selection of text samples across various textual genres and document types, covering a wide range of existing written materials and focusing on the reading needs and reading habits of the target audience—low literacy adults enrolled in vocational education and training centres or adult learning (AL) centres. The collected texts were categorised into the three distinct levels of complexity targeted and defined by the project: very easy, easy, and plain levels. Texts of higher complexity were also included, resulting in the creation of four distinct sub-corpora. The resulting Portuguese dataset consists of 2,186 texts and 942,818 tokens and serves as the foundational source for training and testing the project’s complexity analysis systems. This paper presents a comprehensive overview of the compilation process of the corpus, encompassing its methodological design and the challenges faced. Although some existing Portuguese corpora were used for complexity studies and tool development, these primarily consist of texts classified according to CERF levels and retrieved from didactic materials designed for L2 teaching/learning or texts produced by L2 learners. The corpus presented in this paper introduces a new resource that addresses a significant gap in materials needed to inform and support studies and applications related to text complexity. The resulting dataset provides a novel and important language resource for European Portuguese, with several applications including research on linguistic complexity, development of automatic text complexity and readability assessment systems, and educational purposes.en
dc.description.versionpublishersversion
dc.description.versionpublished
dc.format.extent20
dc.format.extent1054721
dc.identifier.doi10.3991/ijet.v19i08.52023
dc.identifier.issn1863-0383
dc.identifier.otherPURE: 106187097
dc.identifier.otherPURE UUID: 444d8960-d4bc-4d91-80dc-b64abaa4b518
dc.identifier.otherORCID: /0000-0002-4923-7186/work/175650719
dc.identifier.otherORCID: /0000-0002-5617-278X/work/205672282
dc.identifier.otherORCID: /0000-0001-7340-7184/work/208433079
dc.identifier.urihttp://hdl.handle.net/10362/177360
dc.identifier.urlhttps://online-journals.org/index.php/i-jet/article/view/52023
dc.language.isoeng
dc.peerreviewedyes
dc.relationinfo:eu-repo/grantAgreement/FCT/6817 - DCRRNI ID/UIDB%2F03213%2F2020/PT
dc.relationCentre of Linguistics of NOVA University of Lisbon
dc.relationCentre of Linguistics of NOVA University of Lisbon
dc.subjectClassified Corpus
dc.subjectText Complexity
dc.subjectAdult Learning
dc.subjectLow Literacy
dc.titleIntelligent Support for Low Literacy Adultsen
dc.title.subtitleThe European Portuguese iRead4Skills Corpusen
dc.typejournal article
degois.publication.firstPage61
degois.publication.issue8
degois.publication.lastPage81
degois.publication.titleInternational Journal of Emerging Technologies in Learning (iJET)
degois.publication.volume19
dspace.entity.typePublication
oaire.awardNumberUIDB/03213/2020
oaire.awardNumberUIDP/03213/2020
oaire.awardTitleCentre of Linguistics of NOVA University of Lisbon
oaire.awardTitleCentre of Linguistics of NOVA University of Lisbon
oaire.awardURIinfo:eu-repo/grantAgreement/FCT/6817 - DCRRNI ID/UIDB%2F03213%2F2020/PT
oaire.awardURIinfo:eu-repo/grantAgreement/FCT/6817 - DCRRNI ID/UIDP%2F03213%2F2020/PT
oaire.fundingStream6817 - DCRRNI ID
oaire.fundingStream6817 - DCRRNI ID
person.affiliation.nameCentro de Linguística da UNL
person.familyNameCorreia
person.givenNameSusana Mesquita de Deus
person.identifier.ciencia-idE515-2DD9-9EFE
person.identifier.orcidhttps://orcid.org/0000-0001-7340-7184
project.funder.identifierhttp://doi.org/10.13039/501100001871
project.funder.identifierhttp://doi.org/10.13039/501100001871
project.funder.nameFundação para a Ciência e a Tecnologia
project.funder.nameFundação para a Ciência e a Tecnologia
rcaap.rightsopenAccess
relation.isAuthorOfPublication236d1ae1-01c3-4005-a1a5-a4a26513c463
relation.isAuthorOfPublication.latestForDiscovery236d1ae1-01c3-4005-a1a5-a4a26513c463
relation.isProjectOfPublication493fa20c-2660-4a0b-8a41-58f5c58cd372
relation.isProjectOfPublication42811a8d-e313-41cb-989b-9b63edcd66bc
relation.isProjectOfPublication.latestForDiscovery42811a8d-e313-41cb-989b-9b63edcd66bc

Ficheiros

Principais
A mostrar 1 - 1 de 1
A carregar...
Miniatura
Nome:
61_Intelligent_Support_for_Low_Literacy_Adults_The_European_Portuguese_iRead4Skills_Corpus.pdf
Tamanho:
1.01 MB
Formato:
Adobe Portable Document Format