Logo do repositório
 
Publicação

Time and Space Efficient Data Structures for Supporting Machine Translation Tasks

datacite.subject.fosEngenharia e Tecnologia::Engenharia Eletrotécnica, Eletrónica e Informáticapt_PT
dc.contributor.advisorLopes, José
dc.contributor.advisorRusso, Luís
dc.contributor.authorCosta, Jorge André Nogueira da
dc.date.accessioned2018-01-24T10:35:37Z
dc.date.available2018-01-24T10:35:37Z
dc.date.issued2017-12
dc.date.submitted2017
dc.description.abstractThe amount of digital natural language text collections available nowadays is huge and it has been growing at an exponential rate. All this information can be easily accessed by individuals of several nationalities and cultures. This leads to the development of new and innovative techniques and tools, for processing and indexing these texts, in fields of research such as Machine Translation, Natural Language Processing or Cross-Language Information Retrieval. Over the years, a lot of important work has been developed, using efficient data structures, such as suffix arrays, for fast pattern matching and to determine statistics. However, these data structures require a considerable amount of space, around four times the text size, which is a problem considering the amount of bilingual texts available in so many languages. This thesis proposal introduces a two-layer bilingual framework based on compact data structures, for indexing parallel texts, translation memories and bilingual lexica, and their alignments, in pairs of two different languages. Besides a word-based suffix array implementation, this thesis proposal presents a solution based on two byte-codes wavelet trees, one for each text, and bitmaps to represent the alignment. Additionally, it introduces a skip-based bilingual search procedure that speeds up the search time response of the framework, for operations over pairs of word,multi-word or discontiguous phrases. For indexing and querying over aligned parallel corpora, the bilingual framework presents a space consumption around 50% of the alignment-annotated corpora size, against the 160% of the non compressed approach. In terms of search time response, the compressed approach is slower than the one based on suffix arrays as expected. The skip-based bilingual search procedure improves the time response from the original bilingual search algorithm from 1.6x to 2.3x in average. With such space requirements, the framework is able to represent huge amounts of data in main memory, avoiding the considerably slower disk accesses, and to support tasks such as translation, text alignment, word-sense disambiguation or context analysis.pt_PT
dc.identifier.tid101416210
dc.identifier.urihttp://hdl.handle.net/10362/28929
dc.language.isoengpt_PT
dc.subjectBilingual textspt_PT
dc.subjectByte-codesWavelet Treept_PT
dc.subjectSuffix Arraypt_PT
dc.subjectMachine Translationpt_PT
dc.subjectBilingual Frameworkpt_PT
dc.subjectbilingual searchpt_PT
dc.titleTime and Space Efficient Data Structures for Supporting Machine Translation Taskspt_PT
dc.typedoctoral thesis
dspace.entity.typePublication
rcaap.rightsopenAccesspt_PT
rcaap.typedoctoralThesispt_PT
thesis.degree.nameDoutor em Informáticapt_PT

Ficheiros

Principais
A mostrar 1 - 1 de 1
A carregar...
Miniatura
Nome:
Costa_2017.pdf
Tamanho:
2.62 MB
Formato:
Adobe Portable Document Format
Licença
A mostrar 1 - 1 de 1
Miniatura indisponível
Nome:
license.txt
Tamanho:
348 B
Formato:
Item-specific license agreed upon to submission
Descrição: