| Nome: | Descrição: | Tamanho: | Formato: | |
|---|---|---|---|---|
| 3.92 MB | Adobe PDF |
Autores
Resumo(s)
The increasing volume, speed and heterogeneity of data has led companies to invest in Big Data technology, such as Data Lakes, which are central data repositories capable of ingesting structured and unstructured data in a large scale. However, Data Lakes are still not suitable for typical Business Intelligence use cases and analyses as data is stored without a defined schema, which is why most companies still want to keep their existing Enterprise Data Warehouses (EDW). Regarding architectures that combine a Data Lake and an EDW, there are no defined best practices for data storage, metadata management, and data loading from the Data Lake into the EDW, particularly into those based on Data Vault 2.0. There is also a need to understand the impact that a Delta Lake layer can have in optimizing said data loading. This dissertation aims to fill these gaps in the literature and provide the scientific community and banking industry with an efficient architecture for a Data Lake that sources an EDW, and an EDW model based on Data Vault 2.0.
Descrição
Dissertation presented as the partial requirement for obtaining a Master's degree in Information Management, specialization in Knowledge Management and Business Intelligence
Palavras-chave
Data Lake Data Warehouse Architecture Data Vault Delta Lake Metadata
