Publicação
Large language models (LLMs) for legal analysis: RAG and beyond for optimizing domain adaptation in Portuguese legal domain
| datacite.subject.fos | Ciências Sociais::Economia e Gestão | |
| dc.contributor.advisor | Han, Qiwei | |
| dc.contributor.author | Barros, Tiago Mendonça Alencar | |
| dc.date.accessioned | 2026-06-18T13:25:09Z | |
| dc.date.available | 2026-06-18T13:25:09Z | |
| dc.date.issued | 2025-01-10 | |
| dc.date.submitted | 2024-12-17 | |
| dc.description.abstract | This study explores RAG systems tailored to the Portuguese legal domain, highlighting challenges in underrepresented languages. Fixed-size chunking strategies, particularly Token Text Splitter, were found to be most effective, while more advanced techniques like Recursive and Semantic splitting showed little benefits. Larger chunk sizes improved retrieval accuracy and answer quality, though the impact of chunk overlap remains inconclusive. Self-reflection techniques show promising results, particularly for weaker LLMs. Techniques such as adding a pre-post translation proved to be an efficient technique for mitigating language bias. | eng |
| dc.identifier.tid | 203927613 | |
| dc.identifier.uri | http://hdl.handle.net/10362/203859 | |
| dc.language.iso | eng | |
| dc.relation | UID/ECO/00124/2013 | |
| dc.rights.uri | http://creativecommons.org/licenses/by/4.0/ | |
| dc.subject | Retrieval-augmented generation | |
| dc.subject | RAG | |
| dc.subject | Large language models | |
| dc.subject | LLM | |
| dc.subject | Artificial intelligence | |
| dc.subject | AI | |
| dc.subject | Hallucination | |
| dc.subject | Question answering | |
| dc.subject | RAG evaluation | |
| dc.subject | Vector store | |
| dc.subject | Chunking | |
| dc.subject | Legal AI | |
| dc.subject | Knowledge graph | |
| dc.subject | GraphRAG | |
| dc.subject | RDF | |
| dc.subject | Graph-based reasoning | |
| dc.subject | Self-assessment | |
| dc.subject | Self-reflection | |
| dc.subject | Multi-agent systems | |
| dc.subject | MAS | |
| dc.subject | Document reranking | |
| dc.subject | Relevance ranking | |
| dc.subject | Legal information retrieval | |
| dc.subject | Portuguese legal retrieval | |
| dc.subject | Machine translation | |
| dc.subject | Natural language processing | |
| dc.subject | LLM bias | |
| dc.subject | Prompt engineering | |
| dc.subject | Hierarchical indexing | |
| dc.subject | Hierarchical retrieving | |
| dc.subject | Chain-of-thought | |
| dc.title | Large language models (LLMs) for legal analysis: RAG and beyond for optimizing domain adaptation in Portuguese legal domain | eng |
| dc.type | master thesis | |
| dspace.entity.type | Publication | |
| thesis.degree.name | A Work Project, presented as part of the requirements for the Award of a Master’s degree in Business Analytics from the Nova School of Business and Economics |
