Logo do repositório
 
A carregar...
Miniatura
Publicação

Metadata and textual features to advise administrative courts decisions - Predicting judgments of violation processes and identifying similar cases

Utilize este identificador para referenciar este registo.
Nome:Descrição:Tamanho:Formato: 
TCDMAA0161.pdf2.37 MBAdobe PDF Ver/Abrir

Resumo(s)

Many aspects of citizens' lives are affected by the decision-making process carried out by government organizations, such as administrative justice bodies. However, these bodies have limited resources and struggle to keep up with the increasing cases demand. This work project is composed of two studies to assist administrative courts. In the first one, a two-stage cascade classifier predictive model was proposed for advising administrative decisions. The model employs a first-stage prediction from textual features and a second-stage classifier that includes 'proceedings' metadata. The study was conducted using time-based cross-validation, being entirely built on data available before the predicted judgment. It only considers the first document in each proceeding, along with the metadata recorded when the violation is first registered. Consequently, the model endows generality and out-of-sample serviceability. Also, by preserving visibility on the textual features and employing the Shapley Additive Explainability (SHAP), the proposed model provides local explainability. Results show that this is an advantageous procedure when both text and metadata are available. With a weighted F1Score of 0.900, the results outperform the text-only baseline by 1.24% and the metadata-only baseline by 5.63%, with better discriminative properties, evaluated by the Receiver Operating Characteristic (ROC) and Precision-Recall curves. The second study evaluated different combinations of document representations, embeddings, and similarity measures to assess whether these assemblies can identify similar pairs of cases that satisfy the human notion of similarity. A representative dataset of administrative violations was employed, and the models’ outputs were compared to an experts’ gold standard and a baseline model. Adopting the mean Average Recall (mAR) as the primary evaluation metric, the Word2Vec models outperformed the baseline, and together with the TF-IDF models, obtained the best Recall scores overall. Stemming showed to be an essential pre-processing technique to influence the results. Findings also suggest that the choice of the model (Word2Vec) prevails over the setup, i.e., the corpus, the number of dimensions, and the implementation algorithm. Additionally, the performance obtained using sentences as text representation, or BM25 and Doc2Vec as models, showed to be noticeably low. The encouraging results from the studies confirm that both AI-based legal assistance techniques can be instrumental in helping administrative justice bodies improve their decision-making process, both in terms of speed and consistency of decisions.

Descrição

Project Work presented as the partial requirement for obtaining a Master's degree in Data Science and Advanced Analytics, specialization in Business Analytics

Palavras-chave

Administrative decision prediction document similarity cascade generalization legal assistance machine learning natural language processing SDG 16 - Promote peaceful and inclusive societies for sustainable development, provide access to justice for all and build effective, accountable and inclusive institutions at all levels

Contexto Educativo

Citação

Projetos de investigação

Unidades organizacionais

Fascículo