Logo do repositório
 
A carregar...
Miniatura
Publicação

Treatment Simulation Framework for Applying Causal Machine Learning to Cross-Sectional Data

Utilize este identificador para referenciar este registo.
Nome:Descrição:Tamanho:Formato: 
TDDM4831.pdf986.13 KBAdobe PDF Ver/Abrir

Resumo(s)

This thesis offers a methodological investigation into the use of causal machine learning (CML) methods for extensive, cross-sectional observational datasets. Instead of assessing a real-world intervention, the study employs a simulated treatment framework to demonstrate how contemporary causal estimators can be designed, carried out, and understood in nonexperimental settings. The empirical testbed utilizes the publicly accessible 1990 U.S. Census dataset, from which a theoretical binary treatment is created by categorizing weekly working hours, while a categorized income variable serves as a demonstrative outcome. The suggested pipeline integrates feature selection based on mutual information, estimation and matching of propensity scores, and outcome modeling using Random Forest and Neural Network algorithms. Average Treatment Effects (ATE) and Conditional Average Treatment Effects (CATE) are assessed to analyze the performance of various estimators across subgroups characterized by socio-demographic traits. All outcomes are understood solely within the simulated context and are not regarded as significant assertions regarding labor markets or income trends. The findings indicate that causal machine learning processes can be applied to high-dimensional census data, while also highlighting issues concerning covariate imbalance, discretization, and restricted overlap between treated and control groups. The research provides a clear, replicable illustration of a CML workflow for observational data and specifies methodological advancements for future studies.

Descrição

Dissertation presented as the partial requirement for obtaining a Master's degree in Data Driven Marketing, specialization in Data Science for Marketing

Palavras-chave

Causal Machine Learning Propensity Score Matching Treatment Effect Estimation Heterogeneous Effects Observational Data Simulated Treatment Framework SDG 12 - Responsible production and consumption

Contexto Educativo

Citação

Projetos de investigação

Unidades organizacionais

Fascículo