| Nome: | Descrição: | Tamanho: | Formato: | |
|---|---|---|---|---|
| 1.28 MB | Adobe PDF |
Orientador(es)
Resumo(s)
Customer churn has been increasing in insurance, mainly due to technological improvements that allow customers to explore other insurance providers’ offers. Given this, insurance providers need to compete among them, not only to get new customers but, to maintain their own.
This report results from a project developed during an internship in Grupo Ageas Portugal, which has different insurance brands such as Ocidental Seguros. This project's main goal was to model, with a monthly periodicity, customer churn of this latter’s Workers’ Compensation portfolio to improve the company’s competitiveness and, ultimately, profit.
Many of the company’s customer churn happens at their policy renewal time, where the only variable that the company detains control over is the price (premium) variation. Hence, by considering the premium variation and other relevant predictive variables, the goal was to predict the probability of a given customer to churn, allowing the company to optimize the current renewal’s pricing process and maximize this branch’s profit.
Thus, different variables that could influence the company’s customer behavior were collected—one of those was the customer’s location. Given the high dimensionality that such variable would represent and the small dataset available for modeling, clustering analysis is used to create new significant (with fewer dimensions) customer geographical areas.
Different supervised learning algorithms were then evaluated accordingly to their performance in predicting customer churn. The predictive models used were a Gradient Boosting, an Extreme Gradient Boosting, a Logistic Regression, and a Multilayer Perceptron. Given that the number of customers that renew their contracts is much superior to the number of customers who churn, Synthetic Minority Oversampling Technique (SMOTE) was used to create less unbalanced datasets (with synthetic samples) and evaluate the impact on the performance of one of the models.
Lastly, to guarantee a successful integration of the models into the renewal’s pricing process, models were evaluated accordingly to the two business goals. First, by translating the observed evaluation metrics into profit. Secondly, by assuring that the customer’s price elasticity would be captured, assuring a monotonic increasing relationship among the policy’s premium variation and probability of churn.
Descrição
Internship Report presented as the partial requirement for obtaining a Master's degree in Data Science and Advanced Analytics
Palavras-chave
Supervised Learning Classification Customer Churn Prediction Non-Life Insurance Renewal Price Elasticity Clustering Neural Network Logistic Regression Stochastic Gradient Boosting
