| Nome: | Descrição: | Tamanho: | Formato: | |
|---|---|---|---|---|
| 2.29 MB | Adobe PDF |
Autores
Orientador(es)
Resumo(s)
The use of data mining as a way of solving problems from the widest range of areas with the main
purpose of gaining competitive advantage is rising, specially in retail, an extremely competitive
sector that requires an even bigger advantage. Additionally, food loss, beyond representing a huge
waste of resources, can also be considered a major issue to the retail sector due to the financial
losses originated from it. Thus, I proposed to help Pingo Doce, a well-known Portuguese retail
company, to solve their food loss issue which, despite being the major cause of a huge drop in the
company’s profits, has never been solved till this day.
Therefore, this project focuses on the development of a classification algorithm that will allow to
predict future significant losses in several fruits sold in certain Pingo Doce stores. To do so I applied a
Decision Tree algorithm that, due to its representation in the form of if-then rules, will help to
identify the main features that lead to a higher number of losses, namely the period of the year and
the category to which each fruit belongs, among others.
The dataset provided by the company contains variables that measure the quantity and value of
sales, stocks, identified and unidentified losses, over a one-year period, and regarding 81 different
fruits and 20 stores from all over the country. Additionally, I created new variables such as the
criminality rate of the municipality and the climate class of each store, as well as the seasons and the
day of the week in which each observation occurred. All these variables allowed me to create four
different datasets that originated four different Classification Trees.
The results show that, using a dataset with no information regarding stocks and sales, containing
only variables that describe the characteristics of the stores, products and periods of time, as well as
the value of product sold per unit of measurement, i.e. the price per unit of measurement of each
fruit, it is possible to create a Decision Tree that reaches an accuracy of 74% and correctly predicts
82% of the observations that represent significant losses.
The algorithm obtained allowed to identify the variables that are more prone to originate significant
losses, namely: the day of the week, the fruit’s category, the season of the year, the position of that
week in the respective month and the price at which the product is being sold.
Descrição
Project Work presented as the partial requirement for obtaining a Master's degree in Information Management, specialization in Knowledge Management and Business Intelligence
Palavras-chave
Machine Learning Retail Loss Classification Algorithm Classification Tree Decision Tree
