| Nome: | Descrição: | Tamanho: | Formato: | |
|---|---|---|---|---|
| 2 MB | Adobe PDF |
Autores
Orientador(es)
Resumo(s)
Artificial neural networks are typically regarded as black boxes, given how difficult it is for humans to interpret how these models reach their results. One way in which humans attempt to interpret complex systems is by imagining how they behave in hypothetical scenarios. In this work, we propose a method that allows one to modify what an artificial neural network is perceiving regarding specific human-defined concepts of interest, allowing one to test how they behave under such hypothetical scenarios. Through empirical evaluation, we test the proposed method on different models and datasets, assessing its qualities.
Descrição
Publisher Copyright: © 2025 Copyright for this paper by its authors.
Palavras-chave
Counterfactual Explainability Interpretability Neural Networks General Computer Science
