Logo do repositório
 
Publicação

Towards Explaining Actions of Learning Agents

dc.contributor.authorRodrigues, Bruno
dc.contributor.authorKnorr, Matthias
dc.contributor.authorKrippahl, Ludwig
dc.contributor.authorGonçalves, Ricardo
dc.contributor.institutionNOVALincs
dc.date.accessioned2026-03-16T12:19:01Z
dc.date.available2026-03-16T12:19:01Z
dc.date.issued2023-06-01
dc.description.abstractAgents increasingly use Deep Neural Networks to process sensor information and make decisions. While these models have been shown to provide excellent results, they come with the disadvantage that they behave like black boxes, mapping inputs to outputs in a way that is hard for humans to understand. This is a serious disadvantage because it makes it harder to predict how agents will act in unexpected situations, which is especially dangerous when agents have to interact physically with humans, such as self-driving vehicles or industrial robots, but also creates risks for agents such as chat bots and other virtual agents since their actions may result in legal liabilities or reputation damage. Being able to explain decisions taken by these neural networks that guide the agents is important for preventing incorrect behavior and for building trust and providing legal justifications whenever necessary. This applies not only to interactions with humans, but also to multiagent systems. In this paper, we build on a recent framework on Explainable AI that uses small neural networks to map activations from a trained deep neural network to relevant concepts in a logical formalization of the domain, which in turn can be used to provide explanations for the outputs of the original network. Since this framework is applied to the deep neural network at inference time, after training, it can be applied to neural networks used in agents regardless of whether these were trained using supervised or reinforcement learning. We show that a potential bottleneck of the approach, the creation of such mapping networks, can be solved by employing automated neural architecture search. This paves the way towards applying this approach to more advanced use cases of explaining decisions of agents based on deep neural networks, regardless of how these networks were trained.en
dc.description.versionpublishersversion
dc.description.versionpublished
dc.format.extent9
dc.format.extent1006576
dc.identifier.otherPURE: 157622497
dc.identifier.otherPURE UUID: c97f8973-cf35-49a8-a989-5d60995eadab
dc.identifier.otherBibtex: 9770be021479011cd758e7fb9ad8206b
dc.identifier.urihttp://hdl.handle.net/10362/201451
dc.identifier.urlhttps://alaworkshop2023.github.io/papers/ALA2023_paper_14.pdf
dc.language.isoeng
dc.peerreviewedyes
dc.subjectExplanations
dc.subjectNeural Architecture
dc.subjectReinforcement Learning
dc.titleTowards Explaining Actions of Learning Agentsen
dc.typeconference object
degois.publication.titleProceedings of the Adaptive and Learning Agents Workshop (ALA 2023)
dspace.entity.typePublication
rcaap.rightsopenAccess

Ficheiros

Principais
A mostrar 1 - 1 de 1
A carregar...
Miniatura
Nome:
Rodrigues_et_al._2023._Towards_Explaining_Actions_of_Learning_Agents..pdf
Tamanho:
982.98 KB
Formato:
Adobe Portable Document Format