My research is situated in Human-Computer Interaction (HCI), a fundamentally interdisciplinary field that draws on computer science, design, ergonomics, cognitive science, and the humanities and social sciences. As such, my work adopts an interdisciplinary approach.
Main thematic areas
Human-AI interaction
The first area concerns the study of Human-AI Interaction. Within this framework, I analyse, under controlled laboratory conditions, how certain characteristics of AI (e.g. its ability to learn from examples) influence user behaviour (perception and action) during interaction. This area follows a classical HCI approach, drawing on computer science, design, and cognitive science.
Examples of works include:
The second area concerns the impact of artificial intelligence (AI) on the creative and cultural sectors, application domains I have been engaged with since my PhD at IRCAM. This area is grounded in field methodology, combining observation and interviews, and contributes to the social sciences and humanities dimension of HCI.
Examples of works include:
The third area addresses the explainability of AI models. It aims to design and evaluate explainability tools for AI that are aligned with human judgment. This area sits at the intersection of HCI and AI.
Examples of works include:
Current and Past Supervision. I had the chance of supervising PhD students and postdoctoral researchers across a diverse range of fields, including human-AI interaction, embodied interaction, interactive machine learning, music technology, and, more recently, empirical philosophy of AI.
Human testers—often end-users themselves—can judge system errors in context and reveal failures that automated methods miss. Inspired by software testing practices, we conducted an exploratory study in which 15 participants tested a satellite image classifier using an interactive tool that allowed them to collect data and create test cases. While participants shared a common understanding of the model’s behavior and generally adopted a failure-driven approach, we observed significant variability in their testing behaviors, including the number of test cases created, the timing of seeking feedback, the distribution of effort across classes, and the types of failures identified. Although links between specific strategies and outcomes remain unclear, our findings provide a first step toward understanding human testing of ML models and inform future research on human-driven AI auditing.
Sensemaking in User-Driven Algorithm Auditing: A Case Study on Gender Bias in an Image Captioning Model
Behnoosh
Mohammadzadeh, Jules
Françoise, Michèle
Gouiffès, and
1 more author
In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (Best Paper 🏆), Apr 2026
Non-experts increasingly engage in user-driven algorithm auditing, interacting directly with AI systems to probe, document, and reflect on biased behavior. Yet, auditing remains challenging due to model opacity and limited support for navigating and interpreting outputs. This paper explores the design and evaluation of interfaces grounded in the sensemaking framework to support non-experts in auditing gender bias in image captioning. In a between-subjects study, 60 participants audited an image captioning model using one of three interface conditions: a Baseline interface, a Masking Tool for image manipulation, or a Filtering Tool for organizing captions. Our findings show that interface design shaped what participants noticed, how they interpreted model behavior, and supported their hypotheses. The Image Masking Tool enabled fine-grained testing of visual cues and context, while the Text Filtering Tool revealed broader asymmetries in gendered language. We argue that incorporating sensemaking into auditing practices can advance accountability and transparency in machine learning systems.
Artists on a Decade of AI Evolution: An Interview Study of Affordances, Culture, and Artistic Practice with Machine Learning
Téo
Sanchez, Mariya
Dzhimova, Stacy
Hsueh, and
3 more authors
In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, Apr 2026
In the mid-2010s, media artists began developing practices using machine learning (ML) as an artistic medium. Since 2022, the rise of large generative models, the mainstreaming of AI as consumer products, and intensifying ethical disputes have reconfigured the conditions of their artistic practice. This paper aims to understand how artists working with ML over the past decade respond to these shifts, shedding light on how practices, tools, and culture co-evolve. We address this question through thematic analysis of semi-structured interviews with 30 artists active before 2020. Our findings show how artists experience narrowing aesthetics and reduced malleability of post-2020 ML systems, have diverging views on where to locate moral responsibility with large AI models, and face shifting cultural reception that challenges the legibility of their work. We map how artists envision their practice going forward and discuss those orientations with respect to HCI conversations on design and creativity.
Benchmarking XAI Explanations with Human-Aligned Evaluations
Rémi
Kazmierczak, Steve
Azzolin, Eloïse
Berthier, and
9 more authors
Proceedings of the AAAI Conference on Artificial Intelligence, Mar 2026
We introduce PASTA (Perceptual Assessment System for explanaTion of Artificial Intelligence), a novel human-centric framework for evaluating eXplainable AI (XAI) techniques in computer vision. Our first contribution is the creation of the PASTA-dataset, the first large-scale benchmark that spans a diverse set of models and both saliency-based and concept-based explanation methods. This dataset enables robust, comparative analysis of XAI techniques based on human judgment. Our second contribution is an automated, data-driven benchmark that predicts human preferences using the PASTA-dataset. This scoring called PASTA-score method offers scalable, reliable, and consistent evaluation aligned with human perception. Additionally, our benchmark allows for comparisons between explanations across different modalities, an aspect previously unaddressed. We then propose to apply our scoring method to probe the interpretability of existing models and to build more human interpretable XAI methods.
Disentangling Hallucinations: Orthogonal Semantic Projection for Robust Interpretability
Emirhan
Bilgiç, Baptiste
Caramiaux, Zhi
Yan, and
1 more author
In European Conference on Computer Vision (ECCV), Jun 2026
As Vision-Language Models are increasingly deployed in safety-critical applications, the trustworthiness of their explanations becomes crucial. Explainable AI (XAI) methods for Vision-Language Models often suffer from semantic hallucination, where attribution maps highlight prominent image regions even when prompted with incorrect text descriptions (e.g., highlighting a dog when prompted “cat”). Although this problem is widespread, a formal mathematical analysis of XAI methods and CLIP embeddings is largely missing in the literature. We demonstrate that this phenomenon is not specific to a single architecture but is a fundamental consequence of Linear Semantic Leakage in high-dimensional embedding spaces. We propose a unified theoretical framework, Linear Semantic Attribution (LSA), which generalizes across discriminative methods. We introduce OSP, a geometric intervention that utilizes the residual property of OMP to disentangle unique semantic signals from shared concepts. We prove theoretically and demonstrate empirically that OSP minimizes hallucination by orthogonalizing the query vector against distractor concepts, rendering the attribution model blind to shared features while preserving fidelity for correct prompts. Our code is available at: https://github.com/emirhanbilgic/Orthogonal-Semantic-Projection
IR Lens: A Tool for Interpreting Cross-Encoder Models
Mihai
Branga-Peicu, Mathias
Vast, Basile
Van Cooten, and
4 more authors
In Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval, Jun 2026
Transformer-based ranking models, such as MonoBERT, are central to Information Retrieval; yet their inner workings remain largely opaque. This hinders not only our understanding of the systems implementing them, but also our ability to improve them. To alleviate this limitation, we introduce IR Lens, a new interpretability tool tailored to cross-encoders based on two key components: 1) Neuron Integrated Gradients to expose the contributions of model parts at multiple levels, and 2) targeted ablations to support hypothesis tracking. With its interactive graphical interface, IR Lens enables IR practitioners to explore, analyze, and manipulate neuron-level mechanisms in cross-encoders, facilitating a deeper understanding of neural ranking models. By extending the reach of existing interpretability methods, we believe IR Lens has the potential to support the improvement of cross-encoders.
2025
Generative AI in Documentary Photography: Exploring Opportunities and Challenges for Visual Storytelling
Lenny
Martinez, Baptiste
Caramiaux, and Sarah
Fdili Alaoui
In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, Apr 2025
Generative AI is increasingly used to create images from text, but its role in documentary photography remains under-explored. This paper investigates how generative AI can be integrated into documentary practice while maintaining ethical standards. Through interviews with six documentary photographers, we explored their views on AI’s potential to support community-driven storytelling. While AI presents opportunities for creative expression and community involvement, concerns about trust, authenticity, and decontextualization of images persist. Photographers expressed doubts about AI’s ability to accurately represent lived experiences, fearing it could compromise narrative integrity. Our findings suggest that AI tools should be designed to enhance collaboration and transparency in storytelling, complementing rather than replacing traditional documentary methods. This study contributes to the ongoing discourse on AI in photography, advocating for the development of tools that preserve the ethical foundations of documentary storytelling while empowering communities.
Generative AI and Creative Work: Narratives, Values, and Impacts
Baptiste
Caramiaux, Kate
Crawford, Q. Vera
Liao, and
2 more authors
Generative AI has gained a significant foothold in the creative and artistic sectors. In this context, the concept of creative work is influenced by discourses originating from technological stakeholders and mainstream media. The framing of narratives surrounding creativity and artistic production not only reflects a particular vision of culture but also actively contributes to shaping it. In this article, we review online media outlets and analyze the dominant narratives around AI’s impact on creative work that they convey. We found that the discourse promotes creativity freed from its material realisation through human labor. The separation of the idea from its material conditions is achieved by automation, which is the driving force behind productive efficiency assessed as the reduction of time taken to produce. And the withdrawal of the skills typically required in the execution of the creative process is seen as a means for democratising creativity. This discourse tends to correspond to the dominant techno-positivist vision and to assert power over the creative economy and culture.
2024
Comparing Teaching Strategies of a Machine Learning-based Prosthetic Arm
Vaynee
Sungeelee, Nathanaël
Jarrassé, Téo
Sanchez, and
1 more author
In Proceedings of the 29th International Conference on Intelligent User Interfaces, Apr 2024
Pattern-recognition-based arm prostheses rely on recognizing muscle activation to trigger movements. The effectiveness of this approach depends not only on the performance of the machine learner but also on the user’s understanding of its recognition capabilities, allowing them to adapt and work around recognition failures. We investigate how different model training strategies to select gesture classes and record respective muscle contractions impact model accuracy and user comprehension. We report on a lab experiment where participants performed hand gestures to train a classifier under three conditions: (1) the system cues gesture classes randomly (control), (2) the user selects gesture classes (teacher-led), (3) the system queries gesture classes based on their separability (learner-led). After training, we compare the models’ accuracy and test participants’ predictive understanding of the prosthesis’ behavior. We found that teacher-led and learner-led strategies yield faster and greater performance increases, respectively. Combining two evaluation methods, we found that participants developed a more accurate mental model when the system queried the least separable gesture class (learner-led). Our results conclude that, in the context of machine learning-based myoelectric prosthesis control, guiding the user to focus on class separability during training can improve recognition performances and support users’ mental models about the system’s behavior. We discuss our results in light of several research fields : myoelectric prosthesis control, motor learning, human-robot interaction, and interactive machine teaching.
Studying Collaborative Interactive Machine Teaching in Image Classification
Behnoosh
Mohammadzadeh, Jules
Françoise, Michèle
Gouiffès, and
1 more author
In Proceedings of the 29th International Conference on Intelligent User Interfaces, Apr 2024
While human-centered approaches to machine learning explore various human roles within the interaction loop, the notion of Interactive Machine Teaching (IMT) emerged with a focus on leveraging the teaching skills of humans as a teacher to build machine learning systems. However, most systems and studies are devoted to single users. In this article, we study collaborative interactive machine teaching in the context of image classification to analyze how people can structure the teaching process collectively and to understand their experience. Our contributions are threefold. First, we developed a web application called TeachTOK that enables groups of users to curate data and train a model together incrementally. Second, we conducted a study in which ten participants were divided into three teams that competed to build an image classifier in nine days. Qualitative results of participants’ discussions in focus groups reveal the emergence of collaboration patterns in the machine teaching task, how collaboration helps revise teaching strategies and participants’ reflections on their interaction with the TeachTOK application. From these findings we provide implications for the design of more interactive, collaborative and participatory machine learning-based systems.
2022
Deep Learning Uncertainty in Machine Teaching
Téo
Sanchez, Baptiste
Caramiaux, Pierre
Thiel, and
1 more author
In Proceedings of the 27th International Conference on Intelligent User Interfaces (Best Paper 🏆), Mar 2022
Machine Learning models can output confident but incorrect predictions. To address this problem, ML researchers use various techniques to reliably estimate ML uncertainty, usually performed on controlled benchmarks once the model has been trained. We explore how the two types of uncertainty—aleatoric and epistemic—can help non-expert users understand the strengths and weaknesses of a classifier in an interactive setting. We are interested in users’ perception of the difference between aleatoric and epistemic uncertainty and their use to teach and understand the classifier. We conducted an experiment where non-experts train a classifier to recognize card images, and are tested on their ability to predict classifier outcomes. Participants who used either larger or more varied training sets significantly improved their understanding of uncertainty, both epistemic or aleatoric. However, participants who relied on the uncertainty measure to guide their choice of training data did not significantly improve classifier training, nor were they better able to guess the classifier outcome. We identified three specific situations where participants successfully identified the difference between aleatoric and epistemic uncertainty: placing a card in the exact same position as a training card; placing different cards next to each other; and placing a non-card, such as their hand, next to or on top of a card. We discuss our methodology for estimating uncertainty for Interactive Machine Learning systems and question the need for two-level uncertainty in Machine Teaching.
"Explorers of Unknown Planets": Practices and Politics of Artificial Intelligence in Visual Arts
Baptiste
Caramiaux and Sarah
Fdili Alaoui
Proceedings of the ACM on Human-Computer Interaction, Nov 2022
Alongside recent advances in artificial intelligence (AI), a new art practice has emerged in recent years that borrows and transforms these advances in the production of artworks. The actors of this emergent practice are coming from contemporary art, media and digital arts. These artists have developed an original practice of AI within their creative field. In this article, we propose a qualitative study to explore the nature of this practice. We interviewed five internationally renowned artists about how AI is integrated into their work. Through a thematic analysis of the interviews, we first find that their practice relies on crafting algorithms and data as materials. We uncover how they explicitly use this material unpredictability rather than avoid it. Secondly, we highlight the politics of their practice that consist of resisting the culture of AI research, as well as its inherent power dynamics. We also highlight how their relationship with the technology is imbued with ethics and how they rethink their role with respect to the technology. In this paper, we aim to provide the CSCW community with a way to expand the framework in which AI can be understood not only as a tool but also as cultural and political design material.
2021
How do people train a machine? Strategies and (Mis) Understandings
Téo
Sanchez, Baptiste
Caramiaux, Jules
Françoise, and
2 more authors
Proceedings of the ACM on Human-Computer Interaction, Nov 2021