
SR · Video feature
Researcher, Embedded Intelligence — DFKI
I am a researcher at the Embedded Intelligence group at the German Research Center for Artificial Intelligence (DFKI). My work sits at the intersection of human-computer interaction and artificial intelligence.


SR · Video feature


Klinikum Saarbrücken · Press article
I explore how AI can support learning without taking over the learning process. This includes designing tools that strengthen understanding, preserve learner agency and help educators make informed decisions, as well as examining how teaching and assessment should adapt when capable AI systems are readily available.
Selected publications

As distance learning becomes increasingly important and artificial intelligence tools continue to advance, automated systems for individual learning have attracted significant attention. However, the scarcity of open-source online tools that are capable of providing personalized feedback has restricted the widespread implementation of research-based feedback systems. In this work, we present RATsApp, an open-source automated feedback system that incorporates research-based features such as formative feedback. The system focuses on core STEM competencies such as mathematical competence, representational competence, and data literacy. It also allows lecturers to monitor students' progress. RATsApp can be used at different levels of STEM education or research, as it allows the creation and customization of educational content. We present a specific case of its implementation in higher education, where we report the results of a usability survey with 64 undergraduate students using the technology acceptance model 2. Our findings confirm the applicability of the model, revealing that relevance to the course of study, output quality, and ease of use significantly influence perceived usefulness. We also found a linear relation between perceived usefulness and intention to use, which in turn is a significant predictor of frequency of use. Moreover, the formative feedback feature received positive feedback, indicating its potential as an educational tool. As an open-source platform, RATsApp encourages public contributions to its ongoing development.

The increased presence of large language models (LLMs) in educational settings has ignited debates concerning negative repercussions, including overreliance and inadequate task reflection. Our work advocates moderated usage of such models, designed in a way that supports students and encourages critical thinking. We developed two moderated interaction methods with ChatGPT: hint-based assistance and presenting multiple answer choices. In a study with students (N=40) answering physics questions, we compared the effects of our moderated models against two baseline settings: unmoderated ChatGPT access and internet searches. We analyzed the interaction strategies and found that the moderated versions exhibited less unreflected usage, such as copy and paste, compared to the unmoderated condition. However, neither ChatGPT-supported condition could match the ratio of reflected usage present in internet searches. Our research highlights the potential benefits of moderating language models, showing a research direction toward designing effective AI-supported educational strategies.

Individual teaching is among the most successful ways to impart knowledge. Yet, this method is not always feasible due to large numbers of students per educator. Quantum computing serves as a prime example facing this issue, due to the hype surrounding it. Alleviating high workloads for teachers, often accompanied with individual teaching, is crucial for continuous high quality education. Therefore, leveraging Large Language Models (LLMs) such as GPT-4 to generate educational content can be valuable. We conducted two complementary studies exploring the feasibility of using GPT-4 to automatically generate tips for students. In the first one students (N=46) solved four multiple-choice quantum computing questions with either the help of expert-created or LLM-generated tips. To correct for possible biases towards LLMs, we introduced two additional conditions, making some participants believe that they were given expert-created tips, when they were given LLM-generated tips and vice versa. Our second study (N=23) aimed to directly compare the LLM-generated and expert-created tips, evaluating their quality, correctness and helpfulness, with both experienced educators and students participating. Participants in our second study found that the LLM-generated tips were significantly more helpful and pointed better towards relevant concepts than the expert-created tips, while being more prone to giving away the answer. While participants in the first study performed significantly better in answering the quantum computing questions when given tips labeled as LLM-generated, even if they were created by an expert. This phenomenon could be a placebo effect induced by the participants' biases for LLM-generated content. Ultimately, we find that LLM-generated tips are good enough to be used instead of expert tips in the context of quantum computing basics.

Large language models (LLMs) have recently gained popularity. However, the impact of their general availability through ChatGPT on sensitive areas of everyday life, such as education, remains unclear. Nevertheless, the societal impact on established educational methods is already being experienced by both students and educators. Our work focuses on higher physics education and examines problem solving strategies. In a study, students with a background in physics were assigned to solve physics exercises, with one group having access to an internet search engine (N=12) and the other group being allowed to use ChatGPT (N=27). We evaluated their performance, strategies, and interaction with the provided tools. Our results showed that nearly half of the solutions provided with the support of ChatGPT were mistakenly assumed to be correct by the students, indicating that they overly trusted ChatGPT even in their field of expertise. Likewise, in 42% of cases, students used copy and paste to query ChatGPT, an approach only used in 4% of search engine queries, highlighting the stark differences in interaction behavior between the groups and indicating limited reflection when using ChatGPT. In our work, we demonstrated a need to guide students on how to interact with LLMs and create awareness of potential shortcomings for users.

Allowing users of interactive systems to reflect on their task proficiency is often incidental. This is unfortunate, as communicating meaningful task-related proficiency feedback could improve users' awareness of their abilities and their willingness to improve. To highlight the feasibility of this concept, we evaluated how different methods of readability feedback impacted users during a text production task. In general, our results showed that having access to readability feedback allowed participants to reflect on their task solving approach, facilitating the users' understanding of their proficiency. Revision-based methods are less distracting for the user than continuous feedback methods, while still offering high efficacy. Further, feedback should be paired with a subtle form of gamification elements. We envision this reflection-oriented design to user proficiency to be applicable to a variety of interactive systems, allowing for an improved and engaging user experience.
I explore how AI can support learning without taking over the learning process. This includes designing tools that strengthen understanding, preserve learner agency and help educators make informed decisions, as well as examining how teaching and assessment should adapt when capable AI systems are readily available.

Publication · 2024
Lars Krupp, Steffen Steinert, Maximilian Kiefer-Emmanouilidis, Karina E. Avila, Paul Lukowicz, Jochen Kuhn, Stefan Küchemann, Jakob Karolus
ACM International Conference on Mobile and Ubiquitous Multimedia (MUM)
Summary
Studies moderated uses of ChatGPT designed to support learning and critical thinking. Hint-based assistance and multiple-answer formats were compared with unmoderated ChatGPT and internet search in a physics study. Moderated interactions reduced unreflected behaviors such as copy and paste, though they did not reach the reflective usage seen with search engines.
Best Paper Award
The increased presence of large language models (LLMs) in educational settings has ignited debates concerning negative repercussions, including overreliance and inadequate task reflection. Our work advocates moderated usage of such models, designed in a way that supports students and encourages critical thinking. We developed two moderated interaction methods with ChatGPT: hint-based assistance and presenting multiple answer choices. In a study with students (N=40) answering physics questions, we compared the effects of our moderated models against two baseline settings: unmoderated ChatGPT access and internet searches. We analyzed the interaction strategies and found that the moderated versions exhibited less unreflected usage, such as copy and paste, compared to the unmoderated condition. However, neither ChatGPT-supported condition could match the ratio of reflected usage present in internet searches. Our research highlights the potential benefits of moderating language models, showing a research direction toward designing effective AI-supported educational strategies.
Selected publications
5 total