Developers' Experience with Generative AI Beyond Productivity Assessment -- Insights from an Empirical Mixed-Methods Field Study

2026-07-02Software Engineering

Software Engineering
AI summary

The authors studied how professional developers use AI coding assistants like GitHub Copilot during their normal work. They found that developers like using these tools for boring or repetitive tasks and believe they get work done faster. Different ways of interacting with AI, like chat or in-code suggestions, work better for different tasks, but mixing these methods in one task can be less helpful. The study also showed that using AI can sometimes make thinking harder, especially when coding a lot, but better AI quality helps productivity. Overall, the study helped developers understand and use AI tools more intentionally.

Generative AIAI coding assistantGitHub Copilottask efficiencycognitive loadin-code suggestionschat-based promptingmixed-methods studydeveloper productivityAI interaction
Authors
Charlotte Brandebusemeyer, Kerim Zunic, Thomas Zimmermann, Tobias Schimmer, Bert Arnrich
Abstract
With the growing adoption of AI-powered coding assistants, organizations and developers are increasingly seeking to optimize their interaction with these tools. Prior research has largely focused on output quality and productivity gains, with limited attention paid to developers' well-being and interaction experiences. This paper presents a developer-centered empirical mixed-methods study to investigate how professional developers engage with Generative AI (GenAI) in their natural work environment. Controlled data collection sessions are combined with natural work periods. Results show that developers are generally satisfied with GenAI, particularly for monotonous, repetitive, and structured tasks, and report perceived efficiency and productivity gains. Copilot interaction type preferences differ by task type and complexity: While both in-code suggestions and chat-based prompting independently improve task efficiency and reduce perceived workload, combining these interaction types within a single task diminishes benefits. We propose a rule-of-thumb for selecting an interaction type based on task characteristics. During development-heavy tasks, results indicate that perceived cognitive load arises from AI interaction, while perceived productivity depends on AI output quality. Participation in this study positively influenced developers' awareness and intentional use of GenAI tools. These findings demonstrate the value of real-world, mixed-methods study designs to understand GenAI tools and developers' experiences with them.