Understanding How Humans Inject Knowledge into Machine Learning Workflows through Visual Analytics

2026-07-01Human-Computer Interaction

Human-Computer InteractionMachine Learning
AI summary

The authors looked at over 200 research papers that study how people use visual tools to help improve machine learning (ML) processes. They focused on ways humans add their knowledge, like labeling data or tuning settings, through interactive visualizations. By organizing and analyzing these studies, they identified patterns in how visualization supports human involvement in ML. They also proposed models to explain why these visual tools are helpful for building and optimizing ML models.

Visual AnalyticsMachine Learning WorkflowsInteractive VisualizationData LabelingHyper-parameter TuningFeature EngineeringModel ArchitectureInformation-theoretic Cost-benefit AnalysisHuman-in-the-loop
Authors
Yiwen Xing, Philip Beaucamp, Joyraj Chakraborty, Afrah Farea, Yuanzhe Jin, Saiful Khan, Gennady Andrienko, Natalia Andrienko, Min Chen
Abstract
Visual analytics (VA) plays an increasingly important role in supporting machine learning (ML) workflows. In the field of visualization, such approaches and techniques are referred to as VIS4ML. While ML models are mostly learned automatically, the corresponding ML workflows receive a variety of human inputs, such as data labelling, feature engineering, model architecture designing, hyper-parameter tuning, and so on. In this work, we surveyed over 200 VIS4ML papers to gain an understanding of how humans inject their knowledge into ML workflows through interactive visualization. We collected a corpus of VIS4ML papers from the IEEE VIS conferences in the past decade. We developed a coding scheme to facilitate the literature research from four perspectives: characteristics of ML, visualization, interaction, and actions. The analysis of the coded dataset allows us to observe different pathways that transfer human knowledge to ML workflows via interactive visualization. Building on the analysis, we explain the phenomena of VIS4ML using the conceptual model that views VA as model building and the information-theoretic cost-benefit analysis that reasons VA as for optimizing ML workflows. This work provides unequivocal evidence showing the merits of using VA in ML workflows. The full list of surveyed papers, along with all analysis results and figures, is available at https://vis4ml4hd.github.io/ml-knowledge-inject-va/.