Multi-Agent AI System for Radiology Report Structuring and Quality Assurance with Independent Radiologist Evaluation

2026-08-18Computation and Language

Computation and Language
AI summary

The authors developed a computer system using artificial intelligence to organize and check the quality of radiology reports from CT scans. Their system divided reports into clear anatomical sections and looked for errors or missing important information. When tested on hundreds of reports, the system successfully structured all of them and flagged potential issues in about 14%. Two expert radiologists reviewed a smaller set and found that the system’s work was mostly accurate and didn’t miss or add important details. The authors suggest this kind of AI tool could help make radiology reports more consistent and reliable.

radiology reportCT scanartificial intelligencereport structuringquality assurancelarge language modelsfindings sectionimpression sectionregex rulesanatomical sections
Authors
Iryna Hartsock, Cesar Lam, Christopher Otteni, Aliya Qayyum, Robert Gatenby, Cyrillo Araujo, Ghulam Rasool
Abstract
Purpose: To develop and evaluate a locally deployed multi-agent AI system for radiology report structuring and quality assurance. Materials and Methods: This retrospective study included 638 radiology reports from CT examinations of the chest, abdomen, and pelvis dictated by 15 board-certified radiologists in 2023 and 2024. A multi-agent AI pipeline was developed to perform report structuring and quality assurance (QA). The system structured the report into standardized anatomical sections at the sentence level using regex rules and local large language models. It also detected mismatches between the Findings and Impression sections, or within sections; gender-anatomy conflicts; and undocumented communication of critical findings. Two board-certified radiologists independently evaluated a 45-report subset. Results: The multi-agent system structured the Findings sections of all reports (22,270 sentences) into a predefined anatomical format while retaining the original report content. The system flagged 90 (14.1%) reports, most commonly for section mismatches (80 reports, 12.5%). In the radiologist evaluation, both reviewers agreed that 31 (69%) were correctly restructured, 2 reports (4%) were incorrectly restructured, and disagreed on the remaining 12 reports (27%). Both reviewers agreed that no clinically important information was omitted and no fabricated content was introduced. Overall QA performance was rated as "excellent" or "good" in 84% of the evaluated reports, with the remaining reports rated as "fair". Conclusion: A locally deployed multi-agent AI system combined radiology report structuring and quality assurance within a single workflow. The system demonstrated favorable performance in radiologist evaluation. Such systems may support standardization of reporting and quality assurance in radiology practice.