Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction

2026-08-27Computation and Language

Computation and LanguageArtificial Intelligence
AI summary

The authors found that clinical language models often rely on text quirks like templates instead of patient information, making them less reliable in new settings. They developed CAST, a method that uses Sparse Autoencoders to identify and remove these misleading text artifacts during training. This approach allows the model to focus on real clinical concepts and provides a way for humans to audit how the model makes decisions. When tested on hospital discharge notes, CAST improved prediction accuracy and offered clear explanations for its predictions.

Clinical language modelsDeployment shiftText artifactsSparse AutoencodersTransformer activationsLLM interpretationICD-10Mortality predictionModel auditing
Authors
Jin Mu, Guanhua Chen
Abstract
Clinical language models can achieve strong in-hospital accuracy yet fail under deployment shifts because they exploit note-specific artifacts (e.g., templates, separators, boilerplate) that do not reflect patient state. We propose CAST (Concept-guided Artifact Suppression Tuning), an SAE-based framework for auditable clinical text classification. CAST uses Sparse Autoencoders to expose sparse, human-auditable features from intermediate Transformer activations, labels SAE latents with an LLM-assisted interpretation pipeline and ICD-10 retrieval constraints, suppresses verified artifact latents via residual subtraction during fine-tuning, and provides post-hoc per-concept attributions for auditing model decisions. On MIMIC-IV discharge-note mortality prediction, CAST improves over its corresponding fine-tuned encoder baselines and remains competitive with strong LLM baselines, while producing a feature-level audit trail of the clinical concepts that support each prediction and the artifact concepts suppressed during training.