Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows

2026-08-25Computation and Language

Computation and LanguageArtificial Intelligence
AI summary

The authors studied large language models (LLMs) used to analyze financial documents and help with investment decisions. They found that even when LLMs can accurately retrieve information, that information doesn’t always affect their investment judgments, especially as the amount of extra unrelated context grows. This problem happens across different models and tasks and can be reduced but not fully fixed by better models. The authors also showed that how the information is processed and structured (workflow) is important for making sure retrieved facts actually influence decisions. Simply measuring retrieval isn’t enough to judge an AI analyst’s true understanding.

large language modelsfinancial disclosuresinvestment decisionsretrieval-integration gaplong-context analysis10-K filingschunk-and-summarizeworkflow architecturecausal memory interventionsAI evaluation
Authors
Miao Liu, Zhizhe Liu
Abstract
Large language models (LLMs) are increasingly deployed as AI analysts to process financial disclosures and support AI-assisted investment decisions. Yet such systems are usually evaluated by what they can retrieve, not whether retrieved information affects their judgments. We identify a retrieval-integration gap in long-context financial analysis. Holding focal-firm information fixed and varying only unrelated context from 2,000 to 128,000 tokens, we find that a risk disclosure's influence on investment judgments falls to the experimental noise floor even as direct retrieval remains accurate. The pattern replicates across model families and judgment tasks and in experiments removing real disclosures from actual 10-K filings. More capable models postpone but do not eliminate the gap. Causal memory interventions show that compressed summaries and source-text lookup jointly transmit disclosures into judgments. Workflow architecture determines whether this transmission succeeds: chunk-and-summarize pipelines evict relevant information, whereas a targeted, structured restatement adjacent to the decision restores its influence. AI analyst performance is therefore jointly determined by model capability and workflow architecture. Retrieval-based evaluations can certify systems whose investment judgments ignore information they demonstrably retrieved.