On the Structural Limits of Machine Learning Decision Systems: An Information-Theoretic, Interaction-Based, and Stochastic-Dynamical Perspective
2026-08-13 • Machine Learning
Machine Learning
AI summaryⓘ
The authors explain that machine learning methods are limited not just by algorithms but mainly by the underlying data and how it is generated. They use concepts from information theory, like Fano bounds and the Cramér-Rao inequality, to show the minimal error we can expect in classification and estimation tasks. They also point out that assumptions about the data, such as being independent or stable over time, impact the reliability of these methods. Finally, they describe decision systems as dynamic processes influenced by their past states, stressing the need for good data models to improve predictions within these fundamental limits.
Machine LearningInformation TheoryFano BoundCramér-Rao InequalityData-Generating ProcessParametric EstimationMarkov Random FieldsErgodicityLLM-integrated AgentsStochastic Processes
Authors
Nestor R. Barraza, Gabriel Pena
Abstract
Machine learning procedures are commonly evaluated in terms of predictive accuracy and computational efficiency. However, their achievable performance is fundamentally constrained by structural properties of the underlying data-generating process, which are formalized in terms of informational bounds. In this work we examine intrinsic limits of data-driven decision systems from an information-theoretic and interaction-based perspective. We analyze minimal achievable error in classification through Fano-type bounds and precision limits in parametric estimation via the Cramér-Rao inequality, emphasizing that such limits depend on the underlying model rather than on algorithmic sophistication alone. We further discuss how implicit assumptions, such as independence, ergodicity, and distributional stability, affect the validity of inferential procedures. Building on interaction-based modeling principles, we review typical frameworks such as Markov Random Fields and potential based representations for encoding dependence mechanisms. We also describe decision systems, including LLM-integrated agent architectures, as feedback-driven stochastic processes where state-dependent dynamics may induce emergent macroscopic behavior. This perspective highlights the importance of having adequate models for the data as a prerequi- site for expanding predictive capability, and situates algorithmic learning within the informational limits imposed by the models.