Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions
2026-07-03 • Cryptography and Security
Cryptography and SecurityArtificial IntelligenceInformation RetrievalSocial and Information Networks
AI summaryⓘ
The authors reviewed 74 studies on using advanced AI systems called agentic AI for analyzing open-source intelligence (OSINT), which involves gathering publicly available information for security and investigations. They organized the research into 11 key areas and found that many studies talk about AI hallucinating (making up information), but very few measure this problem specifically for OSINT tools. They also noticed that while AI helps well with collecting and analyzing data, it is less developed for verifying and reporting findings. The authors suggest that a good approach is to have AI assist humans in gathering information, while people stay in charge of checking facts and making decisions.
Open-Source Intelligence (OSINT)Large Language Models (LLMs)Agentic AIHallucination in AIRetrieval-Augmented Generation (RAG)Prompt EngineeringKnowledge GraphsAdversarial RobustnessMultimodal IntelligenceHuman-AI Collaboration
Authors
Eduardo Almeida Palmieri, Mohamed Chahine Ghanem, Dipo Dunsin, Zubair Baig, Ed de Quincey, Kim-Kwang Raymond Choo
Abstract
The rapid growth of publicly available digital information has rendered manual open-source intelligence (OSINT) analysis insufficient for modern intelligence, cybersecurity, and cyber investigation. Large language models (LLMs) and agentic AI systems, capable of tool use, multi-step reasoning, and iterative intelligence generation, have emerged as promising solutions, yet evaluation frameworks have not kept pace with reported capabilities. This survey systematically reviews 74 studies and makes four contributions. First, it establishes agentic AI as a distinct analytical category rather than an extension of LLM prompting, organising the literature through an 11-category taxonomy covering LLM foundations, agentic architectures, retrieval-augmented generation (RAG), knowledge graphs, prompt engineering, domain adaptation, evaluation benchmarks, and risk. Second, it identifies the hallucination-validation gap as a corpus-level finding: although hallucination is recognised as a major reliability concern in over twenty studies, end-to-end hallucination is empirically measured in only one OSINT-specific RAG-based system, non-reproducible conditions, while related reasoning and factual-correction studies evaluate general-domain question answering rather than OSINT. Third, it maps existing research to the OSINT lifecycle, showing strong support for collection and analysis but limited coverage of verification, reporting, dissemination, and decision support. Fourth, it derives a ten-point research agenda addressing evaluation, benchmarking, hallucination measurement, adversarial robustness, dark-web coverage, multimodal intelligence, and governance. It concludes that a human-AI co-pilot model, where LLMs assist collection and triage while analysts retain responsibility for verification and decision-making, represents the most defensible near-term deployment architecture.