Author ORCID Identifier
Jay Yang: https://orcid.org/0009-0004-5503-2082
Document Type
Article
Publication Title
IEEE Access
Abstract
In cybersecurity, security analysts constantly face the challenge of mitigating newly discovered vulnerabilities in real-time, with over 300,000 vulnerabilities identified since 1999. The sheer volume of known vulnerabilities complicates the detection of patterns for unknown threats. While LLMs can assist, they often hallucinate and lack alignment with recent threats. Over 40,000 vulnerabilities have been identified in 2024 alone, which are introduced after most popular LLMs’ (e.g., GPT-5) training data cutoff. This raises a major challenge of leveraging LLMs in cybersecurity, where accuracy and up-to-date information are paramount. Therefore, we aim to improve the adaptation of LLMs in vulnerability analysis by mimicking how an analyst performs such tasks. We propose ProveRAG, an LLM-powered system designed to assist in rapidly analyzing vulnerabilities with automated retrieval augmentation of web data while self-evaluating its responses with verifiable evidence. ProveRAG incorporates a self-critique mechanism to help alleviate the omission and hallucination common in the output of LLMs applied in cybersecurity applications. The system cross-references data from verifiable sources (NVD and CWE), giving analysts confidence in the actionable insights provided. Our results indicate that ProveRAG excels in delivering verifiable evidence to the user with over 99% and 97% accuracy in exploitation and mitigation strategies, respectively. ProveRAG guides analysts to secure their systems more effectively by overcoming temporal and context-window limitations while also documenting the process for future audits.
Pages
212815-212826
html
DOI
https://doi.org/10.1109/ACCESS.2025.3638251
Publisher
IEEE
Volume
13
Publication Date
12-2025
Keywords
ProveRAG, LLM, provenance, CVE, CWE, RAG, vulnerability, self-critique, Prevention and mitigation, Computer security, Retrieval augmented generation, Accuracy, Databases, Adaptation models, Training data, Training, Manuals, Large language models
Disciplines
Computer Sciences | Databases and Information Systems | Data Science | Information Security
ISSN
2169-3536
Recommended Citation
R. Fayyazi, S. H. Trueba, M. Zuzak and S. J. Yang, "ProveRAG: Provenance-Driven Vulnerability Analysis With Automated Retrieval-Augmented LLMs," in IEEE Access, vol. 13, pp. 212815-212826, 2025, doi: 10.1109/ACCESS.2025.3638251.
Upload File
wf_yes
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 International License.
Included in
Databases and Information Systems Commons, Data Science Commons, Information Security Commons
