⚠ Official Notice: www.ijisrt.com is the official website of the International Journal of Innovative Science and Research Technology (IJISRT) Journal for research paper submission and publication. Please beware of fake or duplicate websites using the IJISRT name.



Development of the SecurePromptTrace Algorithm for Detecting Prompt Injection, Data Exfiltration, and Tool Misuse in Generative Artificial Intelligence Systems with Comparative Evaluation Against Keyword Filters, Classifier-Based Defenses, and Static Access Controls


Authors : Praise Elojo Attah; Lawrence Anebi Enyejo

Volume/Issue : Volume 11 - 2026, Issue 8 - August


Google Scholar : https://tinyurl.com/4pbamzh6

Scribd : https://tinyurl.com/3zdedptj

DOI : https://doi.org/10.38124/ijisrt/26aug331

Note : A published paper may take 4-5 working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and ResearchGate.


Abstract : Generative artificial intelligence systems increasingly operate as autonomous agents capable of retrieving external information, accessing confidential resources, invoking application programming interfaces, and executing consequential actions. These capabilities introduce substantial security risks because malicious instructions embedded in user prompts, retrieved documents, webpages, emails, tool outputs, or persistent memory may alter an agent’s intended behaviour. Conventional keyword filters, standalone prompt classifiers, and static access-control mechanisms provide limited protection against semantically obfuscated attacks, multi-stage data exfiltration, manipulated tool arguments, and attacks that remain within formally permitted privileges. This paper develops SecurePromptTrace, a novel runtime security algorithm for detecting and controlling prompt injection, sensitive-data exfiltration, and tool misuse in generative artificial intelligence systems. SecurePromptTrace constructs a Dynamic Prompt Provenance Graph that represents trusted instructions, untrusted content, model-generated plans, retrieved data, confidential variables, tool calls, tool arguments, and execution outcomes as provenance-labelled nodes and causal edges. A relation-aware graph attention network analyses instruction dependencies and identifies conflicts between the authenticated user objective and instructions originating from untrusted sources. The graph model is integrated with a DeBERTa-v3 semantic injection classifier, confidential-data taint propagation, cross-layer intent alignment, tool-capability compatibility analysis, and adaptive policy enforcement. A composite threat score combines semantic injection probability, provenance conflict, sensitive-data flow, tool-privilege mismatch, execution-sequence deviation, and predictive uncertainty. According to the calculated risk, the algorithm permits, sanitises, replans, isolates, requests approval for, or blocks an operation. The evaluation framework covers direct and indirect prompt injection, encoded and multilingual attacks, contextwindow attacks, memory poisoning, cross-tool exfiltration, parameter substitution, privilege chaining, and unauthorised tool execution. SecurePromptTrace is compared with Aho–Corasick keyword filtering, regular-expression filtering, a standalone DeBERTa-v3 classifier, PromptShield-style detection, role-based access control, attribute-based access control, and combined static guardrails. Performance is assessed using macro-F1 score, precision, recall, AUROC, AUPRC, attack success rate, data-exfiltration prevention rate, tool-misuse prevention rate, false-positive rate, task-utility retention, computational latency, and memory overhead. Comparative bar charts, ROC and precision–recall curves, confusion matrices, risk-score distributions, ablation graphs, and security–latency Pareto plots are used to demonstrate performance differences. The central hypothesis is that provenance-aware semantic and behavioural tracing will produce significantly lower attack-success and exfiltration rates than content-only or permission-only defenses while preserving legitimate task completion. Statistical superiority will be established through confidence intervals, McNemar tests, bootstrap comparisons, and effect-size analysis.

Keywords : Prompt Injection; Data Exfiltration; Tool Misuse; Generative Artificial Intelligence; Static Access Controls.

References :

  1. Adewale, L. D. (2026). Digital evidence chains for PPAP assurance: AR-guided data capture, AI-verified documentation, and continuous audit automation for secure multi-tier supplier traceability in Industry 4.0 manufacturing. International Journal of Multidisciplinary Evolutionary Research, 7(1), 43–55. https://doi.org/10.54660/ijmer.2026.7.1.43-55
  2. Adewale, L. D. (2026). Smart factories, smarter evidence: Reinventing quality assurance for U.S. manufacturing competitiveness. International Journal of Multidisciplinary Futuristic Development, 7(1), 9–18. https://doi.org/10.54660/IJMFD.2026.7.1.09-18
  3. Agyekum, B., Seffah-Duodu, D., & Gloria, A. (2025). Federated industrial IoT threats detection. International Journal of Scientific Research and Modern Technology, 4(12), 226–244. https://doi.org/10.38124/ijsrmt.v4i12.1552
  4. Alotaibi, A., Mughus, R., & Ahmed, M. (2026). A red teaming framework for large language models: A case study on faithfulness evaluation. Software Quality Journal, 34, Article 42. https://doi.org/10.1007/s11219-026-09779-y
  5. Alzahrani, A. (2026). PromptGuard: A structured framework for injection-resilient language models. Scientific Reports, 16, Article 1277. https://doi.org/10.1038/s41598-025-31086-y
  6. Ayoola, V. B., Ugoaghalam, U. J., Idoko, I. P., Ijiga, O. M., and Olola, T. M. (2024). Effectiveness of social engineering awareness training in mitigating spear phishing risks in financial institutions from a cybersecurity perspective. Global Journal of Engineering and Technology Advances, 20(3), 94–117. https://doi.org/10.30574/gjeta.2024.20.3.0164
  7. Ayoola, V. B., Ugochukwu, U. N., Adeleke, I., Michael, C. I., Adewoye, M. B., and Adeyeye, Y. (2024). Generative AI-driven fraud detection in health care: Enhancing data loss prevention and cybersecurity analytics for real-time protection of patient records. International Journal of Scientific Research and Modern Technology, 3(11), 89–107. https://doi.org/10.38124/ijsrmt.v3i11.112
  8. Boiko, D. A., MacKnight, R., Kline, B., & Gomes, G. (2023). Autonomous chemical research with large language models. Nature, 624, 570–578. https://doi.org/10.1038/s41586-023-06792-0
  9. Briskinfosec, (2026). The Hidden Risk of Data Leakage in AI Code Assistants https://www.briskinfosec.com/blogs/blogsdetail/the-hidden-risk-of-data-leakage-in-ai-code-assistants
  10. Cai, Y. (2026). Prompt injection attacks on educational large language models for higher and vocational education. Scientific Reports, 16, Article 15594. https://doi.org/10.1038/s41598-026-46563-1
  11. Chen, K., Zhou, X., Lin, Y., Feng, S., Shen, L., & Wu, P. (2025). A survey on privacy risks and protection in large language models. Journal of King Saud University–Computer and Information Sciences, 37, Article 163. https://doi.org/10.1007/s44443-025-00177-1
  12. Clusmann, J., Ferber, D., Wiest, I. C., Schneider, C. V., Brinker, T. J., Foersch, S., Truhn, D., & Kather, J. N. (2025). Prompt injection attacks on vision language models in oncology. Nature Communications, 16, Article 1239. https://doi.org/10.1038/s41467-024-55631-x
  13. Debenedetti, E., Shumailov, I., Fan, T., Hayes, J., Carlini, N., Fabian, D., Kern, C., Shi, C., Terzis, A., & Tramèr, F. (2025). Defeating prompt injections by design. arXiv:2503.18813.
  14. Debenedetti, E., Zhang, J., Balunović, M., Beurer-Kellner, L., Fischer, M., & Tramèr, F. (2024). AgentDojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents. Advances in Neural Information Processing Systems, 37.
  15. Dzhaliuk, N., Sabodashko, D., Khoma, V., Khoma, Y., Kolchenko, V., and Podpora, M. (2026). Comparative evaluation of machine learning methods for protecting LLMs from prompt injection attacks. International Journal of Information Security, 25, Article 109. https://doi.org/10.1007/s10207-026-01264-8
  16. Gabla, E. S., Peter-Anyebe, A. C., & Ijiga, O. M. (2025). Assessing machine-learning-enabled anomaly detection models for real-time cyberattack mitigation in optical fiber communication systems. World Journal of Advanced Engineering Technology and Sciences, 17(2), 1–17. https://doi.org/10.30574/wjaets.2025.17.2.1454
  17. Hagendorff, T., Derner, E., & Oliver, N. (2026). Large reasoning models are autonomous jailbreak agents. Nature Communications, 17, Article 1435. https://doi.org/10.1038/s41467-026-69010-1
  18. He, P., Gao, J., & Chen, W. (2021). DeBERTaV3: Improving DeBERTa using ELECTRA-style pre-training with gradient-disentangled embedding sharing. arXiv:2111.09543.
  19. Ibokette, A. I., Ogundare, T. O., Anyebe, A. P., Alao, F. O., Odeh, I. I., and Okafor, F. C. (2024). Mitigating maritime cybersecurity risks using AI-based intrusion detection systems and network automation during extreme environmental conditions. International Journal of Scientific Research and Modern Technology, 3(10), 65–91. https://doi.org/10.38124/ijsrmt.v3i10.73
  20. Idika, C. N., James, U. U., Ijiga, O. M., & Enyejo, L. A. (2023). Digital twin-enabled vulnerability assessment with Zero Trust policy enforcement in smart manufacturing cyber-physical systems. International Journal of Scientific Research in Computer Science, Engineering and Information Technology, 9(6), 475–499. https://doi.org/10.32628/CSEIT23906189
  21. Idoko, I. P., Ijiga, O. M., Enyejo, L. A., Akoh, O., Ebiega, G. I., Odeyemi, M. O., Olatunde, T. I., and Olajide, F. I. (2024). Integrating superhumans and synthetic humans into the Internet of Things and ubiquitous computing: Emerging AI applications and their relevance in the U.S. context. Global Journal of Engineering and Technology Advances, 19(1), 6–36. https://doi.org/10.30574/gjeta.2024.19.1.0055
  22. Igba, E., Olarinoye, H. S., Ezeh, N. V., Sehemba, D. B., Oluhaiyero, Y. S., & Okika, N. (2025). Synthetic data generation using generative AI to combat identity fraud and enhance global financial cybersecurity frameworks. International Journal of Scientific Research and Modern Technology, 4(2), 1–19. https://doi.org/10.5281/zenodo.14928919
  23. Ijiga, O. M., Idoko, I. P., Ebiega, G. I., Olajide, F. I., Olatunde, T. I., and Ukaegbu, C. (2024). Harnessing adversarial machine learning for advanced threat detection: AI-driven strategies in cybersecurity risk assessment and fraud prevention. Open Access Research Journal of Science and Technology, 11(1), 1–24. https://doi.org/10.53022/oarjst.2024.11.1.0060
  24. Ijiga, O. M., Okika, N., Balogun, S. A., Enyejo, L. A., and Agbo, O. J. (2025). A comprehensive review of federated learning architectures for insider threat detection in distributed SQL-based enterprise environments. International Journal of Innovative Science and Research Technology, 10(7), 536–550. https://doi.org/10.38124/ijisrt/25jul392
  25. Jalan, P., Abishethvarman, V., Chandna, B., & Naseem, U. (2026). Survey on LLM safety: Attacks, defenses, alignment, metrics, and guardrails. Machine Learning, 115, Article 130. https://doi.org/10.1007/s10994-026-07060-8
  26. James, U. U. (2026). Multiscale Signal Processing Framework for Zero Trust Security in Next- Generation Wireless Networks International Journal of Wireless and Mobile Networks DOI/Link: https://aircconline.com/ijwmn/V18N3/18326ijwmn01.pdf
  27. James, U. U., and Akujuobi, C. M. (2026). Multiscale signal processing framework for Zero Trust security in next-generation wireless networks. International Journal of Wireless and Mobile Networks, 18(3), 1–21. https://doi.org/10.5121/ijwmn.2026.18301
  28. James, U. U., Olarinoye, H. S., Uchenna, I. R., Idika, C. N., Ngene, O. J., Ijiga, O. M., and Itemuagbor, K. (2025). Combating deepfake threats using X-FACTS explainable CNN framework for enhanced detection and cybersecurity resilience. Advances in Artificial Intelligence and Robotics Research, 1, 41–64. https://doi.org/10.4236/airr.2025.11004
  29. James, U. U., Olarinoye, H. S., Uchenna, I. R., Idika, C. N., Ngene, O. J., Ijiga, O. M., & Itemuagbor, K. (2025). Combating deepfake threats using X-FACTS explainable CNN framework for enhanced detection and cybersecurity resilience. Advances in Artificial Intelligence and Robotics Research, 1, 41–64. https://doi.org/10.4236/airr.2025.11004
  30. James, U. U., Salami, E. O., & Enyejo, L. A. (2025). Real-time policy orchestration for cybersecurity risk management in GRC-aligned financial technology infrastructures. International Journal of Innovative Science and Research Technology, 10(8). https://doi.org/10.38124/ijisrt/25aug1021
  31. James, U. U., Salami, E. O., & Enyejo, L. A. (2025). Real-time policy orchestration for cybersecurity risk management in GRC-aligned financial technology infrastructures. International Journal of Innovative Science and Research Technology, 10(8). https://doi.org/10.38124/ijisrt/25aug1021
  32. James, U. U., Salami, E. O., & Enyejo, L. A. (2025). Real-time policy orchestration for cybersecurity risk management in GRC-aligned financial technology infrastructures. International Journal of Innovative Science and Research Technology, 10(8). https://doi.org/10.38124/ijisrt/25aug1021
  33. Ko, R., Chung, J., Zheng, S., Xiao, C., Kim, T.-W., Onizuka, M., & Shin, W.-Y. (2026). Seven security challenges in cross-domain multi-agent LLM systems. npj Artificial Intelligence. https://doi.org/10.1038/s44387-026-00128-9
  34. Kpogli, S. A., Onwuzurike, M. A., & Ayoola, V. B. (2024). Optimizing Educational Return on Investment Through AI-Driven Curriculum Adaptation and Performance-Based Resource Allocation Models. International Journal of Scientific Research and Modern Technology, 3(9), 141–156. https://doi.org/10.38124/ijsrmt.v3i9.1352
  35. Kpogli, S. A., Onwuzurike, M. A., & Enyejo, J. O. (2024). Integrating artificial intelligence and learning sciences to reduce cognitive load and achievement gaps in data-driven K–12 instructional systems. International Journal of Scientific Research in Computer Science, Engineering and Information Technology, 10(6), 2569–2589. https://doi.org/10.32628/CSEIT25113575
  36. Kwarteng, R. A., Idoko, I. P., Ijiga, O. M., & Enyejo, L. A. (2020). Integrating cybersecurity awareness and access control into organizational IT operations for risk reduction. International Journal of Scientific Research in Computer Science, Engineering and Information Technology, 6(1), 243–261. https://doi.org/10.32628/CSEIT23906128
  37. Milani, A., Franzoni, V., & Florindi, E. (2026). Indirect prompt injection in large language models. Neural Computing and Applications, 38, Article 530. https://doi.org/10.1007/s00521-026-12266-x
  38. Okika, N., Okoh, O. F., & Etuk, E. E. (2025). Mitigating insider threats and social engineering tactics in advanced persistent threat operations through behavioral analytics and cybersecurity training. International Journal of Advance Research Publication and Reviews, 2(3), 11–27.
  39. Ononiwu, M., Azonuche, T. I., Imoh, P. O., & Enyejo, J. O. (2024). Evaluating blockchain content monetization platforms for autism-focused streaming with cybersecurity and scalable microservice architectures. Iconic Research and Engineering Journals, 8(1), 722–740.
  40. Onwuzurike, M. A. & Enyejo, J. O. (2026). A Business Intelligence Framework for AI Powered Educational Platforms Linking Learning Analytics to Strategic Decision Making in K-12 Schools International Journal of Recent Research in Commerce Economics and Management (IJRRCEM) Vol. 13, Issue 2, pp: (21-42), DOI: https://doi.org/ 10.5281/zenodo.19510038
  41. Onwuzurike, M. A. (2023). Human-Centered Design of Intelligent Tutoring Systems Integrating Behavioral Analytics and Inclusive Pedagogical Principles for Early Learners International Journal of Scientific Research in Science, Engineering and Technology Volume 10, Issue 3, Page Number 720-738, doi : https://doi.org/10.32628/IJSRSET2310330
  42. Onwuzurike, M. A., & Kpogli, S. A. (2022). Data-informed strategic management of EdTech startups leveraging artificial intelligence for sustainable K–12 learning innovation. International Journal of Scientific Research and Modern Technology, 1(12), 187–200. https://doi.org/10.38124/ijsrmt.v1i12.1117
  43. Onwuzurike, M. A., & Kpogli, S. A. (2025). Predictive modeling of student engagement and behavioral outcomes using machine learning techniques in technology-enhanced classrooms. International Journal of Scientific Research in Humanities and Social Sciences, 2(6), 58–79. https://doi.org/10.32628/IJSRHSS2525135
  44. Onwuzurike, M. A., & Raphael, F. O. (2025). Ethical governance models for artificial intelligence deployment in K–12 education: Balancing algorithmic personalization, accountability and child protection policy. International Journal of Scientific Research and Modern Technology, 4(8), 193–208. https://doi.org/10.38124/ijsrmt.v4i8.1271
  45. Onwuzurike, M. A., Enyejo, J. O., & Peter-Anyebe, A. C. (2026). Design and evaluation of real-time adaptive learning algorithms for personalized K–12 curriculum optimization using student performance analytics. World Journal of Advance Multidisciplinary Research, 3(3), 21–36. https://doi.org/10.5281/zenodo.19131296
  46. Onwuzurike, M. A., Igba, E. (2023). Applying explainable machine learning models to educational data for transparent decision support in curriculum design and student assessment. Journal of Frontiers in Multidisciplinary Research. 2023;4(1):585–599. doi:10.54660/.JFMR.2023.4.1.585-599
  47. Shilov, I., Meeus, M., & de Montjoye, Y.-A. (2026). The mosaic memory of large language models. Nature Communications, 17, Article 2142. https://doi.org/10.1038/s41467-026-68603-0
  48. Srinivasan, J., Regi, S. A., Anbarasan, A. K., Suresh, A., Vetriselvi, T., & Venu, S. (2026). Detection and analysis of prompt injection in Indian multilingual large language models. Scientific Reports, 16, Article 16208. https://doi.org/10.1038/s41598-026-43883-0
  49. Ussher-Eke, D., Onoja, D. A., Ijiga, O. M., & Enyejo, L. A. (2025). Strengthening human resource compliance and ethical oversight through cybersecurity awareness and policy enforcement. International Journal of Management and Commerce Innovations, 13(1), 376–393. https://doi.org/10.5281/zenodo.16539315
  50. Wang, W., Liu, J., Gu, B., Wang, C., Qu, Y. N., & Feng, D. (2026). Securing LLM-in-the-loop software for empirical study of risks, mitigations, and utility trade-offs in a safety-critical case. Empirical Software Engineering, 31(4), Article 87. https://doi.org/10.1007/s10664-026-10820-8
  51. Webb, T., Mondal, S. S., & Momennejad, I. (2025). A brain-inspired agentic architecture to improve planning with LLMs. Nature Communications, 16, Article 8633. https://doi.org/10.1038/s41467-025-63804-5
  52. Yang, Y., Jin, Q., Huang, F., and Lu, Z. (2025). Adversarial prompt and fine-tuning attacks threaten medical large language models. Nature Communications, 16, Article 9011. https://doi.org/10.1038/s41467-025-64062-1

Generative artificial intelligence systems increasingly operate as autonomous agents capable of retrieving external information, accessing confidential resources, invoking application programming interfaces, and executing consequential actions. These capabilities introduce substantial security risks because malicious instructions embedded in user prompts, retrieved documents, webpages, emails, tool outputs, or persistent memory may alter an agent’s intended behaviour. Conventional keyword filters, standalone prompt classifiers, and static access-control mechanisms provide limited protection against semantically obfuscated attacks, multi-stage data exfiltration, manipulated tool arguments, and attacks that remain within formally permitted privileges. This paper develops SecurePromptTrace, a novel runtime security algorithm for detecting and controlling prompt injection, sensitive-data exfiltration, and tool misuse in generative artificial intelligence systems. SecurePromptTrace constructs a Dynamic Prompt Provenance Graph that represents trusted instructions, untrusted content, model-generated plans, retrieved data, confidential variables, tool calls, tool arguments, and execution outcomes as provenance-labelled nodes and causal edges. A relation-aware graph attention network analyses instruction dependencies and identifies conflicts between the authenticated user objective and instructions originating from untrusted sources. The graph model is integrated with a DeBERTa-v3 semantic injection classifier, confidential-data taint propagation, cross-layer intent alignment, tool-capability compatibility analysis, and adaptive policy enforcement. A composite threat score combines semantic injection probability, provenance conflict, sensitive-data flow, tool-privilege mismatch, execution-sequence deviation, and predictive uncertainty. According to the calculated risk, the algorithm permits, sanitises, replans, isolates, requests approval for, or blocks an operation. The evaluation framework covers direct and indirect prompt injection, encoded and multilingual attacks, contextwindow attacks, memory poisoning, cross-tool exfiltration, parameter substitution, privilege chaining, and unauthorised tool execution. SecurePromptTrace is compared with Aho–Corasick keyword filtering, regular-expression filtering, a standalone DeBERTa-v3 classifier, PromptShield-style detection, role-based access control, attribute-based access control, and combined static guardrails. Performance is assessed using macro-F1 score, precision, recall, AUROC, AUPRC, attack success rate, data-exfiltration prevention rate, tool-misuse prevention rate, false-positive rate, task-utility retention, computational latency, and memory overhead. Comparative bar charts, ROC and precision–recall curves, confusion matrices, risk-score distributions, ablation graphs, and security–latency Pareto plots are used to demonstrate performance differences. The central hypothesis is that provenance-aware semantic and behavioural tracing will produce significantly lower attack-success and exfiltration rates than content-only or permission-only defenses while preserving legitimate task completion. Statistical superiority will be established through confidence intervals, McNemar tests, bootstrap comparisons, and effect-size analysis.

Keywords : Prompt Injection; Data Exfiltration; Tool Misuse; Generative Artificial Intelligence; Static Access Controls.

Paper Submission Last Date
31 - August - 2026

SUBMIT YOUR PAPER CALL FOR PAPERS
Video Explanation for Published paper

Never miss an update from Papermashup

Get notified about the latest tutorials and downloads.

Subscribe by Email

Get alerts directly into your inbox after each post and stay updated.
Subscribe
OR

Subscribe by RSS

Add our RSS to your feedreader to get regular updates from us.
Subscribe