The Double-Edged Sword of AI Pair Programmers: A Systematic Literature Review of Security Vulnerabilities in AI-Generated Code and Agentic Development Environments
DOI:
https://doi.org/10.66823/2egmd793Keywords:
AI-assisted software development, GitHub Copilot, Cursor, code large language models, secure code generation, prompt injection, agentic AI, Model Context Protocol, software supply chain, systematic literature reviewAbstract
The role of AI pair programmers has expanded from local code completion to active participation in the development environment. Contemporary tools can interpret repository context, edit multiple files, call package managers, execute terminal commands, and communicate with external services. This review synthesizes security evidence concerning GitHub Copilot, ChatGPT-based coding, code large language models, Cursor-style agentic editors, command-line coding agents, and Model Context Protocol ecosystems. A protocol-driven search, supplemented by backward and forward snowballing, covered literature and technical evidence available through 12 July 2026. The verified corpus comprised 216 verified records, including peer-reviewed studies, preprints, benchmarks, standards, and clearly identified technical disclosures. The evidence supports useful roles in vulnerability discovery and repair, test generation, and secure-coding guidance, but it also documents recurring injection flaws, unsafe memory and file handling, weak cryptography, authentication errors, secret exposure, hallucinated dependencies, and incomplete patches. Whether these weaknesses persist depends partly on human factors, including expertise, prompt framing, review effort, automation bias, and the authority users assign to the assistant. Most generated-code defects remain familiar CWE classes. Agentic systems add orchestration-level concerns: indirect prompt injection, context poisoning, tool and protocol supply-chain attacks, permission amplification, approval spoofing, cross-file persistence, and autonomous execution. We synthesize these findings in a unified taxonomy, an evidence-based threat model, a lifecycle security model, and a continuous-assurance architecture built on structured context, least privilege, sandboxing, provenance, conventional SAST/SCA/secret scanning, security tests, and mandatory human authorization for high-impact actions. Without such governance, faster production may be offset by accumulating validation debt and downstream incident risk.
Downloads
Published
Data Availability Statement
No data is used during the development of this work
Issue
Section
License
Copyright (c) 2026 Journal of Sustainable Smart Systems in Education & Environment

This work is licensed under a Creative Commons Attribution 4.0 International License.


