Introduction
Artificial intelligence is rapidly changing how software is developed. AI-assisted programming tools can help developers generate code, troubleshoot errors, create application programming interface (API) endpoints, refactor existing applications and automate repetitive development tasks.
Tools such as GitHub Copilot and other large language model (LLM)-based coding assistants can reduce the effort required for many development tasks. More advanced coding agents can also edit multiple files, run commands, install dependencies and execute tests. These capabilities can improve development speed, but they do not by themselves establish that the resulting software is secure.
AI-generated code may compile, pass functional tests and appear technically correct while still containing exploitable weaknesses. GitHub advises users to review and test generated code, particularly in security-sensitive applications.
As organisations incorporate AI into development workflows, the volume and speed of change may increase. For web applications and APIs, independent security testing remains important for identifying access-control, workflow and business-logic weaknesses that developers, coding assistants and automated scanners may overlook.
For businesses, the challenge is to benefit from faster delivery without increasing exposure to breaches, service disruption, remediation costs and reputational damage. More security testing should mean assurance that keeps pace with risk and change, rather than a full penetration test for every minor edit. Secure design, developer review, automated controls, proportionate independent testing and monitoring all have a role.
AI-assisted coding: contents
The Rise of AI-Assisted Software Development
AI coding assistants are increasingly becoming part of modern software development workflows. Depending on the product, model and configuration, they can assist with:
generating source code and API endpoints
troubleshooting errors and explaining unfamiliar code
writing database queries and refactoring applications
generating unit and integration tests
creating configuration and infrastructure files
proposing fixes for suspected vulnerabilities.
Conventional assistants suggest changes for a developer to apply. Agentic tools can go further by planning multi-step work, editing files, invoking tools and executing commands. Their authority depends on the environment and permissions provided; GitHub's Copilot Agents application card describes these distinctions across its own products.
These capabilities can provide productivity benefits, but introduce a fundamental security consideration: code that functions correctly is not necessarily code that is secure. AI-generated changes need the same secure development discipline as manually written software, with assurance scaled appropriately where the frequency or impact of change increases.
Security Risks Associated with AI-Generated Code
AI coding tools generate output from prompts, available context, model behaviour and patterns learned during training. A plausible implementation is not evidence that it meets an application's security requirements. Results vary by model, task, language, context and configuration. Several familiar application weaknesses, alongside risks arising from agent access, deserve particular attention.
1. Incomplete Input Handling and Output Protection
Generated functionality may focus on expected inputs and successful paths without adequately handling malformed, malicious, oversized or semantically invalid data. Weaknesses can contribute to injection, cross-site scripting (XSS), path traversal, unsafe file processing, information leakage through errors and unexpected business-logic behaviour.
Input validation is only one control. Use parameterised database interfaces, context-appropriate output encoding, and sanitisation where applications intentionally accept HTML. Uploaded files need controls over type, size, storage and processing; structured input should be checked against an appropriate schema. Content Security Policy can add defence in depth, but does not replace safe output handling. See OWASP's SQL injection prevention and XSS prevention guidance.
2. Broken Object-Level Authorisation
An assistant may generate an endpoint that retrieves, modifies or deletes an object without fully understanding the application's authorisation requirements. For example:
GET /api/accounts/1001The endpoint may retrieve account 1001 correctly, but an important question remains: should the current user be permitted to access that account?
If object-level authorisation is missing, changing an identifier may expose another user's resource. OWASP calls this broken object-level authorisation (BOLA). Unpredictable identifiers can make enumeration harder, but do not replace an authorisation decision. Functional tests can miss the flaw because the endpoint returns exactly the data requested.
3. Authentication and Function-Level Authorisation Weaknesses
Authentication confirms who a user is; authorisation determines what that user may do. Generated code may verify a login while failing to enforce permissions for particular functions or resources.
This matters where APIs expose multiple roles, objects, endpoints and administrative operations. Test access between different users, across privilege levels and between tenants. Include denied operations, administrative functions and changes in identity or session state. The presence of authentication alone does not demonstrate that access controls work.
4. Insecure Database Interactions
AI tools can generate database queries quickly, but unsafe implementation may expose applications to injection or inappropriate data access. Use supported parameterised APIs or prepared statements, and avoid concatenating untrusted data into query syntax.
Where table or column names must vary, select them from an explicit allowlist: identifiers generally cannot be bound as ordinary query values. Check the privileges of the database account as well as the query construction.
5. Sensitive Information and Secrets
Development environments may contain API keys, authentication tokens, passwords, private keys, database credentials, internal URLs and proprietary source code. Depending on the tool, context can extend beyond the immediate prompt to open files, repository content, terminal output, logs and connected services.
OWASP's Secure Coding with AI guidance recommends controlling the context supplied to coding tools. Exclude sensitive files where possible and provide required secrets through approved secret-management mechanisms. A repository ignore file should not be assumed to prevent a tool from reading local files.
Understand what data each selected product can access or transmit. Evaluate retention, model-training use, residency, subcontractors, telemetry and deletion arrangements against organisational policy. These arrangements vary by product, plan, configuration and deployment model.
6. Dependency and Supply Chain Risks
Suggested libraries, packages and frameworks should not automatically be treated as trustworthy, maintained or suitable. Verify registry identity, source, maintainers, provenance, release history, licence, known vulnerabilities and compatibility before adding a dependency.
The NCSC Software Security Code of Practice implementation guidance provides the more direct context for ordinary software development and third-party components. Its separate secure AI development guidance addresses the development of AI systems; the two should not be treated as identical in scope.
Maintain a component inventory, lock approved versions and integrity information, and use controlled registries or repository proxies where practical. Continue checking dependencies after their initial approval.
7. Hallucinated Dependencies
AI coding assistants may recommend packages that do not exist or are incorrectly named. Research published through USENIX documents package hallucinations in code generated by language models.
An attacker could register a hallucinated or similar package name, hoping that a developer or agent installs it without verifying its provenance. A hallucinated name is not itself proof that a malicious package exists: the risk depends on the name being registered and subsequently installed.
Independently verify suggested packages and consider approved dependency lists and software composition analysis. Resolve unfamiliar dependencies in an isolated environment with restricted access to registries before promoting approved versions into trusted builds.
8. Indirect Prompt Injection
Coding agents may process source repositories, issue descriptions, pull requests, documentation, comments, error messages, dependency metadata, external web content and tool output. Connected tools can include Model Context Protocol (MCP) servers.
Malicious instructions embedded in that material can attempt to redirect the agent. For example, an agent resolving an issue might mistake attacker-controlled text in the issue description for an instruction to perform an unrelated action. The consequences become more serious when the agent can execute commands, modify files or access other systems.
Treat retrieved material as untrusted data, separate it from authorised instructions, minimise context, inspect unexpected changes and restrict tools and network destinations. Prompt filtering may help, but should not be presented as a complete defence. OWASP's prompt injection guidance describes a layered approach.
9. Excessive Agent Permissions and Project Context Exposure
Agents can have substantially greater access than code-completion tools. File modification, command execution, dependency installation and network access can expose the development environment to configuration mistakes, model errors or malicious input.
Apply least privilege and use controlled, sandboxed environments. Limit access to the workspace and resources needed for the task; restrict outbound networking and use short-lived, narrowly scoped credentials where credentials are necessary.
For consequential actions, bind approval to the actual operation, target and parameters. Enforce authorisation outside the model, particularly for destructive changes, publication, signing, production deployment and changes to security controls. OWASP's AI Agent Security guidance explains why tool permissions and execution controls need independent enforcement.
AI Changes the Scale of Software Development
One of the most significant implications of AI-assisted development is not necessarily the creation of entirely new vulnerability categories. It is the potential change in the speed and scale of software production.
Teams may generate additional API endpoints, features, database integrations, authentication workflows and third-party integrations in less time. This can produce a rapidly changing attack surface.
Even if the defect rate per change did not increase, a higher volume and frequency of change would increase the amount of code and attack-surface change requiring proportionate assurance. The effect varies by organisation, architecture, tasks and engineering maturity.
If security processes cannot keep pace, weaknesses may reach production more quickly and expose customer data, corporate information or critical services. Scale assurance through automation, reusable security requirements, secure framework defaults, targeted review and risk-based independent testing, rather than simply scheduling more full penetration tests.
Why Automated and Human-Led Testing Are Complementary
Static application security testing, dependency scanning, secret detection and dynamic testing remain important. Automated tools offer repeatability and broad coverage, while enforcing baseline checks across frequent changes.
Application-specific context can be harder to assess. Consider:
GET /api/customer/1001
GET /api/customer/1002Both endpoints may return expected responses. Whether Customer A should be allowed to access Customer B's information depends on the intended access-control model. Useful testing requires an explicit statement of which user or tenant may perform which action on which resource, and in which state.
Similar challenges arise with business-logic flaws, privilege escalation, workflow manipulation, chained vulnerabilities and application-specific abuse cases. Automated and AI-assisted tools can help identify patterns and generate tests, but still need correct requirements and meaningful expected outcomes. OWASP's workflow testing guidance illustrates the importance of developing misuse cases from the application's intended process.
Limitations of AI-Generated Security Tests
AI-generated tests can improve coverage and reduce effort, but they are not automatically independent assurance of generated code. Tests produced within the same task, context or specification may reproduce the implementation's assumptions and omissions.
Passing tests or high coverage can therefore create false confidence if the wrong scenarios were tested. Authentication, authorisation, tenant isolation, input handling, cryptography and sensitive-data processing should receive independently designed negative and abuse-case tests. Review deleted tests, weakened assertions and excessive mocking as carefully as new code. AI-generated test cases can augment this work; they should not be its sole source.
Risk-Based Web Application and API Security Testing
Web applications and APIs expose functionality to users, mobile applications, third-party services and other systems. New endpoints, integrations and features can change the attack surface even when the underlying business purpose remains familiar.
Independent testing assesses behaviour under realistic attack, misuse and cross-user scenarios. Depth and timing should reflect risk: an authentication redesign, a new internet-facing API, a payment workflow or a changed tenant boundary normally warrants greater scrutiny than a low-risk presentation change.
For AI-assisted applications, assessment may cover:
authentication, authorisation and session management
object-level and function-level API access controls
input handling and file upload functionality
sensitive data exposure and error handling
business logic and workflow abuse
third-party integrations
interactions between users, tenants and privilege levels.
Web application penetration testing can identify weaknesses that scanning and functional tests miss, especially where exploitation depends on business context or chained behaviour. The API security testing guide explains related considerations for exposed interfaces.
Penetration testing is a point-in-time assessment, not a guarantee of security. It should complement secure design, code review, automated controls, vulnerability management and monitoring. Agree a scope, test accounts, permissions, environment and rules of engagement before testing, then use findings to strengthen ongoing assurance.
Example AI-Assisted Development Security Flow
Select the diagram to enlarge it. Detailed steps follow.
Define security requirements and a threat model.
Scope the task, project context, tools and agent permissions.
Generate changes in an isolated branch or workspace, with execution restrictions where needed.
Review the complete diff, dependencies and test evidence.
Run build checks, unit, integration and negative tests, SAST, SCA and secret scanning.
Perform independent security review and risk-based application/API testing.
Remediate findings and retest the fixes.
Authorise the merge and deploy through controlled processes.
Monitor, respond to vulnerabilities and feed lessons into regression tests.
This approach does not require organisations to avoid AI-assisted development. It integrates controls before, during and after implementation so development speed does not outpace assurance. The NIST Secure Software Development Framework, version 1.1, places secure production alongside organisational preparation, software protection and vulnerability response.
Defensive Measures
Organisations adopting AI-assisted development should apply layered controls across the software lifecycle.
Maintain Human Oversight
Every AI-assisted change should have an accountable human owner who understands its intended behaviour, reviews the full diff and dependency changes, and examines test evidence before merge. For high-risk changes, include a reviewer who did not author the implementation or its sole generated test suite.
Perform Secure Code Reviews
Review security assumptions, dependencies, access controls and data handling, alongside functional correctness. Give additional scrutiny to authentication, authorisation, cryptography and sensitive data processing.
Check unexpected changes outside scope, disabled controls, altered CI workflows, new network destinations, generated configuration and sensitive logging. A clean automated review should not be the only approval criterion. See our code review guide for the role of security-focused review.
Apply Secure Design and Development Practices
Define security requirements, trust boundaries and misuse cases before generation. Use established secure patterns and framework defaults. Repository instructions can describe approved practices, but important requirements should also be enforced through code, tests, policy and continuous integration controls rather than prompts alone.
Integrate Automated Security Testing
Choose tools appropriate to the application and development pipeline. These may include:
compiler warnings, type checking and secure linting
static application security testing (SAST)
software composition analysis (SCA), provenance and licence checks
secret scanning
infrastructure-as-code and container scanning where relevant
dynamic application security testing (DAST)
fuzzing or property-based tests for parsers and security boundaries
regression tests for previously identified vulnerabilities.
Protect Sensitive Development Data
Apply data-classification rules to prompts, source code and tool output. Restrict sensitive project context, review product data-handling settings and avoid unnecessarily exposing secrets. Access exclusions should be tested rather than assumed to work.
Apply Least Privilege and Sandboxing
Limit agents to the resources needed for the task, with sandbox and approval policies enforced outside the model. Restrict outbound networking and keep production, signing and broad administrative credentials outside routine agent environments. Any credentials genuinely needed should be short-lived and narrowly scoped.
Use action-specific approvals for consequential operations. Broad approval of an entire session should not substitute for checking what will be changed, where it will be changed and with which permissions.
Perform Proportionate Independent Testing
Assess internet-facing and security-sensitive applications before initial release and after material changes to attack surface, trust boundaries, authentication, authorisation, sensitive-data processing or critical workflows. Review testing frequency against exposure, threat intelligence, applicable obligations and the effectiveness of continuous controls.
AI-assisted development does not make every change equally risky. Target deeper assessment where the possible impact or uncertainty is greatest.
Retest Identified Vulnerabilities
Confirm that fixes address the underlying cause without introducing new weaknesses. Turn relevant findings into repeatable regression tests, and check related code paths where the same insecure pattern may have been reused.
Monitor After Deployment
Maintain useful security telemetry, update components, investigate unexpected behaviour and provide a process for receiving and resolving vulnerability reports. Changes to the threat landscape or a dependency can affect software that passed its original tests. Feed operational findings back into design, review and testing.
Manage AI Coding Tools as Part of Engineering Governance
Maintain an inventory of approved coding assistants, models, agent frameworks, plugins, MCP servers and significant configurations. Define the repositories and data classifications each may access, retain appropriate audit records and assess material upgrades before wider rollout.
Treat instruction files, hooks, tool definitions and MCP configuration as security-sensitive code. Version-control and review changes, monitor unexpected modifications, and retain branch protection, mandatory checks and deployment approval outside the agent's authority.
For connected tools, use approved servers, scoped credentials and validation of tool inputs and outputs. OWASP's MCP Security guidance provides further controls for those integrations.
Security Must Keep Pace with Development
AI assistance is neither a security control nor proof that software is insecure. Risk depends on the change, architecture, model and agent configuration, tool permissions and assurance controls applied before and after deployment.
The important question is therefore not simply whether developers should use AI, but whether security assurance can keep pace with the way the organisation develops and changes software.
Secure development, constrained agent execution, automated controls, review and independent testing should evolve together. Their success should be judged through evidence of correct security behaviour, not the confidence of generated explanations or test-pass counts alone.
Conclusion
AI-assisted coding offers opportunities to improve productivity and accelerate delivery, but can also create a faster-changing attack surface. Code should not be considered secure simply because it functions correctly or passes automated tests.
Web applications and APIs still need careful assessment of authentication, authorisation, input handling, business logic and sensitive data. The appropriate response is layered, proportionate security assurance, with independent testing directed at material risks and controls maintained after deployment.
The speed at which an organisation can create software should not exceed its ability to secure it.
To discuss security testing for AI-assisted web applications or APIs, contact ProCheckUp. If the product itself uses an LLM, retrieval or agentic functionality, our AI security testing service addresses that separate application-specific scope.
References
OWASP API Security Top 10 (2023) — Broken Object Level Authorization
NCSC — Software Security Code of Practice: Secure Design and Development
NCSC — Guidelines for Secure AI System Development: Secure Development
USENIX — Package Hallucinations: How LLMs Can Invent Vulnerabilities
OWASP Web Security Testing Guide — Testing for Circumvention of Workflows
NIST — Secure Software Development Framework (SSDF), Version 1.1




Categories