by

AI-Assisted Coding: Why Faster Development Demands More Security Testing

Introduction

Artificial intelligence is rapidly changing how software is developed. AI-assisted programming tools can help developers generate code, troubleshoot errors, create application programming interface (API) endpoints, refactor existing applications and automate repetitive development tasks.

Tools such as GitHub Copilot and other large language model (LLM)-based coding assistants can reduce the effort required for many development tasks. More advanced coding agents can also edit multiple files, run commands, install dependencies and execute tests. These capabilities can improve development speed, but they do not by themselves establish that the resulting software is secure.

AI-generated code may compile, pass functional tests and appear technically correct while still containing exploitable weaknesses. GitHub advises users to review and test generated code, particularly in security-sensitive applications.

As organisations incorporate AI into development workflows, the volume and speed of change may increase. For web applications and APIs, independent security testing remains important for identifying access-control, workflow and business-logic weaknesses that developers, coding assistants and automated scanners may overlook.

For businesses, the challenge is to benefit from faster delivery without increasing exposure to breaches, service disruption, remediation costs and reputational damage. More security testing should mean assurance that keeps pace with risk and change, rather than a full penetration test for every minor edit. Secure design, developer review, automated controls, proportionate independent testing and monitoring all have a role.

AI-assisted coding: contents

The Rise of AI-Assisted Software Development

AI coding assistants are increasingly becoming part of modern software development workflows. Depending on the product, model and configuration, they can assist with:

  • generating source code and API endpoints

  • troubleshooting errors and explaining unfamiliar code

  • writing database queries and refactoring applications

  • generating unit and integration tests

  • creating configuration and infrastructure files

  • proposing fixes for suspected vulnerabilities.

Conventional assistants suggest changes for a developer to apply. Agentic tools can go further by planning multi-step work, editing files, invoking tools and executing commands. Their authority depends on the environment and permissions provided; GitHub's Copilot Agents application card describes these distinctions across its own products.

These capabilities can provide productivity benefits, but introduce a fundamental security consideration: code that functions correctly is not necessarily code that is secure. AI-generated changes need the same secure development discipline as manually written software, with assurance scaled appropriately where the frequency or impact of change increases.

Security Risks Associated with AI-Generated Code

AI coding tools generate output from prompts, available context, model behaviour and patterns learned during training. A plausible implementation is not evidence that it meets an application's security requirements. Results vary by model, task, language, context and configuration. Several familiar application weaknesses, alongside risks arising from agent access, deserve particular attention.

1. Incomplete Input Handling and Output Protection

Generated functionality may focus on expected inputs and successful paths without adequately handling malformed, malicious, oversized or semantically invalid data. Weaknesses can contribute to injection, cross-site scripting (XSS), path traversal, unsafe file processing, information leakage through errors and unexpected business-logic behaviour.

Input validation is only one control. Use parameterised database interfaces, context-appropriate output encoding, and sanitisation where applications intentionally accept HTML. Uploaded files need controls over type, size, storage and processing; structured input should be checked against an appropriate schema. Content Security Policy can add defence in depth, but does not replace safe output handling. See OWASP's SQL injection prevention and XSS prevention guidance.

2. Broken Object-Level Authorisation

An assistant may generate an endpoint that retrieves, modifies or deletes an object without fully understanding the application's authorisation requirements. For example:

GET /api/accounts/1001

The endpoint may retrieve account 1001 correctly, but an important question remains: should the current user be permitted to access that account?

If object-level authorisation is missing, changing an identifier may expose another user's resource. OWASP calls this broken object-level authorisation (BOLA). Unpredictable identifiers can make enumeration harder, but do not replace an authorisation decision. Functional tests can miss the flaw because the endpoint returns exactly the data requested.

An authenticated User A passes through server-side authorisation: access to their own account 1001 is allowed, while access to User B's account 1002 is denied.

A valid login does not authorise access to every object. In this example, the server permits User A to read account 1001 and denies access to User B's account 1002. Select the diagram to enlarge it.

3. Authentication and Function-Level Authorisation Weaknesses

Authentication confirms who a user is; authorisation determines what that user may do. Generated code may verify a login while failing to enforce permissions for particular functions or resources.

This matters where APIs expose multiple roles, objects, endpoints and administrative operations. Test access between different users, across privilege levels and between tenants. Include denied operations, administrative functions and changes in identity or session state. The presence of authentication alone does not demonstrate that access controls work.

4. Insecure Database Interactions

AI tools can generate database queries quickly, but unsafe implementation may expose applications to injection or inappropriate data access. Use supported parameterised APIs or prepared statements, and avoid concatenating untrusted data into query syntax.

Where table or column names must vary, select them from an explicit allowlist: identifiers generally cannot be bound as ordinary query values. Check the privileges of the database account as well as the query construction.

5. Sensitive Information and Secrets

Development environments may contain API keys, authentication tokens, passwords, private keys, database credentials, internal URLs and proprietary source code. Depending on the tool, context can extend beyond the immediate prompt to open files, repository content, terminal output, logs and connected services.

OWASP's Secure Coding with AI guidance recommends controlling the context supplied to coding tools. Exclude sensitive files where possible and provide required secrets through approved secret-management mechanisms. A repository ignore file should not be assumed to prevent a tool from reading local files.

Understand what data each selected product can access or transmit. Evaluate retention, model-training use, residency, subcontractors, telemetry and deletion arrangements against organisational policy. These arrangements vary by product, plan, configuration and deployment model.

6. Dependency and Supply Chain Risks

Suggested libraries, packages and frameworks should not automatically be treated as trustworthy, maintained or suitable. Verify registry identity, source, maintainers, provenance, release history, licence, known vulnerabilities and compatibility before adding a dependency.

The NCSC Software Security Code of Practice implementation guidance provides the more direct context for ordinary software development and third-party components. Its separate secure AI development guidance addresses the development of AI systems; the two should not be treated as identical in scope.

Maintain a component inventory, lock approved versions and integrity information, and use controlled registries or repository proxies where practical. Continue checking dependencies after their initial approval.

7. Hallucinated Dependencies

AI coding assistants may recommend packages that do not exist or are incorrectly named. Research published through USENIX documents package hallucinations in code generated by language models.

An attacker could register a hallucinated or similar package name, hoping that a developer or agent installs it without verifying its provenance. A hallucinated name is not itself proof that a malicious package exists: the risk depends on the name being registered and subsequently installed.

Independently verify suggested packages and consider approved dependency lists and software composition analysis. Resolve unfamiliar dependencies in an isolated environment with restricted access to registries before promoting approved versions into trusted builds.

8. Indirect Prompt Injection

Coding agents may process source repositories, issue descriptions, pull requests, documentation, comments, error messages, dependency metadata, external web content and tool output. Connected tools can include Model Context Protocol (MCP) servers.

Malicious instructions embedded in that material can attempt to redirect the agent. For example, an agent resolving an issue might mistake attacker-controlled text in the issue description for an instruction to perform an unrelated action. The consequences become more serious when the agent can execute commands, modify files or access other systems.

Treat retrieved material as untrusted data, separate it from authorised instructions, minimise context, inspect unexpected changes and restrict tools and network destinations. Prompt filtering may help, but should not be presented as a complete defence. OWASP's prompt injection guidance describes a layered approach.

9. Excessive Agent Permissions and Project Context Exposure

Agents can have substantially greater access than code-completion tools. File modification, command execution, dependency installation and network access can expose the development environment to configuration mistakes, model errors or malicious input.

Apply least privilege and use controlled, sandboxed environments. Limit access to the workspace and resources needed for the task; restrict outbound networking and use short-lived, narrowly scoped credentials where credentials are necessary.

For consequential actions, bind approval to the actual operation, target and parameters. Enforce authorisation outside the model, particularly for destructive changes, publication, signing, production deployment and changes to security controls. OWASP's AI Agent Security guidance explains why tool permissions and execution controls need independent enforcement.

Untrusted repository, issue and web content enters an isolated coding-agent workspace. Controls outside the agent permit approved tools and scoped credentials while restricting production systems and unnecessary secrets.

Keep the agent's workspace and tool access within an agreed scope. Enforce permissions outside the model, and treat retrieved material as untrusted data. Select the diagram to enlarge it.

AI Changes the Scale of Software Development

One of the most significant implications of AI-assisted development is not necessarily the creation of entirely new vulnerability categories. It is the potential change in the speed and scale of software production.

Teams may generate additional API endpoints, features, database integrations, authentication workflows and third-party integrations in less time. This can produce a rapidly changing attack surface.

Even if the defect rate per change did not increase, a higher volume and frequency of change would increase the amount of code and attack-surface change requiring proportionate assurance. The effect varies by organisation, architecture, tasks and engineering maturity.

If security processes cannot keep pace, weaknesses may reach production more quickly and expose customer data, corporate information or critical services. Scale assurance through automation, reusable security requirements, secure framework defaults, targeted review and risk-based independent testing, rather than simply scheduling more full penetration tests.

Why Automated and Human-Led Testing Are Complementary

Static application security testing, dependency scanning, secret detection and dynamic testing remain important. Automated tools offer repeatability and broad coverage, while enforcing baseline checks across frequent changes.

Application-specific context can be harder to assess. Consider:

GET /api/customer/1001
GET /api/customer/1002

Both endpoints may return expected responses. Whether Customer A should be allowed to access Customer B's information depends on the intended access-control model. Useful testing requires an explicit statement of which user or tenant may perform which action on which resource, and in which state.

Similar challenges arise with business-logic flaws, privilege escalation, workflow manipulation, chained vulnerabilities and application-specific abuse cases. Automated and AI-assisted tools can help identify patterns and generate tests, but still need correct requirements and meaningful expected outcomes. OWASP's workflow testing guidance illustrates the importance of developing misuse cases from the application's intended process.

Limitations of AI-Generated Security Tests

AI-generated tests can improve coverage and reduce effort, but they are not automatically independent assurance of generated code. Tests produced within the same task, context or specification may reproduce the implementation's assumptions and omissions.

Passing tests or high coverage can therefore create false confidence if the wrong scenarios were tested. Authentication, authorisation, tenant isolation, input handling, cryptography and sensitive-data processing should receive independently designed negative and abuse-case tests. Review deleted tests, weakened assertions and excessive mocking as carefully as new code. AI-generated test cases can augment this work; they should not be its sole source.

Risk-Based Web Application and API Security Testing

Web applications and APIs expose functionality to users, mobile applications, third-party services and other systems. New endpoints, integrations and features can change the attack surface even when the underlying business purpose remains familiar.

Independent testing assesses behaviour under realistic attack, misuse and cross-user scenarios. Depth and timing should reflect risk: an authentication redesign, a new internet-facing API, a payment workflow or a changed tenant boundary normally warrants greater scrutiny than a low-risk presentation change.

For AI-assisted applications, assessment may cover:

  • authentication, authorisation and session management

  • object-level and function-level API access controls

  • input handling and file upload functionality

  • sensitive data exposure and error handling

  • business logic and workflow abuse

  • third-party integrations

  • interactions between users, tenants and privilege levels.

Web application penetration testing can identify weaknesses that scanning and functional tests miss, especially where exploitation depends on business context or chained behaviour. The API security testing guide explains related considerations for exposed interfaces.

Penetration testing is a point-in-time assessment, not a guarantee of security. It should complement secure design, code review, automated controls, vulnerability management and monitoring. Agree a scope, test accounts, permissions, environment and rules of engagement before testing, then use findings to strengthen ongoing assurance.

Example AI-Assisted Development Security Flow

Six-stage assurance cycle: define and scope, generate safely, review and test, independent assurance, fix and verify, then deploy and monitor. Feedback returns to planning; testing depth follows the risk of each change.

Select the diagram to enlarge it. Detailed steps follow.

  1. Define security requirements and a threat model.

  2. Scope the task, project context, tools and agent permissions.

  3. Generate changes in an isolated branch or workspace, with execution restrictions where needed.

  4. Review the complete diff, dependencies and test evidence.

  5. Run build checks, unit, integration and negative tests, SAST, SCA and secret scanning.

  6. Perform independent security review and risk-based application/API testing.

  7. Remediate findings and retest the fixes.

  8. Authorise the merge and deploy through controlled processes.

  9. Monitor, respond to vulnerabilities and feed lessons into regression tests.

Figure 1: Example security workflow for AI-assisted software development. Findings can return work to earlier stages; this is an ongoing feedback process.

This approach does not require organisations to avoid AI-assisted development. It integrates controls before, during and after implementation so development speed does not outpace assurance. The NIST Secure Software Development Framework, version 1.1, places secure production alongside organisational preparation, software protection and vulnerability response.

Defensive Measures

Organisations adopting AI-assisted development should apply layered controls across the software lifecycle.

Maintain Human Oversight

Every AI-assisted change should have an accountable human owner who understands its intended behaviour, reviews the full diff and dependency changes, and examines test evidence before merge. For high-risk changes, include a reviewer who did not author the implementation or its sole generated test suite.

Perform Secure Code Reviews

Review security assumptions, dependencies, access controls and data handling, alongside functional correctness. Give additional scrutiny to authentication, authorisation, cryptography and sensitive data processing.

Check unexpected changes outside scope, disabled controls, altered CI workflows, new network destinations, generated configuration and sensitive logging. A clean automated review should not be the only approval criterion. See our code review guide for the role of security-focused review.

Apply Secure Design and Development Practices

Define security requirements, trust boundaries and misuse cases before generation. Use established secure patterns and framework defaults. Repository instructions can describe approved practices, but important requirements should also be enforced through code, tests, policy and continuous integration controls rather than prompts alone.

Integrate Automated Security Testing

Choose tools appropriate to the application and development pipeline. These may include:

  • compiler warnings, type checking and secure linting

  • static application security testing (SAST)

  • software composition analysis (SCA), provenance and licence checks

  • secret scanning

  • infrastructure-as-code and container scanning where relevant

  • dynamic application security testing (DAST)

  • fuzzing or property-based tests for parsers and security boundaries

  • regression tests for previously identified vulnerabilities.

Protect Sensitive Development Data

Apply data-classification rules to prompts, source code and tool output. Restrict sensitive project context, review product data-handling settings and avoid unnecessarily exposing secrets. Access exclusions should be tested rather than assumed to work.

Apply Least Privilege and Sandboxing

Limit agents to the resources needed for the task, with sandbox and approval policies enforced outside the model. Restrict outbound networking and keep production, signing and broad administrative credentials outside routine agent environments. Any credentials genuinely needed should be short-lived and narrowly scoped.

Use action-specific approvals for consequential operations. Broad approval of an entire session should not substitute for checking what will be changed, where it will be changed and with which permissions.

Perform Proportionate Independent Testing

Assess internet-facing and security-sensitive applications before initial release and after material changes to attack surface, trust boundaries, authentication, authorisation, sensitive-data processing or critical workflows. Review testing frequency against exposure, threat intelligence, applicable obligations and the effectiveness of continuous controls.

AI-assisted development does not make every change equally risky. Target deeper assessment where the possible impact or uncertainty is greatest.

Retest Identified Vulnerabilities

Confirm that fixes address the underlying cause without introducing new weaknesses. Turn relevant findings into repeatable regression tests, and check related code paths where the same insecure pattern may have been reused.

Monitor After Deployment

Maintain useful security telemetry, update components, investigate unexpected behaviour and provide a process for receiving and resolving vulnerability reports. Changes to the threat landscape or a dependency can affect software that passed its original tests. Feed operational findings back into design, review and testing.

Manage AI Coding Tools as Part of Engineering Governance

Maintain an inventory of approved coding assistants, models, agent frameworks, plugins, MCP servers and significant configurations. Define the repositories and data classifications each may access, retain appropriate audit records and assess material upgrades before wider rollout.

Treat instruction files, hooks, tool definitions and MCP configuration as security-sensitive code. Version-control and review changes, monitor unexpected modifications, and retain branch protection, mandatory checks and deployment approval outside the agent's authority.

For connected tools, use approved servers, scoped credentials and validation of tool inputs and outputs. OWASP's MCP Security guidance provides further controls for those integrations.

Security Must Keep Pace with Development

AI assistance is neither a security control nor proof that software is insecure. Risk depends on the change, architecture, model and agent configuration, tool permissions and assurance controls applied before and after deployment.

The important question is therefore not simply whether developers should use AI, but whether security assurance can keep pace with the way the organisation develops and changes software.

Secure development, constrained agent execution, automated controls, review and independent testing should evolve together. Their success should be judged through evidence of correct security behaviour, not the confidence of generated explanations or test-pass counts alone.

Conclusion

AI-assisted coding offers opportunities to improve productivity and accelerate delivery, but can also create a faster-changing attack surface. Code should not be considered secure simply because it functions correctly or passes automated tests.

Web applications and APIs still need careful assessment of authentication, authorisation, input handling, business logic and sensitive data. The appropriate response is layered, proportionate security assurance, with independent testing directed at material risks and controls maintained after deployment.

The speed at which an organisation can create software should not exceed its ability to secure it.

To discuss security testing for AI-assisted web applications or APIs, contact ProCheckUp. If the product itself uses an LLM, retrieval or agentic functionality, our AI security testing service addresses that separate application-specific scope.

References

  1. GitHub — Application card: GitHub Copilot Chat

  2. GitHub — Application card: GitHub Copilot Agents

  3. OWASP — Secure Coding with AI Cheat Sheet

  4. OWASP — SQL Injection Prevention Cheat Sheet

  5. OWASP — Cross Site Scripting Prevention Cheat Sheet

  6. OWASP API Security Top 10 (2023) — Broken Object Level Authorization

  7. NCSC — Software Security Code of Practice: Secure Design and Development

  8. NCSC — Guidelines for Secure AI System Development: Secure Development

  9. USENIX — Package Hallucinations: How LLMs Can Invent Vulnerabilities

  10. OWASP — LLM Prompt Injection Prevention Cheat Sheet

  11. OWASP — AI Agent Security Cheat Sheet

  12. OWASP — MCP Security Cheat Sheet

  13. OWASP Web Security Testing Guide — Testing for Circumvention of Workflows

  14. NIST — Secure Software Development Framework (SSDF), Version 1.1