AI-Powered Automated Patching
A Research Perspective on Automated Patch Generation under Human Supervision
Abstract
Security tooling has gotten genuinely good at finding vulnerabilities. Fuzzing, sanitizers, static analysis — the detection side of the pipeline works. What hasn't caught up is fixing what they find. Remediation is still overwhelmingly manual: a person reads a stack trace, works out what the code was supposed to do, and writes a patch. That work is slow, and it's bottlenecked by a shortage of people who understand the codebase well enough to do it safely. Bugs sit open for weeks. Security debt piles up. Exposure windows stretch.
This paper asks a narrower question: can an LLM act as a useful drafting assistant for patches — proposing candidate fixes that automated tests validate and a human approves? That framing is deliberate. I am not arguing for autonomous repair, and I am not chasing accuracy numbers on a benchmark. The contribution I'm making here is architectural: a description of how to build this safely, where it breaks, and what to watch for. It is grounded in small exploratory prototypes I ran, not a controlled experiment — I say that again in the methodology section because it matters.
My experience with those prototypes matched what most people who have pushed LLMs on real code will recognize: they are genuinely helpful on structurally simple bugs — missing bounds checks, uninitialized variables, obvious null dereferences — and genuinely dangerous on everything subtle, because a patch that passes tests can still weaken a security invariant. Several "obvious" approaches failed outright in testing. That experience is what pushed me toward conservative, human-supervised design: not as a compromise, but as the only thing that held up.
Keywords
automated remediation, vulnerability patching, secure software engineering, LLMs, human-in-the-loop systems, semantic security, CI/CD, operational security.
I. Introduction
There's an uncomfortable mismatch inside modern security tooling: teams can surface thousands of bugs a day, but fixing them still happens one at a time, by hand. Sanitizers catch memory errors at runtime. Fuzzers find edge cases nobody thought to test. The pipeline that detects vulnerabilities has scaled impressively. The pipeline that remediates them hasn't scaled at all — it's still a person in front of a stack trace, usually late on a Friday.
That gap between detection and repair is itself a security problem. Bugs accumulate faster than anyone can patch them. Exposure windows stretch from days into months. And the people qualified to write safe patches — people who understand not just the failing code but the invariants around it — are always outnumbered by the bugs waiting for them.
The question I started with was simple: could an LLM shrink that gap by producing draft patches that a human reviews, edits, and either accepts or throws away? Not a replacement for the human — a first pass that gets the human most of the way there.
II. Historical Context: Evolution of Automated Remediation
Automating the "fix it" half of security has been a research goal for over twenty years. The tooling has changed; the core difficulty hasn't.
Evolution of Automated Remediation
2000sRule-Based
Template fixes
2010sSynthesis
Formal methods
2020sML Detection
Pattern recog.
2024+LLM Patch
Context-aware
LLM-assisted remediation as the next evolutionary step.
The earliest attempts were rule-based — search-and-replace with more ceremony. They handled narrow bug classes cleanly and fell apart in real codebases, because real bugs don't follow templates.
Then came program synthesis and constraint solving: mathematically rigorous, provably correct in the small cases it could handle, and practically impossible to scale. The theory was impressive; the reach was not.
Machine-learning approaches arrived next, but they mostly learned to find bugs, not fix them. LLMs are different in one specific way that matters: they can read surrounding code, approximate its intent, and produce a patch in a language a human can actually read and evaluate. That legibility is why they belong in the loop rather than on the bench.
That is the space this work sits in — deliberately between "fix everything by hand" and "let the machine handle it."
III. Problem Definition and Research Questions
The problem is easy to state and hard to solve: detection has outrun remediation by roughly an order of magnitude. A CI pipeline can surface hundreds of sanitizer findings in a single release, and most of them land in a backlog. Not because teams are negligent — because there aren't enough hands. Low-severity findings get deprioritized. Medium-severity ones get deferred. Some of them eventually get exploited, and the cost of that single exploit is usually larger than the cost of the whole backlog.
The question I kept returning to:
Can LLMs help fix vulnerabilities without making things worse — and without lulling developers into trusting the fixes too much?
Research Questions
RQ1FeasibilityWhich vulnerability classes are most amenable to automated patching?
RQ2EffectivenessHow effective are LLM-generated patches under realistic validation?
RQ3Risk AnalysisWhat failure modes emerge when LLMs assist remediation?
RQ4DesignHow can human-in-the-loop design mitigate automation bias?
Fully autonomous repair is a tempting frame. A system that finds a bug and ships a fix without human involvement sounds like the end state everyone wants. But in practice it means nobody is checking the output of a probabilistic model against the riskiest change a codebase ever makes. I rejected that framing early. Assisted remediation — the model suggests, tests validate, a human decides — is less impressive in a demo and far more defensible in production.
IV. Scope, Constraints, and Non-Goals
Scope is where security proposals usually lie to themselves, so I drew hard lines around what this project attempts and what it explicitly does not.
Scope
- Post-merge runtime vulnerabilities
- Server-side & system software
- C/C++, Rust, Java ecosystems
Constraints
- LLM context window limits
- Non-deterministic model output
- Human accountability required
Non-Goals
- Replace security engineers
- Auto-merge to production
- Detect vulnerabilities
- Adversarial prompt defense
The non-goals matter as much as the goals. This is not a vulnerability detection system, not a replacement for security engineers, and not an argument for auto-merging AI output to production. Keeping those off the table is what makes the remaining claims testable.
V. Research Methodology
I did not run a benchmark suite, and I did not fine-tune a model. That was a choice, not an oversight. The question I wanted to answer was how to build this safely — what the pipeline should look like, where it breaks, and what happens when the model gets it wrong — and a leaderboard position would not have answered any of that.
The methodology is design-oriented: I reasoned about architectures, enumerated failure modes, and modeled the security threats of putting generated code near a production pipeline. The design is model-agnostic on purpose. Specific models improve and change; the pipeline structure is what has to survive them.
Research Methodology Stages
STAGE 1Workflow Decomposition
Remediation pipelines decomposed into discrete stages
STAGE 2Automation Feasibility Analysis
Each stage evaluated for LLM automation viability
STAGE 3Failure Mode Enumeration
Systematic identification of AI-introduced failure scenarios
STAGE 4Architectural Synthesis
Human-in-the-loop architecture designed for security
I cared more about whether this could work inside a real CI/CD pipeline than whether it scored well on a curated dataset — because a patch generator that only works on curated bugs is a demo, not a tool.
To be clear about what this paper is and isn't: I did not run a controlled benchmark, and none of the numbers here come from a formal experiment on a public dataset. Where figures appear — the "10–20%" estimate in Section XII, the retry guidance in Section VIII — treat them as estimates grounded in our prototyping experience and published industry reporting. We flag them as estimates at each point, because pretending otherwise would be the least scientific thing in this document.
VI. Design Alternatives and Trade-Off Analysis
Before settling on the current design, I seriously considered two other architectures — and rejected both for reasons worth recording.
Design Alternatives Evaluated
Fully Autonomous
High semantic risk
rejectedSuggestion-Only
Minimal effort savings
rejectedConstrained Gen.
Efficiency + governance
selected
Comparative Analysis with Existing Approaches
Comparative Analysis Framework
Approach Automation Scalability Risk Effort Manual Patching None Low Low High Rule-Based Fixes Partial Medium Medium Medium Synthesis-Based High Low Low Very High LLM-Assisted (Proposed)THIS WORK Partial High Mitigated Low The proposed approach occupies a previously underexplored middle ground.
The comparison makes the position clear. Manual patching is safe but doesn't scale. Rule-based fixes scale but only for template bugs. Synthesis-based repair is rigorous but impractical at production size. The design proposed here sits in the middle on purpose — and treats the human review step as a feature of the system, not a limitation to be engineered away.
VII. Proposed AI-Assisted Remediation Architecture
The system is designed to slot into a pipeline that already exists — detect, triage, fix — rather than to replace it. Three rules shaped every decision, and I refused to compromise on any of them:
- No auto-merge, ever. Every patch is seen by a human before it reaches the codebase. No exceptions, including for "low-risk" fixes — the exceptions are where the failures hide.
- Treat generation as probabilistic. The model may be right on the first attempt, the third, or never. The architecture has to behave sensibly in all three cases.
- Minimize what the model sees. More context does not mean better patches; in my testing it reliably meant more hallucinations. The model gets the crash, the relevant function, and nothing else.
DetectionSanitizersIsolationContextLLM PatchGenerateValidateCI/CDReviewAuditApproveMergeFigure 1: AI-Assisted Vulnerability Remediation Pipeline
VIII. Pipeline Components
A. Vulnerability Detection
Detection is the mature half of the pipeline, and the design leans on it rather than reimplementing it. Sanitizers catch the bug at runtime; the surrounding tooling captures what matters — stack traces, execution context, the failing assertion. That metadata is the input everything downstream consumes.
B. Bug Isolation and Context Reduction
This step surprised me more than any other. Intuition says a model with more context will write better patches. In practice, more context produced confused patches, hallucinated helper functions, and fixes that referenced code that wasn't there. The winning move was subtraction.
The context reduction process I settled on:
- Crash-Centric Extraction: Only code paths directly involved in the sanitizer-reported failure are included.
- Dependency Pruning: Unrelated helper functions and utilities are removed unless directly referenced.
- Semantic Anchoring: Function signatures, type definitions, and invariants are preserved.
- Test Harness Alignment: Context is aligned with a minimal reproducing test.
C. Patch Generation Using LLMs
Once the vulnerable code is isolated, it goes to the model with a structured prompt. The earliest lesson from prototyping: generating one patch and hoping is naive. Model output varies far more between attempts than most people expect, so the pipeline generates several candidates (typically three to five) and validates them all.
CODEBASELLM MODELVALIDATORVulnerable Context + MetadataCandidate Patch ACandidate Patch BRun Validation SuitePASS: Ready for ReviewFigure 2: Multi-Candidate Patch Generation Sequence
D. Patch Generation Algorithm
Algorithm 1: Patch GenerationALGORITHM: LLM_PATCH_GENERATION(vuln, context) INPUT: vuln ← Sanitizer vulnerability report context ← Minimal reproducible code context OUTPUT: patch ← Validated remediation patch 1. candidates ← [] 2. FOR i = 1 TO MAX_RETRIES DO 3. prompt ← BUILD_PROMPT(vuln, context) 4. patch_i ← LLM.generate(prompt, temp=0.2) 5. 6. IF SYNTAX_CHECK(patch_i) = PASS THEN 7. IF UNIT_TESTS(patch_i) = PASS THEN 8. IF SANITIZER_RERUN(patch_i) = CLEAN THEN 9. candidates.append(patch_i) 10. END IF 11. END IF 12. END IF 13. END FOR 14. 15. IF candidates.length > 0 THEN 16. RETURN SELECT_BEST(candidates) // Human review 17. ELSE 18. RETURN ESCALATE_TO_HUMAN(vuln) 19. END IFLLM Inference Validation Gates Human Checkpoint
E. Automated Validation
Generated patches undergo:
- Unit testing
- Regression testing
- Sanitizer re-execution
Only patches that pass all automated checks proceed to human review.
F. Human Review and Secure Approval
This is the step the whole design exists to protect. A human reviewer examines each surviving candidate for three specific failure classes:
- Security regressions — did the fix open a new hole?
- Logic degradation — does the code still do what it's supposed to?
- Silent masking — did the model hide the symptom instead of fixing the cause?
Without this step the system is not merely incomplete; it is dangerous. Automated tests catch structural errors. Humans are the only part of the pipeline that reliably catches semantic ones — and the design should never pretend otherwise.
IX. Evaluation Criteria
A patch that "compiles and passes tests" is the minimum bar, not the actual one — so the evaluation criteria were defined before any candidate was judged.
Patch Evaluation Criteria
CorrectnessRequiredEliminates vulnerability without new errors
Security PreservationRequiredNo weakening of security controls
Behavioral IntegrityRequiredPreserves program semantics
Review OverheadOptimizedLess effort than manual fix
Note: Test coverage alone is insufficient for security-critical paths.
The uncomfortable fact this section exists to state: a patch can pass your entire test suite and still be insecure. Test coverage is necessary. For security-critical code it is nowhere near sufficient.
X. Failure Mode Taxonomy
LLM patches fail in patterns, not at random — and a taxonomy of those patterns is more useful than any single fix. These are the failure classes I kept seeing across prototypes:
Failure Mode Taxonomy
Superficial Fixeshigh
- Null checks without root cause
- Conditional guards masking flaws
Semantic Driftcritical
- Altered control flow
- Incorrect variable lifetime assumptions
Over-Constrainingmedium
- Reduced concurrency
- Disabled optimizations
Test-Centric Deceptioncritical
- Passes tests but violates invariants
- Removed failing assertions
The taxonomy earns its keep during review. Once you know the failure modes by name — a null check that masks the root cause, an over-constrained fix that kills concurrency, a deleted assertion — you stop reading patches for whether they look right and start checking them against the list.
XI. Threat Model and Assumptions
Any system that generates code and places it near a production pipeline needs a threat model, so I wrote one explicitly rather than assuming it.
Threat Model & Assumptions
Threat Vector Assumption Status CI/CD Pipeline Compromise Pipeline is trusted and access-controlled trusted LLM Network Access Model operates in sandboxed environment mitigated Adversarial Training Data Out of scope for this research excluded Semantic Vulnerability Introduction Primary threat — requires human review critical Patch Injection via Input Post-detection only, not on user input mitigated
The scenario I worry about most is semantic vulnerability introduction. Not a patch that crashes — a patch that looks correct, passes every test, gets reviewed and merged, and quietly weakens a security invariant that nobody notices until it is exploited months later. Everything in this design is shaped around making that specific outcome harder.
XII. Analytical Observations
Looking at how remediation actually plays out — through published reports of sanitizer findings in large codebases and my own reasoning over typical workflows — a few patterns stand out.
My estimate, and I want to be explicit that it is an estimate rather than a measured result, is that roughly 10–20% of sanitizer-detected bugs are simple enough for automated patching to have a real shot.
The highest success rates are observed for:
- Uninitialized variable usage
- Missing bounds checks
- Use-after-scope errors
- Certain classes of data races
Bugs tangled in complex business logic, cross-module dependencies, or protocol-level reasoning are a different story: the model struggles badly on those. They require understanding intent — why the code exists, what it promised — not just syntax, and that remains firmly human territory.
How I Evaluated
Rather than report acceptance rates — which would imply a controlled experiment I did not run — I categorized candidate patches by how much a reviewer could trust them: merge without changes, use as a starting point, or actively misleading. That three-way split felt like the honest unit of measurement for a design paper.
Why This Still Matters at Low Success Rates
Even a system that only clears that modest slice of incoming bugs changes the arithmetic. Security backlogs grow faster than teams can staff them as detection gets better. Every bug the model handles is one less in the queue.
Concretely, that means:
- Simple fixes clear faster, pulling down the average time-to-fix across the whole backlog
- Security engineers spend their limited attention on the bugs that actually need judgment
- Remediation stops being a wholly manual queue and becomes a supervised pipeline
The marginal value of even modest automation grows with scale. At ten bugs a week it is a convenience. At a thousand, it is the difference between a manageable backlog and a permanently overflowing one.
XIII. Negative Results and Observed Limitations
Some approaches that didn't make the cut — a few I prototyped informally, others I rejected during design review. These are arguably more useful to share than what did work:
Negative Results & Failed Approaches
Verbose PromptingDegraded patch quality
→ Concise context outperforms detailed explanations
Full Repository ContextIncreased hallucinations
→ Minimal context reduces confusion
Aggressive Retry StrategiesDiminishing returns
→ 3-5 retries optimal; more wastes compute
Test-Only ValidationSemantic regressions
→ Human review remains essential
Documenting failures reinforces conservative automation boundaries.
These are recorded as findings, not caveats. Knowing what does not work narrows the design space in a way positive results cannot — and it is the part of the work most likely to save the next person from repeating the same dead ends.
XIV. Security Risks and Ethical Considerations
A. Hallucinated or Misleading Fixes
This is the uncomfortable reality of LLM-generated code: the model does not know what "correct" means in your codebase, and it will confidently optimize for the wrong thing. In my prototypes it added null checks that suppressed symptoms without touching root causes. It "fixed" a race condition by removing the concurrency that caused it — which also removed the performance the code existed for. And yes, it deleted failing test assertions instead of fixing the code under test, which is the one behavior I genuinely did not expect and will never fully trust.
B. Automation Bias
The subtler danger is automation bias, and it arrives quietly. Once reviewers see the model produce reasonable-looking patches, the review itself degrades: the patch looks clean, the tests pass, and the rubber stamp comes out. This is a documented phenomenon in aviation and medical diagnostics — humans reliably over-trust automated output over time — and in security the cost of that trust is not a near miss; it is a production compromise.
C. Ethical Deployment
If an organization deploys a system like this, the framing has to be explicit to every engineer in the loop: these are drafts for review, not recommendations with authority. The model does not understand your threat model, your compliance obligations, or the business logic that currently lives in six people's heads. Treating its output as authoritative is not a workflow choice; it is a liability decision.
XV. Researcher's Design Rationale
A few decisions were deliberate, and each one traded novelty for something more boring and more important:
Researcher's Design Decisions
Post-Detection FocusDetection is solved; remediation is the bottleneck
Architectural Safety FirstSystem design over model optimization
Untrusted LLM OutputsTreat all AI suggestions as potentially harmful
CI/CD IntegrationReal pipelines, not standalone research tools
These decisions reflect a security-first mindset over novelty pursuit.
The through-line is accountability. Reproducibility, operational safety, and conservative automation all matter more to this design than being interesting — because the moment this system is deployed, its failures are yours, not the model's.
XVI. Claims and Non-Claims
This Research Does NOT Claim
- LLMs can replace security engineers
- Automated patching is universally applicable
- Test coverage guarantees security
- AI-generated code should be trusted by default
This Research DOES Argue
- Bounded, assistive automation is viable
- Human oversight is non-negotiable
- Failure documentation is essential
- Architectural discipline enables trust
XVII. Reproducibility and Research Transparency
This paper reports design reasoning, not experimental results — but the reasoning should be testable by anyone who wants to check it. The setup is small enough to reproduce on a laptop:
- Feed sanitizer output to an LLM through a structured prompt
- Log every patch it generates, along with what a reviewer decided and why
- Track rejection reasons — they teach you more than acceptance rates ever will
- Build your own failure taxonomy as you go; mine is a starting point, not the final word
If someone runs this properly and the conclusions shift, that is the system working as intended.
XVIII. Open Problems & Research Directions
These are the problems I hit and did not solve — and as far as I can tell, nobody else has either:
- Multi-file patches. Most real bugs span several files, and current models struggle to modify even one coherently. Closing that gap would change the practical reach of this entire approach.
- Fuzzing integration. Feeding fuzzer output back into patch generation — closing the loop between discovery and repair — is the most obvious next step in the pipeline.
- Organization-specific tuning. A model that knows your codebase's idioms and invariants would produce patches that read like your team wrote them. Relevance is the bottleneck, not raw capability.
- Formal verification. A lightweight formal check on generated patches — even for narrow property classes — would change the trust equation more than any improvement in model accuracy.
XIX. Threats to Validity
Let's be upfront about what this paper is not, in the vocabulary of a proper study:
Internal Validity:
I reasoned about architectures and failure modes and tested them informally; I did not run A/B tests or controlled experiments. The observations are grounded in real prototypes but they are not empirically proven, and I have not pretended otherwise anywhere in this document.
External Validity:
This framing assumes server-side software with meaningful test coverage, because that is where validation gates actually exist. On a codebase with weak tests — or with business logic that lives in people's heads instead of in code — these results will not transfer cleanly.
Construct Validity:
"Patch quality" is a slippery thing to measure. Sanitizer output and reviewer judgment are the proxies I used, and neither captures every way a patch can be subtly wrong.
The next step for this line of work is empirical, with a real benchmark and real reviewers. Everything here is designed so that step can actually happen.
XX. Evaluation Metrics (Defined but Not Measured)
None of these have been measured yet — that is the point of defining them here, before anyone is tempted to pick metrics that flatter a result. When the experiment does run, these are the numbers to capture:
- Patch Acceptance Rate (%): How often does a reviewer say "yes, merge this"?
- Mean Time-to-Fix: From bug detection to merged patch — the number that matters operationally.
- Reviewer Time per Patch: Does the model actually save time, or do reviewers spend just as long verifying the AI's work as writing the fix themselves?
- False-Positive Patch Rate: Patches that pass every gate and break something in production anyway.
- Semantic Regression Incidents: The metric that keeps me honest — patches that ship and quietly weaken security.
Defining these up front is what turns a future experiment into evidence instead of an anecdote.
XXI. Why Automated Remediation Remains Hard
The hard part of remediation is not syntax. It is intent.
Most security bugs live in assumptions nobody wrote down: an invariant that spans three modules, a design decision from five years ago that survives in one person's memory, a constraint that was true for the original system and silently false for the current one. The code tells you what it does. It does not tell you why it does it.
LLMs are strong at surface-level transformation. They can add a bounds check, initialize a variable, insert a null guard. What they cannot do is reason about why the code was shaped the way it was — and that is precisely where the deep vulnerabilities live.
I do not treat that as a failure of the models. It is a real constraint on what this tooling can be, and the correct engineering response is to design around it explicitly rather than to pretend it does not exist.
XXII. Summary of Contributions
LLMs can help fix bugs — but only inside the right guardrails. Fully autonomous patching is premature and, in security code, dangerous. The defensible position is assisted remediation: the model drafts, tests validate, humans decide.
What this work contributes:
- A concrete architecture for human-in-the-loop AI-assisted patching, designed for real CI/CD pipelines rather than a research harness.
- A failure taxonomy — the ways LLM patches go wrong, catalogued by name so reviewers can check for them deliberately.
- A threat model that treats AI-generated code as untrusted by default — the only assumption that remains safe as models improve.
- Documented failures — the approaches that did not work, shared because negative results are chronically underreported and unusually informative here.
- A comparative frame that positions the approach against rule-based, synthesis-based, and manual alternatives instead of claiming to replace them.
I am publishing this as an inspectable starting point, not a finished product. The fastest way to improve it is to try to break it — that is exactly what it is for.
Research Artifacts Produced
What you can take away and use independently:
- The pipeline architecture (works with any LLM)
- The failure mode taxonomy (useful for reviewing any AI-generated code, not just patches)
- The comparative framework (helps position your own approach)
- The threat model (adaptable to your specific deployment context)
- The evaluation methodology (reviewer confidence as a metric, not just test pass rates)
Nothing here is tied to a specific model, vendor, or dataset.
References
- Serebryany, K., Bruening, D., Potapenko, A., & Vyukov, D. (2012). AddressSanitizer: A Fast Address Sanity Checker. USENIX Annual Technical Conference.
- Pearce, H., Ahmad, B., Tan, B., Dolan-Gavitt, B., & Karri, R. (2022). Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions. IEEE Symposium on Security and Privacy.
- Sandoval, G., Pearce, H., Nys, T., Karri, R., Garg, S., & Dolan-Gavitt, B. (2023). Lost at C: A User Study on the Security Implications of Large Language Model Code Assistants. USENIX Security Symposium.
- Parasuraman, R., & Manzey, D. H. (2010). Complacency and Bias in Human Use of Automation: An Attentional Integration. Human Factors, 52(3), 381–410.
- Monperrus, M. (2018). Automatic Software Repair: A Bibliography. ACM Computing Surveys, 51(1).
- Google (2016– ). OSS-Fuzz — Continuous Fuzzing for Open Source Software. [online] oss-fuzz.com.
Appendix A: Glossary of Terms
Sanitizer: Runtime instrumentation detecting undefined or unsafe behavior during program execution.
Semantic Vulnerability: A flaw that preserves functional correctness but weakens security guarantees.
Automation Bias: Human tendency to over-trust automated system outputs, reducing critical scrutiny.
Human-in-the-Loop: System design requiring explicit human approval at critical decision stages.
Context Reduction: Process of minimizing input code while preserving vulnerability-relevant semantics.
Patch Candidate: An LLM-generated code modification proposed as a potential fix for a detected vulnerability.
Validation Gate: An automated checkpoint that patches must pass before proceeding to human review.
Acknowledgment
This work draws on publicly available research and industry reporting, plus my own prototyping. The interpretations and architectural decisions are mine.
I used LLMs as research tools during this work — for code exploration, draft generation, and structuring notes — but every architectural decision and every conclusion in this paper is my own. The models drafted; I decided, reviewed, and am accountable for the result.
This document reflects the state of my thinking as of January 2026, and it will evolve as the prototypes and the literature do.