By Aparna Ash Himmatramka
A developer using an AI assistant can ship a feature in an afternoon. In many organizations, the security review for that feature can still take a week.
Believe it or not, most AppSec programs are still built this way. BSIMM16, which studied 111 organizations, found security feature review in 80.2% of them. Those reviews made sense when teams shipped quarterly. Today they’re the bottleneck. What this means is that teams only review their crown jewels, and everything else goes unexamined.
Having worked in the AppSec field for more than 8 years now, I see this problem repeated over and over again. Until I realized it’s not a technical problem, it’s a gap of how we measure success for these programs.
The usual response is to make reviews faster: better triage, more AI in the review queue, one more hire. It helps for a while, but it doesn’t fix the real problem. Any program that needs a human decision for every change will grow with engineering output, and engineering output is now growing faster than any security team can hire.
So the goal shouldn’t be faster reviews. It should be needing fewer of them.
The model: humans set policy, the environment enforces it
In this model, no security engineer approves or blocks a deployment. Instead, the security team builds the systems that make those decisions automatically. We need a couple of pillars established for this model to run smoothly:
- Secure defaults. Templates, libraries, and infrastructure modules make the safe choice the easy choice, so fewer vulnerabilities get written in the first place.
- Automatic risk scoring. Every change gets a risk score based on what it touches and how critical the asset is. Low-risk changes move fast. High-risk changes get a deeper look.
- AI-filtered findings. Models grounded in your own codebase remove false positives and suggest fixes that look like your developers wrote them.
- Assurance scores instead of approvals. Every service has a health score that the deployment pipeline reads directly. Healthy services ship freely. Unhealthy services face limits proportional to their risk.
- Feedback loops. Recurring findings become new secure defaults. When a security engineer makes a call on a new issue, that call becomes a rule, so nobody has to make it twice.
The last one is what makes everything else work. It’s surprising how many programs fix the same bug class over and over across different teams. A program that learns eliminates the class.
What this looks like in practice
Every piece of this model has already worked at scale somewhere. What’s new is putting them together, with AI connecting them.
Netflix replaced reviews with a paved road. Their AppSec team concluded that per-app security assessments didn’t scale anymore and invested in secure-by-default frameworks and self-service guidance instead. One example is Wall-E, a gateway that handles authentication for web applications. The team rewrote their checklists to push developers toward Wall-E as a pre-approved default. When a team couldn’t use it, the security team extended Wall-E instead of granting an exception. By 2021, Wall-E fronted over 350 applications, and its adoption rate was on every leader’s security risk dashboard. Notice that accountability came from a metric, not an approval.
Google shows what happens when you design a vulnerability class out instead of catching it earlier. This isn’t shift left. Shift left moves detection closer to the developer, but someone still has to find and fix every bug. Google moved responsibility for safety away from the developer and into its languages, libraries, and frameworks. Before 2012, frontends like Gmail had a few dozen XSS vulnerabilities a year. After refactoring onto safe-by-design frameworks, that dropped to near zero. This month, an internal Google AI agent found over 500 XSS vulnerabilities across Google’s broader web applications, and only two across the hundreds built on its hardened frameworks.
And this isn’t just for tech companies. Capital One, a heavily regulated bank, built its own policy-as-code guardrails in 2016 to monitor and automatically fix cloud configurations against enterprise policy.
In none of these cases did the review get faster. The need for it went away.
Three lessons for security leaders
1. Accountability doesn’t require approval. The first question every executive asks is: if nobody signs off, who is accountable? The answer is consequences tied to risk. A service that stays unhealthy loses privileges that would increase its exposure, like frequent deploys, new dependencies, or autoscaling. At planning time, security health sits right next to reliability, and teams below the threshold get capacity reserved for recovery. When the constraint comes through engineering planning, it’s no longer a negotiation with the security team.
2. Measure what the program no longer has to do. Counting completed reviews rewards the wrong thing. The single most useful metric is the human touch rate: the percentage of changes that need any security engineer involvement. It should go down every quarter. Pair it with time to
remediate critical findings and the false positive rate. A falling touch rate is also how security coverage grows without headcount growing at the same pace.
3. Plan for the AI being wrong. The worst case is a false positive that blocks a team who knows the finding isn’t real. Therefore, give developers a fast way to escalate. The disputed finding should stop counting against them immediately, a security engineer should review it within hours, and every decision should go back into the model. Also, be honest about what automation can’t do. It doesn’t reliably catch business logic flaws, design complex authorization, or decide how much risk the business should accept. That’s where your senior security talent belongs, and this model gives them the time for it.
Where to start
Don’t start with AI. Models need your organization’s data to be accurate, and that data comes from the basics: consistent scanning, an inventory of services with criticality and data classification, and simple rule-based risk scoring. Make assurance scores visible before you enforce them. Add machine learning once you have several months of history. Expect about a year before the program starts improving on its own.
The takeaway
AppSec used to be measured by how many reviews it did. At AI speed, the better measure is how few reviews the organization needs while risk goes down and delivery speeds up. Make the secure path the easiest path, and let the environment do the enforcing.
About the Author
Aparna Ash Himmatramka leads an application security engineering team at Amazon.com. She will present this model at Hacker Halted 2026. The views expressed are her own.

