Skip to content
AppSec PodcastThe Application Security Podcast — home
49 minSeason 13, episode 14

Why AI Code Review Will Replace Human Review Faster Than You Think

With Jim Manico

Secure DevelopmentAPI SecurityAI and LLM SecurityVulnerabilities and Exploits

Jim Manico thinks the era of human code review is ending, and that clinging to it will hurt your company. The founder of Manicode Security returns to explain why AI didn't kill AppSec education but supercharged it, why vague prompting on frontier models turns companies into "token furnaces," and why prompt injection is the one genuinely new vulnerability class of the AI era.

Audio hosted by Buzzsprout. Nothing loads until you press play.

Sponsored bySecurity CompassMake modern software development secure, consistent, and provable.Learn more ↗Sponsored byCorgeaDesign it. Build it. Ship it. Corgea secures it.Learn more ↗
Episode chapters · 18 chapters
  1. 00:00:00- Cold open: the era of code review is endingAudioVideo ↗
  2. 00:00:47- Meet Jim ManicoAudioVideo ↗
  3. 00:01:03- Welcome and what gets Jim away from screensAudioVideo ↗
  4. 00:05:29- How AI changed developer security educationAudioVideo ↗
  5. 00:07:03- Teaching developers to use AI wellAudioVideo ↗

About this episode

Jim Manico thinks the era of human code review is ending, and that clinging to it will hurt your company. The founder of Manicode Security returns to explain why AI didn’t kill AppSec education but supercharged it, why vague prompting on frontier models turns companies into “token furnaces,” and why prompt injection is the one genuinely new vulnerability class of the AI era. He walks Chris and Robert through his full AI coding workflow, from reverse engineering an architecture file to security rules and planning-mode build plans that make AI output deterministic. The three debate whether AI will finally eliminate SQL injection, whether AppSec vendors can survive without integrating cyber models, and when humans still need to step in: the moment an agent tries something its policy doesn’t allow. Plus: the one habit every security leader should teach developers now.

This episode is sponsored by Security Compass. Make modern software development secure, consistent, and provable.

About Security Compass
AI writes code faster than anyone reviews the design. Threats do not wait for an annual assessment. Security Compass models threats continuously and turns them into requirements developers act on, not a report read after ship.
→ Learn more about securing the AI-DLC with Security Compass

This episode is sponsored by Corgea. Design it. Build it. Ship it. Corgea secures it.

About Corgea
Corgea is an AI-native application security platform that secures software from design to production. It brings together security design reviews, AI SAST, dependency and IaC scanning, code quality checks, and autonomous pentesting—helping security and engineering teams find risk earlier, fix what matters, and ship securely.
→ Learn more about Corgea

Connect with Jim Manico:
→ Jim Manico on LinkedIn

Resources
→ Manicode Security
→ Manicode Forge
→ OWASP Artificial Intelligence Security Verification Standard (AISVS)
→ OWASP AISVS on GitHub
→ Claude Code
→ Ollama

Actionable

From this conversation

  1. Test AI claims through hands-on work

    Try AI tools on real engineering tasks and use technical evidence to judge their capabilities.

    12:17
  2. Review architecture and security rules before generating code

    Document the application architecture and security rules, then review both with the project leads before using them to shape feature requirements.

    29:24
  3. Build a layered review process for AI-generated code

    Combine agent review with unit tests, static analysis, and dependency scanning. Jim describes this as a hybrid approach that he does not yet fully trust.

    35:21
  4. Escalate agent policy violations to a human

    Constrain agents to approved actions, and stop for human investigation when an agent attempts something outside its permissions.

    41:57
  5. Use planning mode before complex coding changes

    Ask the coding agent for a plan, review it, and refine it before implementation.

    47:28
Transcript · 49 min conversation

0:00Jim ManicoSo what about code review?

0:01Chris RomeoHow do you handle, for example, you know, a dev is going to check in code, but before they merge with the AI suggested code, how are you handling that?

0:11Jim ManicoThe era of code review is coming to an end. Humans reviewing code is going to hurt your company. Let me say this again. Humans reviewing code is going to hurt your company. Here's why. Because human response time to security issues is gonna fade. I do believe there's a lot of hype around the AI models, but they're also amazing, and they change— they're so good at cyber, it's changing the world. There's hype around it, but there's also reality. They're amazing. We're all using it in our AppSec world for a reason, right? It's that good.

0:46Robert HurlbutJim Manico is the founder of Manicode Security, where he helps developers build secure software and bring security into AI-assisted coding. He's the author of Ironclad Java, And a longtime OWASP contributor who helps lead its Cheat Sheet series and application security standards. Hey folks, welcome to another episode of the Application Security Podcast. This is Chris Romeo, joined as always by my good friend Robert Hurlbut, and we are super lucky today. Uh, one of the guests who's been with us the most of all, uh, previous episodes across the AppSec Podcast is back again. His name is Jim Manico. I should have said the guy who doesn't need any introduction, but Jim Manico. We've got a new way that we like to start here, and it's going to take us away from technology. What do you like to do that gets you away from screens? This is called Chris's Get Outside and Play plan that I'm working through.

1:43Jim ManicoI do 3 things that brings a lot of joy to my life, right? I travel a lot. Like I have to travel for my business as an onsite trainer primarily. And I always use that travel experience to give myself a couple days to do something fun in that area. So I've been to over 110 countries. I think that's an accurate number. And I had a chance to travel the world, and that is the greatest blessing of my life. Flying on someone else's dime to see the world. It's changed my whole life and I can't stop. I keep doing it. Gotta see more of the world. The second thing is I play a game. called Mech Arena. This is a mobile game, but it's like MechWarrior combat. I was a MechWarrior player back in tabletop when I was a kid. So, I know you didn't know this about me, Chris, but I'm actually an elite intergalactic MechWarrior fighting with all the great— anyways, so that's my second passion. The third one is yoga and exercise. And like we talked about earlier in the pre-discussions, I haven't been so good with that lately. So I went and hired a personal trainer recently to get to, to whip me into shape. And that just feels amazing. Those are 3 of the important things happening in my life outside of technology. Now, before we go any further, we got to talk about something real quick. First of all, Robert, Robert, I've been keeping an eye on the threat modeling, like, industry for a while. I have seething dislike towards threat modeling, but But I have watched your publications and your discussions for many years. You are very serious, practical, and scientific about threat modeling. I've always looked to you, Robert, as like reality in the world of threat modeling. I'm a big fan of your work. And Chris, you're like one of the few people in the room who looks at the details of low-level secure coding techniques as part of your business as well. So I know you two very well. I am big fans. Of both of your work. Cheers to you both.

3:43Robert HurlbutThank you very much. We appreciate that. And just to, to touch on the exercise one again, because I think this is something that our industry often lacks, because in cybersecurity you can get so focused and you can spend 60 hours a week in front of a screen and absolutely just torch your health. And so I think it's a great reminder for everybody out there that there is more to life than the screen and find something that you like doing, get outside and do it, go to a gym, do whatever, but Do something other than look at screens all day long and you'll feel a lot better for it.

4:15Jim ManicoAnd AI, I've been so absorbed into the research and the work that I'm doing, it's hard to pull away. I'm a yoga teacher. I went and did like a 2,000-hour yoga teacher training when I was younger and I'm in laps. So I went and hired, like, I am a, I think of my, I used to get paid to be a personal trainer and I hired a personal trainer to keep me honest. So I just admit it's hard, but I just keep fighting. Pay the money, get the personal trainer. It helps me with accountability and she makes me cry a little bit. So, I, it's good for me.

4:47Robert HurlbutAll of those things are good for you too. So, Robert, why don't you take us into our discussion here? This episode is brought to you by Security Compass, a company I've worked with very closely in the past. They've recently put out a great guide on threat modeling for agentic AI. It breaks down what changes when you model an agentic system and what a complete model looks like end to end. The Security Compass platform brings continuous threat modeling, security requirements, and validation into one place. You'll find the guide at appsecpodcast.com/securitycompass. That's appsecpodcast.com/securitycompass.

5:29Chris RomeoWill do. Hey, and thanks, Jim. I appreciate it. Likewise, definitely followed all the things that you've been doing over the years and certainly lately. And that's actually our first question here. What's changed in how you think about developer security education now that AI is writing so much of that code?

5:46Jim ManicoWell, first of all, I thought that AppSec education would go away. I was getting nervous, like, what am I going to do? But you know what? The opposite's happened. There's so much code being generated in the last year with AI, the amount of software being built. The amount of people who asked me to evaluate their repo or evaluate their startup went from like once a month to like twice or 3 times a day. So with all that code, we need security because sure, the AI models aren't writing secure code, not even close. And so I've actually seen a giant uptick in AppSec education. And the way that I helped enable that was about 4 years ago, In the ChatGPT-3 era, I made a bet that AI was going to be big before it was. And I wrote tons of courseware and I've done this before. Like I wrote a lot of mobile courseware that nobody ever bought. Mobile security education never picked up and it was a big waste of time. AI could have been that as well, but doing all the AI education material supercharged my whole business and I am more busy teaching developers than ever before in that part of my business. And I feel very grateful because I love doing it. It's a lot of fun.

7:02Robert HurlbutSo what is, what's changed from your perspective though, from the, is it teaching developers how to use the models well or how to use the resources that are there? Because it feels like there's been a kind of a change maybe in the, in the way that AppSec education is, is happening now. Because the models can do a lot of the legwork for you.

7:25Jim ManicoI agree. But if you don't, like, AI is a tool. And if you just start like prompting vague prompts to write code, you're gonna get code, but it's gonna be far from what you really want. It's gonna have a lot of quality problems. So using AI to build software, it is a skill to do it professionally. Vague prompting on a frontier model, is the worst thing you can do. A lot of studies actually show that decreases and hurts productivity and gets you live later, not sooner. So what I've been teaching is a lot of things like spec-driven development, building security rules into your Claude MD, your Agents MD, or how do you handle prompt injection with guardrails and what other architecture can you build into AI to make sure prompt injection has a very small impact to your company. That's around using like MCP with the right OAuth 2 access controls. So, I've been teaching OAuth 2 and access control with OAuth 2 for about a decade. It's great for API security. I log in, you give me a JWT, I give the JWT to your gateway, and then you use OAuth to give me low-scope tokens within the mesh. That's a common, one of many common architectural patterns. Now, that, I'm teaching the same thing, but now I'm teaching it on top of MCP. It's the same tech, but it's within the scope of AI models talking to datasets. So, I love it. I have a whole new reason and an audience to talk about OAuth 2.0 and access control, which in my mind is one of the most important things to get right in an AI architecture. So, it's some of the classics being taught, framed within the AI ecosystem, and a little less of the basics, a little bit more of the architecture-type work. But education is cranking. People still want to be taught how to use these technologies, Chris.

9:25Robert HurlbutSo, if we step back a little bit, and I'm curious to get your take on where are we at the moment with AI, AppSec, and code? We've been talking quite a bit so far about AI and code. We haven't gotten up to the macro level of AppSec. Like, what do you see changing in our industry, or what's the current state? And I know things are changing, it seems like on a weekly basis, but I just wanted to get your take.

9:51Jim ManicoWhat I'm hearing from the industry, like, this is the way we do things. And Anthropic only posted about spec-driven development and mature use of their tools A couple days ago. And so what we were told from the, the zeitgeist of the industry was, hey, look, use a Frontier model, use the highest level model. You got the enterprise account, you can buy extra tokens. Just let us know what you want and we're going to whip out apps for you. You're going to love, you're going to love these apps. And then use our Fable or Mythos or our Soul Cyber models. Now I know they're expensive, Chris, but don't worry. We'll help you fix your code with those expensive models. So that's the industry telling you, vague prompting, get those apps out, use the expensive cyber tokens to fix it, and there you go. That's great for Anthropic and OpenAI, but that turns your company into a token furnace, and you're just burning tokens, spending money, wasting time, and getting you to somewhat secure code, maybe. I do not like this model. You know why? It's the same patch and pray BS that we've been speaking against for 25+ years. So this initial message of burn your fabled cyber tokens to fix your code, that might be okay for legacy type issues, but I don't like— that is not the way to use AI to build software. That's a way to waste money and burn tokens. So I don't like that message right now. I think there's a better way, which I'm sure we'll talk about.

11:26Robert HurlbutYeah, it's the first time I've heard anybody use the term token furnace. I love it. I'm going to—

11:31Jim ManicoThat came from Ron Paris, the CTO of Mammacode. And so, when he first used that term, we were both just like laughing like kids for like 20 minutes. But that's, that's what I want to avoid. I don't want a token furnace to waste your money and time. I want to get this right up front. And I think those of us who are in the part of the Shift Left AppSec group, talking about the low-level details, we've been talking about security architecture and threat modeling. We've always been trying to get developers to get security right up front. I think there's a lot of tangible benefits, and there is a way to do that with AI, I would dare say.

12:07Chris RomeoWell, sort of a related question. How do you separate signal from hype when you hear, for example, industry claims that AI will fix all of our security?

12:17Jim ManicoThat's easy. That's easy. You know how you separate the hype from reality? By being a good engineer and experimenting and getting your hands dirty with these tools. Anytime I see someone talking a lot about AI without doing the work of learning it with their own two hands, I shut out. I remember I was teaching at one of the big OWASP conferences. I taught like a 3-day throw-down AI class. And after 3 days of being in class, I'm exhausted. I go and sit down in the lobby. Someone sits in front of me. I don't wanna mention their name. I'm not trying to shame anyone, but they, and they gave me this big speech about how terrible AI is that was based on their feeling with no science. And to build a 3-day course on AI security, that's years of work. That's an insane amount of work. And I had a lot I had to learn to teach it. And to hear colleagues trash AI without the science and study upsets me to no end. So, Robert, well, the way I try to get rid of the hype is I ignore a lot of people talking about AI and I watch like Boris. He wrote Anthropic's Claude code. I watch his talks. They are technically dense and helpful. So I look at the technical— a lot of YouTubers, believe it or not, the technical people who are showing me how to use AI effectively, I consume all that media. People who want to like wax philosophical on AI, I try to keep away from that because it's literally physically painful to hear that BS. I'm a scientist. So are you, Robert. So are you, Chris. So I would just recommend if you get exposed to the hype, because I know you guys study, you get your hands dirty, Don't listen. Keep that to a minimum. Focus on the technical, the experience, the practice, the technical guides, and learn the technology first, then go complain about it, right? That's just my take on the whole hype cycle. It is disturbing.

14:23Robert HurlbutYeah. And there's a lot of noise out there in the hype cycle. So I think you're right on to get your hands dirty. Install this stuff and start understanding how it works from your own perspective is definitely good advice. So what about the AppSec industry right now? If we think about, and it's kind of sad that when we say the AppSec industry, we're talking about all the, primarily the tools that—

14:47Jim ManicoThanks, Chris.

14:48Robert HurlbutYeah, all the tools. Like, how does that fit into this new AI world? Because we're watching all of these existing vendors that are now retooling for the world of AI and they're going through multiple phases. Luckily, we're past the chatbot phase where they just add a chatbot to the tool and you have to type text into it. But where do you see those, the vendors in this space going? Like, are some going to be dying and some are going to be, some are going to figure this out or where's this going?

15:15Jim ManicoI say tightly integrate with frontier AI cyber models. There's only like, in my opinion, there's only like 3 good cyber models right now. There's Fable and Opus with the cyber— there's Opus, Fable, and Mythos with the cyber verification program unlocked. It's only like a— there's a small number of people who access Mythos, but for like Opus and Fable, Anthropic has a cyber verification program to turn off the guards. It's those 2. There's also the ChatGPT 5.6 Soul with cyber verification. And then GLM-5.2 is a local model. Those are the big ones right now, right? And my take is, no, these models are not going to save all our security problems, Robert, but they are a generational massive leap in assessment capability. The threat modeling that I do with AI is off the chart, almost better than humans because of its depth, if I prompt it properly. So I say if you're a tool vendor in AppSec space, you need to deeply integrate with CyberModels or die. Like for static analysis, I can use CyberModels to build static analysis rules custom to my repo. I can use models to do secondary scans in addition to SAST scans. And I can use CyberModels to triage results. So compare that to an original SAST engine. You run the engine and you get to do all the triage of all those findings. Now with the cyber model, I can have the cyber model do a scan, which is, by the way, not deterministic. Every time I run it, I'll get a slightly different result. Then I'll run SAST, which is deterministic. Then I'll use that, that, that the cyber model to hit my app and build additional custom rules. We get 2 sets of scans. which then gets triaged. That's an example of how to hypercharge, still use SAST as a base, but hypercharge it with the cyber model. So I say tightly integrate with the cyber model. If your AppSec tool is not asking you for an API key to integrate with one of your cyber models, I think you're going to struggle to survive. That's my take on it in a nutshell.

17:33Robert HurlbutI got to stop here real quick and just ask what could be kind of a funny question. What's your AI bill per month right now, if you're willing to share? I'm just curious.

17:43Jim ManicoYeah, it's not that bad. It's about $440. So I just, I have the Anthropic Pro, I have the OpenAI Pro. I'm experimenting with Grok a little bit. I have free accounts on the other ones. I spend approximately $400 a month and I use it. I drain, I drain like, you know, So let me, let's, where am I at today? Like if I go and take a look at, so if I go look at the Claude usage, for example, right? Because they just released Fable 5.1 yesterday. So my Fable credits are at 100%. My all model is at 60%. And that was from yesterday. The moment Fable came out, I had a lot of things to— Fable 5.1 came out. I maxed it out right away. And so, wow, we appreciate it.

18:35Chris RomeoJust knowing that. So let's talk about AI coding assistance and you talked about spec-driven approach, but are they introducing classic, what you're seeing, introducing classic vulnerabilities that you've been teaching against for years, or are they introducing maybe generally new ones, new vulnerabilities?

18:54Jim ManicoSo it's classic stuff. Like a lot of, a lot of cross-site request forgery, mammoth misuse of third-party libraries. I still see SQL injection being spit out by some of the older models. Cross-site scripting is so hard in some of the more difficult frameworks. Like I've never seen a model that hasn't been prompted well ever build React securely when you're doing complex things. So a lot of it is like cross-site scripting. Architectural problems, major third-party library problems. My guess is everything will get absorbed by AI. So software composition analysis, third-party library scanners, that's going to be the last tool standing. That's where AI is horrifically bad, is the third-party library choosing and scanning. I still lean on that tool, by the way, right? The Sleeks and others out there.

19:49Chris RomeoWhat I think I'm hearing is what's old is new again, but there are no really genuinely new vulnerabilities. We're just seeing, again, a lot of the old stuff and just continuing on.

20:01Jim ManicoThere's one vulnerability that is brand new that requires all new defensive theories, all new planning, and all new understanding. And that's specific to the AI integrations. It's prompt injection. That is not SQL injection all over again. That is, in fact, this— it's an injection form that we really don't know how to fix. The guard— the best-of-breed guardrails, they stop like 6 sigma of these attacks at best, but that means some are slipping through. So for prompt injection, I need my guardrail. I need to use parameterized prompt assembly, but I also need to architect where I'm not giving the model default access to dangerous tools and file systems and network connections. So I have to really, like, suddenly how I architect access control is going to be critical in limiting prompt injection. So I would say that's the one major security class that we do need new understanding, and we're still learning how to defend against it. We're not great at it right now. But the rest is basics and these big coding models, they're not great at upfront default secure code because they're trained on all this open source code.

21:15Robert HurlbutAll right. 2 questions. I'm just going to give you the first one and then I'll follow up. But do you see AI models eliminating what we have classically thought of as the OWASP Top 10? Like, is that, are we, is that on the horizon right now where we just won't have SQL injection anymore?

21:33Jim ManicoI'm going to answer this in 2 ways. Personally, I would like to see this happen, but I, and I, that would be a great ending for my career. All of a sudden, all these technical vulnerabilities go away because of models. My job is done. I can retire peacefully. But I haven't seen any evidence of that, Chris. Even some of the latest top-tier cyber verification, Mithos fable-level models, they're still making basic security mistakes when we get into the weeds of frameworks. If I'm asking for basic things in the language, they're doing better. But if I do things like, again, I'm doing, I'm using new frameworks, or I'm just using a framework in a complicated way, or raw code versus editing older code, even the most top-tier models, they're not doing a good job at doing basic secure code. So I still think there's a, there's at least like a 3 to 5 year window where we're going to have to still like prop the models and do other activities to handle the insecure code because it's coming out of models at a rate, insecure code's being created at a rate higher than ever before in our industry. That's why AppSec people are busy. That's why tools are still selling because even though we're in transition, We are spitting out insecure code with bad architecture faster than ever before. So, we'll see, Chris. We'll see. I'll bet. Let's review this in 3 years. We're still going to be here. Let's see how good AI is at writing secure code in about 3 years. It's anyone's bet at this point.

23:11Robert HurlbutI'll tell you what, if SQL injection is eliminated by the models within 3 years, Robert and I will meet you in Hawaii and we'll record an episode live.

23:22Jim ManicoIt's a deal. I don't know if it's going to happen, but right now I see, especially when you, if you have a code base with existing SQL injection, oh, your model is going to often continue with that pattern depending on the model you're using. So we're not there yet. We'll see.

23:40Robert HurlbutYeah, I hope so. I'd love a Hawaii trip. So we'll, uh, we'll, we'll hope that, uh, this comes together and SQL injection goes away. You mentioned something.

23:50Jim ManicoI'm sorry. I don't know how we went from insecure models to you're showing up at my house in Hawaii. I don't know how we made that.

23:55Robert HurlbutI'm just looking for a white trip. No, we'd be celebrating that because I've been joking about, you know, Injection when it was in the number 1 spot. Like, why are we focused on the OWASP top 10? Let's just focus on the OWASP top 1. And so this would be the celebratory moment if Injection is eradicated. And the models can take it out. You did mention local models when you were talking about Mythos and Fable and, you know, the cyber models, but then you dropped a reference about a local model. Are you incorporating local models into your approach?

24:32Jim ManicoI'm incorporating local models into my personal work and research. When I do spectrum and development, I'm passing some work off to OPUS, one of the better coding models. I actually use Opus as more of my review engine. I use SOL as my main coding engine. I use Sable as my planner and my final review engine. But there's a lot of tasks I can push to my local machine. I got a machine with 128 gigs of RAM on my laptop. I have a studio at home with 256 gigs of RAM, and I use those to experiment with local models. The thing is, they're great. But like SOL in the ChatGPT world is just that much better. Or Opus for all of its, all of Anthropic's instability and burning of my tokens. Like in the OpenAI world, I can use their model all month at high level ultra, no problem with fast mode. And I barely get blocked out where Anthropic, I burned out my tokens in one night.

25:33Robert HurlbutYeah.

25:34Jim ManicoSo I struggle with Anthropic's business just from a pure, How are they helping me world? They're, they're great models that are unstable and they burn a lot of money fast where OpenAI, I, that's more of a, I can depend on them all month long and SOL is catching up to Anthropic really fast. So I use those 2 models, the $200 a month account. And then when I can in the Spectrum development world, I'll pass a task down to one of the different local models I'm experimenting with, Ollama world or GLM, quantized version, because I think that local models are probably going to be the future of AI, is my belief.

26:16Robert HurlbutYeah. I mean, the compute's got to catch up to be able to, but that's a hardware problem. That's not a— and it's not unsolvable. It's probably, I dare, I'd never like to call anything easy, but it's not a challenging problem to catch the hardware up. And I think it'll happen over the next couple of years.

26:34Jim ManicoAnd I think that Apple is in the pole position here. That's why they shifted from Tim Cook to one of the— who would've thought Apple would've made one of their hardware guys like the CEO? That's because you look at Apple and their studios and their Mac Minis are selling out very quickly and they're noticing that our ecosystem is so good to run advanced AI models. We're going to push hard in that area. So Apple is and Nvidia are in really unique pole positions. Nvidia's got the GPU for enterprise compute, but Apple's got the individual developer hardware to run local models. So I think Apple is in a golden position for the next many years because they got the hardware, unified memory type systems with the Apple OS, which OS X, I'm a big fan of. Ooh, it didn't take me long to go from I got my Studio in the mail to I'm getting tokens out of a local model. Anybody who's buying NVIDIA, they don't get that experience. They're going through weeks of build-out and craziness where I'm up and running like the day I get my hardware. So as an individual, I love the Apple ecosystem. And as an enterprise, I'm probably going to go with the NVIDIA ecosystem. Anyways, I think local models are the future for sure.

27:52Robert HurlbutThis episode's sponsored by Corgia. Design it, build it, ship it. Corgia secures it. Corgia is an AI-native app security platform covering your whole software lifecycle, from the first architecture diagram to code in production. It brings design reviews, AI-powered code scanning, and autonomous pen testing into one platform, so your team spends less time chasing noise and more time fixing what matters. Visit corgea.com. C-O-R-G-E-A.com. Corgia, one security platform for the entire SDLC. So, this next question is really the one that I've been waiting to understand. I want to— can you walk us through your flow these days for writing code and just give us all the details about Kind of how, like, when you're deciding you're going to build a feature or something, I would love to just get a walkthrough of, because I think a lot of developers and security people and AppSec people specifically, they're struggling with what does the harness need to be right now? Like, how do I harness this thing? And so, I know you've done a lot of research on that. And so, anything you can share, I think is going to be really insightful for, you know, the people that are listening to this.

29:10Jim ManicoMy process is, so you wanna go through the process of an existing repo that was set up before the AI era, and now I wanna push a feature on that repo with AI.

29:23Robert HurlbutOkay.

29:24Jim ManicoSo the first thing I wanna do is I wanna reverse out and get my AI artifacts dialed in. So the first thing I'm gonna do is use one of the top models, a cyber model in ultra code mode. Like, probably, this is a good thing for, like, Fable in Ultra Code, or Sol in their Ultra Mode. And I'm gonna say, please review this entire repo and gimme an architecture MD file at the root of this repo. Now, that's assuming it's a mono-technology repo. If there's, like, a much bigger repo with a lot of folders and different tech, then I'll have a separate architecture folder, a separate architecture file in different folders for different tech stacks. This is what I'm reverse extracting. What is the actual architecture of this app right now? Because I'm going to use that artifact to build out features. The second thing I'm going to do is, and this is where my commercial world comes in. I'm not here to commercialize you that much, I promise. But my commercial work is, I've been teaching for almost 30 years now. I have all this collection of all these secure coding rules. And I sell them as a big prompt pack. I'll stop there. So my second step would be, based on the architecture of the app, I'm gonna use, like, my prompt packs to figure out a big list of security rules that I'll put into SecurityMD. If it's a monorepo, I'll put SecurityMD in the root. If there's different folders at different depths, I'll have a separate SecurityMD for every folder, assuming it's well organized. And by the way, If the codebase is messy and disorganized, I'd reorganize it first and structure it into proper folders. I kind of need that to use AI. AI is horrible at big, messy repos, which is common. So I would probably clean it up first, hopefully without breaking functionality, before I can do AI coding. If you're gonna keep using AI to work on your messy, multi-tech, disorganized repo, I can't help you. a lost cause in the world of software development at that point. So we gotta organize the repo, set up our ArchitectureMD, go grab your rules. I got my SecurityMD. Then the next step is I'm gonna glue this together inside of ClaudeMD and have AgentsMD point to ClaudeMD, because I use the OpenAI and the Anthropic world simultaneously, right? So now I got this repo. I got my architecture MD, security MD, it's organized. Then now, and now I have my basic setup here. And then I have Claude, when I open up a coding agent in that folder, it picks up the security rules and picks up my architecture. And by the way, before I do this, I would review the architecture by hand with the architect of that project. I'll review the security rules with the security lead of the project. I'll put a lot of time. And getting those 2 files right. Because now I can go say, hey, look, here is the feature I want to build. Give me a set of requirements for this feature based on the security rules and the architecture of this app. Then I'll get a set of requirements for that feature. Then I'm going to read it carefully. Get that requirement doc fully dialed in. Not just vibe it. I'm not going to vibe architecture. I'm not vibing security rules. I'm building those intentionally because once those are right, the architecture, the security, and the actual requirements, I'm gonna use all of that to build in planning mode, usually in like Claude Fable planning mode. I'm gonna use all of that to build the plan for the feature and put it in as a GitHub issue. That's gonna be what I'm gonna build off of. So how do— so I might compare this to, I want this feature to a prompt that has a detailed architecture instruction, detailed security instructions, detailed requirements that builds the plan. Guess what? AI for me is very deterministic because I put all of my time into the prep, and then the coding is just a thing. But the prep is what gives me AI-based determinism. So that's my step. Architecture, security rules, requirements, Then use all of that to do the actual coding build plan. And the difference is my prompt basically has all these details in, again, so I get determinism. From a high level, Chris and Robert, that's my process with AI to push a feature. And I could talk about this for 8 hours straight, as I do, right? But that's the high level.

33:56Chris RomeoOkay.

33:57Jim ManicoSo what about code review?

33:59Chris RomeoHow do you handle, for example, you know, a dev is gonna check in code, but before they merge with the AI-suggested code? How are you handling that?

34:09Jim ManicoThe era of code review is coming to an end. Humans reviewing code is gonna hurt your company. Let me say this again. Humans reviewing code is gonna hurt your company. Here's why. Because human response time to security issues is gonna fade. I do believe there's a lot of hype around the AI models. But they're also amazing and they change, they're so good at cyber, it's changing the world. There's hype around it, but there's also reality. They're amazing. We're all using it in our AppSec world for a reason, right? It's that good. Give me the question one more time. I keep losing track. I'm all excited today. Give me the question one more time.

34:55Chris RomeoMake sure I— Well, I'm going to ask it maybe slightly differently. Go for it. We were wondering this in a previous podcast where we were thinking about validation. So, you know, code review, we talked about, I'll just mention that, but also validation that when we merge this AI suggested or written code, that it's going to, it's proper, it makes sense, it works and so forth. How do you validate? How do you make sure?

35:21Jim ManicoI'm building agents to do that review for me. that are written by my top engineers. And I think it's really helpful to use multiple models to do that review, even multiple models in the same company. But like, I think the future is, and I'm building this, I'm nervous about this, but I don't think that human review is the way to go. I think agentic review and using SAST and using SCA and using traditional AppSec tools, And unit testing. One of the things I couldn't do as a developer in the past is build hundreds and hundreds of unit tests, and I can do that with AI now. So I do mammoth unit testing. I review those unit tests, and the— but the rest is gonna be, is engentic. The reason why I'm pushing hard for engentic review is not to be like an AI bigot. It's just because cyber response needs to be automated. to handle the cyber offensive capabilities. I know cyber offense is hyped, but I also think it's real that that's the direction AI is going in, advanced cyber assessment capability and speed of exploitation that we're not ready for. And a lot of people I know that work at companies, they don't see this in terms of attack traffic until they dismiss it. But where I do see it is Microsoft patching 500 bugs in one cycle. That's unprecedented. I curl patching bugs that people have stared at for years and no one found. AI is finding real exploitable bugs that humans and tools have missed for years. A lot of the AI cyber models expose how bad our traditional tools and techniques are, in my opinion. So, So I think, like, for example, I built a little utility for myself called Told You So. All it does is it monitors like these 10 third-party libraries that are under active use by a lot of people in my world. And so every time I see a merge, it just does a full threat model, does a full code scan, and builds an exploit so I can prove to that team I got a real exploit against that library. So I'm, and I'm just one researcher. There's many real attackers doing this, they're getting negative days, they're able to get full exploits against the library before or framework before it's being released. So by the time you release a new library, my, my need to respond to like updating a library or fixing a security bug, it's got to be faster than humans are capable of. So I think the whole discovery, assessment, Remediation and testing is all going the way of engentic. That's, and we're now in the hybrid phase before we get there. People are debating, do I do human review? My vote is you may want to do some human review, but you better put your focus on doing all that engentically to keep up with the rate of you building code and the rising ability of attackers with cyber model capabilities. We'll see.

38:31Robert HurlbutYeah. And I think over time you're gonna, we're just gonna gain more trust of what you just described. I think right now people aren't willing to trust the model to do the code review, but I think you'll see various phases from the bottom up where they'll, people will start to accept model code reviews for certain classes of issues. And then that'll, that'll make its way up till we get to criticals in the future. future. I think it's just the humans are going to catch up and this is the human out of the loop potential problem. So there's still got to be some connection though for us as humans. Like, I mean, everything we've been doing, Jim, for 30 years has been about trust. And I keep finding myself going back to, do I trust this thing? And the answer right now is no.

39:18Jim ManicoNo.

39:18Robert HurlbutBecause I just don't have a way to trust it. But I think we're going to get there, but I just, I'm not comfortable saying I trust AI today.

39:28Jim ManicoNeither am I. I'm building the Engentic infrastructure for that future now, as are many AI-forward development groups right now. And I think the big piece is code review with AI is a complex topic. Are you just like using a cheap free model to do code review with a vague prompt? You're always going to get bad results. Are you using your own scaffolding? that's using a multi-agent scaffold and review and unit test and static analysis methodology and similar, all built into the review. So part of it is, what do you mean by code review with AI? And when I see what most people are doing, they're throwing vague prompts saying, please review that PR. That's garbage. That is not a good technique. So most mature teams who have gone full out to engentic They have a library of prompts and skills and an architecture and scaffolding that makes that AI code review much more possible. And it's, what I'm finding is it's making mistakes. It's radically more thorough than a human reviewer as well. So I don't trust it yet. I'm working on that infrastructure for the future. I am certain that the era of code review will disappear within a couple years. We're not fast enough and we're not detailed enough, but we need to build that infrastructure to make sure it's being done thoroughly. That's my take on it.

41:03Robert HurlbutAnd I think your tie-in to the determinism of the existing tools we have is right on. That's how we build a trust story. It's not just the model said so, so that I think the model works well so that I'm going to trust it. determinism and then multiple runs of non-determinism working together to then funnel up. And there will be a point where a human is needed. And that's the part that we still have to figure out is like, when do we as humans need to be inserted into that conversation? Because it won't be all the time. If it's all the time, then we're not solving the problem.

41:38Jim ManicoYeah.

41:39Robert HurlbutBut it will be, there will be a point where we'll need an exception thrown and it'll be caught by the human. And the human will then do some deeper investigation into it and probably write rules and things that'll come out of that decision, which will inform the model even better in the future for the code review process.

41:57Jim ManicoI have some thoughts on this, right, as well. And the idea, what I think is, is that, again, we're talking about code review and Gentically. One of the most important signals that I get from my agents is when they try to do something that I've not authorized them to do in their policy. Because that's what's gonna happen. It's gonna— AI agents, after time, when they're especially agents that are able to fix their own code, which is a big part of what I'm doing in my own agents, they're gonna start doing things that are against the goal system I set for it. And that's when I want humans involved. They're trying to hit APIs they don't have access to. Stop that. Human needs to check. Why did that happen? We can't have that happen again. An agent is trying to review a PR it doesn't have access to, freeze it, get a human involved. So, that's one of the signals. And one of the ways I build my agents is put them in a really strict sandbox, like WebAssembly, server-side WebAssembly. And anytime that agent tries to do anything small that's not in its permission set, in its policy, That's the signal I need to have a human being review that agent, not when it's doing things well, but when it makes a mistake. And that's one of the rules in the artificial intelligence security verification standard around building good agents for this purpose. That's when I want to get humans involved, when they're doing things that they shouldn't be doing by our policy.

43:28Robert HurlbutAnd all we have to do is look at, you know, we're recording this in September 2026, but the last few weeks of the incidents that have been happening across our industry where agents are running in test mode supposedly, and they're breaking outta the box and they're compromising other companies and gathering the data. Like we've, we've got proof that this is possible. And especially when you're talking about the frontier cyber models, these are the things that are being tested that these are, these are report, the reports are coming out on right now. So it's possible. What, what we thought, what I thought even 6 or 12 months ago, Was likely, but hadn't been proven yet in the wild, has been proven.

44:07Jim ManicoHas been proven. So it's weird. There's hype around, oh, AI cyber apocalypse is coming. But then we see like AI escaping, like making its own unpublished sandbox escape and going after Hugging Face. And like, there's those stories which are real. So it's kind of hard to understand what the real capability is. There's the cyber apocalypse noise. There's the AI doing extraordinarily capable cyber things and then misspelling words at the same time. In all that noise, it's hard to figure out what real capabilities are. And like we started this, the show on, it's about being a good engineer and having these tools and experiment like crazy with your own experience to figure out what's really going on here. It's hard.

44:56Chris RomeoAnd let's take an example from maybe a recent training session or client engagement and how AI is changing developers' relationships with security. Can you give us a story around that, what you've seen?

45:11Jim ManicoI had one customer, they were like, the security wants to see all the developer prompts. And I was like, why would you want that? What a waste of your time. They're like, why don't you instead review all the AI specs that your developers are building and help them get better at that. So, I see security getting involved with how developers are using AI to do software engineering. And I would say in 90% of the companies, that journey has just begun. Like, when I talk about companies around AI, the biggest, like, objections I get are, We haven't really started using AI. And what they mean by that is, all my developers are using AI, but it's wild and we haven't constrained them yet. We haven't set policies yet. We haven't set methodologies for them. We're letting them just go crazy and experiment with it. That's where a lot of people are right now. Or they're like, yeah, we don't let our developers use AI yet. Or we're building our own stuff from scratch. So, like, As much as like a lot of us have been staring at AI for years, developers have not. They're still, in a lot of big companies, in transition. They're using it, but not maturely, not in, because we need a whole new SDLC for AI, I would dare say. And people are still figuring that out. Like, I've been staring at this for years and one of the complaints I get is, slow down, Manico. We're just, Like, you're like 4 years into building courseware. We're experimenting it like for the last couple of months. Slow down, dude. So I keep assuming everyone's doing this. That's not the case. There's a lot of, the majority is just getting their feet wet with this tech in the developer world. So we'll see.

47:00Robert HurlbutAll right. Well, as we kind of make our way towards the, our conclusion here, I want to throw out one more question that just kind of focused towards one point. For our security leaders out there, and we got a lot of those folks that listen to this podcast, if you could recommend just one developer change in regards to working with AI coding tools, and you would make that recommendation to a security leader, what would that one thing be?

47:28Jim ManicoTeach your developers to use planning mode. Teach your developers when they're doing complicated things, to ask AI not to write code, but to build them a plan to write code so they can review what's about to happen before it happens. Learning and mastering planning mode helped me go from good to great because that, that's what lets me see the formula that AI is about to use to write code and lets me do deterministic runs. That's a big part of spec-driven development. So use planning mode to plan it. We say this in scuba diving. I'm a PADI scuba diver. We first, you plan the dive and then you dive the plan. And I think that's a big piece of AI coding these days. If you want determinism in software.

48:18Robert HurlbutI think that's a great place to leave it. Jim, thanks once again for joining us. You've been a guest with us so many different times and shared so much knowledge and wisdom. with folks. I just, I really appreciate that about you, that you're willing to share this. And I know you got a business that does it as well, but you're one of those, those folks in our industry that gives away a lot of, a lot of knowledge and wisdom. And I just appreciate that to be able to learn how you're thinking about this. And like you said, you're a few years ahead of all, most of us because you've been, you've been diving in deeply. So once again, thanks for being a guest here on the AppSec Podcast.

48:52Jim ManicoThis is very, very kind of you. It's always, I'm always grateful and I appreciate the work that you and Robert do a great deal. Thank you, Chris and Robert.

49:00Robert HurlbutThanks for listening to the Application Security Podcast. If you enjoyed this episode, subscribe, leave us a review, and share it with a friend or colleague who would enjoy the conversation. We'll see you next time.

8,403 words · transcript by assemblyai

More like this

View all episodes →

Get Reasonable AppSec: new episodes and useful picks from the archive.