--- title: "Brad Geesaman - Redefining AppSec with AI: Shrinking Toil, Expanding Impact - How LLMs are able to reduce toil in triage-heavy AppSec workflows" url: https://appsecpodcast.com/brad-geesaman-redefining-appsec-with-ai-shrinking-toil-expanding-impact-how-llms-are-able-to-reduce-toil-in-triage-heavy-appsec-workflows/ date: 2025-10-28 duration_seconds: 2539 season: 12 episode: 20 guests: ["Brad Geesaman"] topics: ["AI and LLM Security", "Vulnerabilities and Exploits"] audio: https://www.buzzsprout.com/1730684/episodes/18065472-brad-geesaman-redefining-appsec-with-ai-shrinking-toil-expanding-impact-how-llms-are-able-to-reduce-toil-in-triage-heavy-appsec-workflows.mp3 video: https://www.youtube.com/watch?v=S65QBk1-tcM transcript: true --- # Brad Geesaman - Redefining AppSec with AI: Shrinking Toil, Expanding Impact - How LLMs are able to reduce toil in triage-heavy AppSec workflows *October 28, 2025 · 42 min · Season 12, episode 20* with [Brad Geesaman](https://appsecpodcast.com/guests/brad-geesaman/) on [AI and LLM Security](https://appsecpodcast.com/topics/ai-security/), [Vulnerabilities and Exploits](https://appsecpodcast.com/topics/vulnerabilities/) [Audio](https://www.buzzsprout.com/1730684/episodes/18065472-brad-geesaman-redefining-appsec-with-ai-shrinking-toil-expanding-impact-how-llms-are-able-to-reduce-toil-in-triage-heavy-appsec-workflows.mp3) · [Video](https://www.youtube.com/watch?v=S65QBk1-tcM) ## Show notes AppSec teams are drowning in repetitive triage while the work that requires judgment keeps piling up. Brad Geesaman, Principal Security Engineer at Ghost Security, explains how large language models can shrink that toil without handing security decisions to an unreliable black box. He walks through using LLMs for classification, evidence gathering, and contextual analysis, with humans retaining final authority. Brad and Chris examine prompt engineering, trust, market disruption, and the limits of incumbent tools built around producing ever-larger finding queues. They also explore AI-assisted remediation, code drift, and the changing day-to-day work of AppSec engineers. The result is a pragmatic model for gaining leverage from AI while preserving the expertise, accountability, and skepticism that effective security still demands. The Application Security Podcast is brought to you by [Security Journey](https://www.securityjourney.com/). About Security Journey We provide application security training for not just your developers, but for all roles in your SDLC. → [Learn more about Security Journey](https://www.securityjourney.com/) Connect with Brad Geesaman: → [Brad Geesaman on LinkedIn](https://www.linkedin.com/in/bradgeesaman/) → [Ghost Security Reaper](https://github.com/ghostsecurity/reaper) Mentioned in this episode: → [Ghost Security](https://ghostsecurity.com) → [Reaper](https://github.com/ghostsecurity/reaper) → [Security Compass](https://www.securitycompass.com) → [OWASP ZAP](https://www.zaproxy.org) → [Burp Suite Professional](https://portswigger.net/burp/pro) → [SQL Slammer](https://en.wikipedia.org/wiki/SQL_Slammer) → [Code Red](https://en.wikipedia.org/wiki/Code_Red_\(computer_worm\)) → [Nimda](https://en.wikipedia.org/wiki/Nimda) → [Exodus Communications](https://en.wikipedia.org/wiki/Exodus_Communications) → [NetWitness](https://en.wikipedia.org/wiki/Netwitness) Chapters: 00:00 Meet Brad Geesaman 03:01 What toil means in AppSec 05:20 Why triage drains security teams 06:13 Where AI can create leverage 09:29 Does an LLM need custom training? 11:51 Prompt engineering for useful results 13:33 Humans remain at the center 15:23 Trusting probabilistic systems 19:30 A seismic shift in AppSec tooling 20:18 Escaping the pile of findings 24:00 How incumbent vendors are responding 25:46 Why platform shifts leave openings 28:31 The AppSec engineer's changing day 33:04 Moving from triage to code changes 35:36 AI-generated code and application drift 38:49 What Brad hopes comes next 41:51 Closing thoughts ## Transcript *7,600 words · assemblyai* **0:00 Chris Romeo:** Brad Giesemann, Principal Security Engineer at Ghost, joins us today to explore how AI and large language models are transforming the world of application security. We start by discussing the concept of toil, the repetitive, exhausting work that drains AppSec teams as they struggle to keep up with mountains of security findings and alerts. Brad shares his insights on how LLMs can provide meaningful leverage by handling the heavy lifting of triage, classification, and evidence gathering while keeping humans firmly in the loop for final decisions. **0:30 Brad Geesaman:** The Application Security Podcast is brought to you by Security Journey. **0:35 Robert Hurlbut:** We provide application security training for not just your developers, but for all roles in your SDLC. **0:40 Brad Geesaman:** Learn more at securityjourney.com. **0:41 Robert Hurlbut:** Hey folks, welcome to another episode of the Application Security Podcast. This is Chris Romeo. I am flying solo for this interview today. I am a VP at Security Compass and a general partner at Curve Ventures and have been recording the podcast for what seems like roughly 55 years, but maybe not quite that long. We've been at it for almost 10, Robert and I. So, excited to be joined by Brad today. Brad, we'd like to jump right into your security origin story. Our audience is always on the edge of their seat. So, how'd you get into cybersecurity/application security? **1:30 Brad Geesaman:** Yeah, thanks, Chris. Thanks for having me on. I appreciate it. Security origin story. Back in the SaaSer Slammer days, I worked at Symantec MSS. I was a SOC engineer fighting fires those evenings, those late-night shifts. I did a little bit of network and application security pen testing. Did a little bit of sales engineering, selling firewalls and IDSs for a while. Did some consulting, built capture the flag exercises professionally for 4 years, did the whole cloud shift, moved into Kubernetes security, did a few stints there, and now I'm focused on application security at Ghost. So principal security engineer here, diving my, you know, dipping my toes a little bit into the AI space when it comes to code security. **2:17 Robert Hurlbut:** Very cool. I was working incident response Slammer was kind of the tail end of my— I was doing incident response for Code Red, Nimda, all of those fun things at Exodus Communications where we had hundreds of customers that none of them were patched. And those were, I would like to say they were good times, but I don't think I would go that far. They were challenging times. **2:38 Brad Geesaman:** Yeah, that's, I will say that you cut your teeth on that and you don't know what's what until sort of in hindsight. But that network, there's a lot of NetWitness RSA folks that came out of that. So, meet Uran and love Uran and group, that's sort of the foundational network. That group is still pretty tight. So, all's not lost, right? Fighting those fires. **3:01 Robert Hurlbut:** Yeah, very, very definitely true. So, let's dive in here. I've heard people use this word toil. I've heard a couple different people use this in the context of AppSec. I'd love to get your perspective on, like, is this a problem today when we're thinking about the average AppSec person who's working in maybe a medium or large company, probably on a team that's too small? What is this idea of toil and how does it play out? **3:34 Brad Geesaman:** Yeah, we hear that a lot as different things. It sort of comes in bits and pieces, but if you start hearing stories after stories from talking with AppSec teams, toil, at least in my definition, definition. I like to try to put it succinctly as repetitive work without proper leverage. And what we mean by leverage, obviously the repetitive work, it's doing the same thing over and over again without making a dent, like bailing out a boat that's filling up with water with a thimble or a Solo cup. Just more water keeps pouring in no matter what you do, no matter how fast you go, you don't feel like you're making a dent. And I put that as like, it's sort of part of the job and I have a lot of respect for people who do this. in face of that, like they know that they're outnumbered, maybe 100 to 1 or 300 to 1 in terms of dev to AppSec. But it's not just developers, it's the amount of code that they're responsible for. And the fact that you're not necessarily directly able to make changes to the code to improve its posture, you're getting others or you're convincing others. There's like an extrovert battery that you're having to put, you know, not only triage everything with tools that you're getting hundreds of alerts on, Saying, which 5 do I care about? But then you're spending all this energy convincing other folks inside your organization that are incentivized on shipping features to fix those things instead. And you just, you could just see how the battery drains and it leads obviously to things like burnout over time. But that's what I mean by like toil. Like you don't step back and go, you know what? This week was better than last week. Almost every week that comes by, it's, well, there's 5,000 more alerts or there were You know, 50 more CVEs and just not able to keep up. And that's like, I just think of that as the culmination of toil. **5:20 Robert Hurlbut:** Yeah. I mean, I see and understand exactly what you're describing now because it seems like it's a common problem that a lot of folks are dealing with in that every week there's another pile of issues that are coming in and they just feel like we're We're never going to get anywhere. We're never going to be able to get to the point where we reach some level ground. Like, it's— I like your example of the boat. I mean, you're bailing out the water and the hole's getting bigger every week, the hole, or every minute the hole's getting bigger. **5:53 Brad Geesaman:** Exactly. **5:53 Robert Hurlbut:** And the boat's slowly sinking and you're moving faster, but you can only move so fast as a human being. So yeah, definitely understand where you're coming from. A lot of folks are looking at AI and LLMs as the answer to a lot of different things. **6:13 Brad Geesaman:** I'm— **6:13 Robert Hurlbut:** I do dare— I dare say everything. It seems like everything, like AI and LLM, is being positioned. It's where we're seeing all the startup funding in the last year or two. You know, if you don't have AI attached to your pitch, people aren't even looking at what you're doing. It's like it's It seems to be the most important thing. So how does this, how do these things influence, or how are AppSec teams embracing these concepts of AI and LLM today and what you're seeing in the field? **6:42 Brad Geesaman:** Yeah, to your point, AI to get funding, but I think that there's reality, there's quality within the hype that gets drowned out. And if you sort of peel it back, there's a couple really, really strong strengths that align with AppSec problems that LLMs can do that can help with that leverage equation. And it's not everything. You just can't slap an LLM on it, solve it. There's a lot of care and thought to it. But if you break down workflows and you think piece by piece and you try to put the LLM as the smallest possible component in that workflow, you can achieve decent results with high consistency for that thing. If you just say, here, LLM, look at my code base or something, you're going to get what you expect. But if you are very, very careful and you do a lot of context engineering, which is to say, what are you going to provide to the LLM and what is it supposed to look for? And you kind of put baby in a corner, you can achieve some very interesting results at scale. And so the things that it's really good at that I think AppSec engineers maybe don't think about is coverage. You know, most folks have a T-shaped skillset. You know, I know broadly security, I know deep into AppSec, I know maybe 1 or 3, up to 3 programming languages, but do you know every programming language? Do you know every framework and the nuances of a framework with that language? terms of security vulnerabilities? Do you know every vulnerability class? I don't. I'm not going to say I do. Nobody here does. And so if you have a baseline amount of understanding, but you can deepen that and augment and help your assessment of things like, what's this code doing? That's a, that's a thing it can do very well. Like, what's going on here? Like, help me make sense of this. What about the input? Help me trace this through. Like, help me understand the code. Like you had an expert developer next to you that can make you that much more efficient and that much more confident. Is it going to solve everything? No. Is it going to reason about everything correctly? **8:51 Chris Romeo:** No. **8:51 Brad Geesaman:** You still have to be there. You have to guide it through. But if you can break your workflows down to classification steps, annotation, summarization, that's a lot of what triage is. So, you know, what is, what is this? Give me 3 or 4 labels that are meaningful. And now all of a sudden my dataset I can sort on those elements and go, well, here's my top 10 just by that classification. Now I'm working with a smaller dataset. That right there is huge leverage. And so, doing that across languages, frameworks, and across all vulnerability types is like, that's asking the impossible of the AppSec engineer. So, quite frankly, I don't know how we're going to do it without a little bit of help. **9:29 Robert Hurlbut:** So, when you think about using LLMs to assist the AppSec process. This is something that requires customized training of the LLM, though, right? This isn't something that LLMs can do just out of the box today. **9:53 Brad Geesaman:** The frontier models can. I mean, smaller local models, maybe just not as high as quality or consistency, but the ability to say, hey, let's look at these 4 files. These are, they're related this way, this way, and this way. This calls this, this calls this. And present that to an LLM, a strong LLM, and say, tell me about this behavior. And it's not so much like find the vulnerability. It is, I'm looking for, or you are looking for the vulnerability. It has to have these 7 criteria. **10:24 Robert Hurlbut:** Mm-hmm. **10:25 Brad Geesaman:** So point to me and show me that this is the success, like this is user-supplied input. It is not properly validated. It's passed over here. I'm talking about a SQL injection, of course. It's passed in here. It's directly string concatenated, there's no extra ORM support, et cetera. That is a very strong or a high confidence indication of a SQL injection. If you just go, show me the SQL injection, it's going to go, sure, and find you 50 probables. But if you have specific criteria that you're looking for, it can really help shrink the time for you to go, yep, I agree, as opposed to, well, let me start. Where, where am I looking in this code base? It knows how to find where those pieces are. knows how to track down where that import came from and present you the, you know, the slice of context that it's meaningful for that specific code. And you can go, yeah, that's, that's, that's correct. I can validate that. So it's not necessarily going to do all of the work, but it can do a lot of the legwork to present to you, like, maybe a junior AppSec person to a senior AppSec person. The junior is like, I think this, I think I've got something. And senior would say, well, show me the evidence. And the junior shows them the 5 pieces of evidence. The senior's like, you're absolutely right. That's, that's a thing. And so, is it great at 50 files? Not necessarily, but most application security vulns are within 10 files or so. And as long as it's not huge, huge, huge, huge files, and you just jam that in, it can reason about this goes to here, it goes to here. **11:51 Robert Hurlbut:** So there's some significant prompt engineering then. that goes into setting it up. Because like you said, if I just ask the LLM, is there SQL injection in these 10 files, it's gonna then define SQL injection according to its own data and everything. **12:12 Chris Romeo:** Yes. **12:13 Robert Hurlbut:** And so it sounds like the model that you're prescribing here is prompt engineering to prepare the model to really know what to look for in those, the various things that, you know, the 5 things that are going to cause you to say, this is really a SQL injection. And so, it sounds like that's, that's really what you're describing though, is it's not just code going to the LLM. **12:36 Brad Geesaman:** Yeah, absolutely. They're like, we actually try to minimize it in our work. We try to minimize the use to put it to a specific use case. So, we do a lot of upfront legwork, and I know other folks do this as well. They try to provide, here's all the context. Here are some examples. Here's for this framework or this language. Here's things that you're looking for, patterns that we know about in this, you know, like ORM for SQL injection, for example. You know, there's typical things here. You're looking for things like this, plus some other items. You have to have these criteria. It's not just jam the repo into an LLM and go find me the good stuff, because you're going to get what you're going to get. **13:13 Robert Hurlbut:** So. **13:14 Brad Geesaman:** Yeah. Yeah. So that's not to say you can't do that for snippets. I mean, you can take a file and go, tell me what's going on here, and it won't run you astray. But multi-file is like what most apps are. So that's where, you know, tracing from source to sync and things like that, you know, the LLM can really assist with. **13:33 Robert Hurlbut:** So it sounds like you're prescribing more of an augmentation kind of model of where the human's still at the center of the decision, which I happen to be a firm believer in this philosophy at this stage. The human's at the center of the decision, but the LLM is able to, like you said, summarize, pinpoint, even demonstrate flow inside of a file so that then you can, you can, you could do all of the things that the LLM's doing as an AppSec engineer sitting in the chair. But it might take you 30 minutes to gather, trace, stop, think, trace something. Why the heck does it do this? Let me go figure out where it's calling that. It goes to this weird file over here. Turns out it doesn't do anything bad. It's just not maybe the best way that the application could have been laid out across multiple files. So, was that a fair assessment of kind of where you are? **14:32 Brad Geesaman:** 100%. 100% human in the loop. There's, there's never a, like a, I say never, capital N, never. Uh, you're, you're, you're going to have the best results when a human is able to be presented with all of the evidence and they believe it, right? They're able to validate that. So that takes the 98% of the legwork and the energy depletion. How many of these, how many, you know, cold issues can you run down and triage all by yourself? versus how many can you review in that same amount of time? It's a 10x order of magnitude. You go, yes, yes, yes. All right, let's go. Let's go have a conversation. How are we gonna best fix that strategically? As opposed to, well, gee, what does Java do here? I don't know. Is that importing right? Like, you're just down in the muck. You're not making forward progress. **15:23 Robert Hurlbut:** So let's deal with what I think of as one of the the biggest issues when we talk about integrating AI and LLMs, especially into our security world, and that is trust. So you spent a lot of time digging into this stuff, far more than I have. But as a, you know, almost 30-year cybersecurity person, I'm looking at all of these AI, LLM places where we're using them. **15:56 Brad Geesaman:** Yeah. **15:56 Robert Hurlbut:** and saying, but how much do I trust this at the end of the day? So where do you, how do you handle the issue of trust when we start to use these new computing technologies that frankly are only a couple of years old at this point? Like we don't really have a long track record with any of this stuff and it's working, it's going at light speed. Like all you gotta do is wait 6 months and it's twice as good as it was, you know? Yeah. **16:23 Brad Geesaman:** I mean, I, I, you're, you're right. I've been working very closely with it and I've built up a level of trust. There's certainly not absolute trust or blind trust. And when I think trust, I think transparency and evidence. Can't have it without both in this case. Right. So I'm telling you what I'm doing. If I'm, if I'm personifying a system that has an LLM in it, I'm saying, I'm looking for these things. This is what I found. things, this is what I found. And if I suggest that there's a finding, then there's, there's actually interesting characteristics of having another LLM validate that the other LLM's generation was matching up. It's using sort of strength to weakness to feed into a strength. Uh, so something can generate something and, and be close to correct. And the other can look at it and say, did you do 5 things? Did you find the user input? Did you find the concat? Did you you find the lack of validation? They know, then that's not an issue and put that as evidence. But the final output is still up for the human to say, I agree. We're just serving that context up on a silver platter. We're trying to make it so that it's like the best human-led code review, a security review, like the best write-up you could ever possibly have and go, there's nothing to argue with here. I looked at the code. I double-checked all the work. And I didn't have to exert a ton of battery to be able to do that. We're trying to apply that methodology to as much of the triage process, as much of the prioritization process of surfacing quality findings, right? I think that's just general in AppSec, apply it to anything, dynamic validation, static source code, looking up secrets, dependency validation. Like, there's a lot of that, that muscle, that legwork that has to be done. And I think there's a lot of you know, low-hanging fruit there if you do that well with transparency and evidence to back it up. Because ultimately, as the AppSec team, or we find with AppSec teams, they're in a silo from the devs. I mean, there's hundreds of dev teams and companies and small AppSec teams that are dedicated, and they're fundamentally communicating or leveraging their political capital to go get somebody else to stop doing features and fix their thing. And so there's a trust element. that you're, you know, you're slicing off a little bit of your trust and you're putting it down on the table with every one of those findings. For you to trust the finding, you have to validate it so that you can broker your trust to the devs and not look like a silly fool when they go, hey, that's not an issue. At least it stands on its merit of like, well, these things are objectively true. We could have a discussion of whether we fix it or not, but nobody's looking at that and going, what are you talking about? That's not even an issue. works, trying to avoid that. So transparency and evidence is how we build that over time. It's not something that's earned like immediately. It's gotta be like repetitive. This is definitely showing value. That's a bug I would not have found. We've had those sort of eureka moments and like, I never would've thought of that. And it's actually legitimate. I can exploit it. When you have enough of that and you have diminishing amounts of false positives, then that's how you start to build that trust. **19:30 Robert Hurlbut:** So, I know that AI is revolutionizing many different sectors of technology, but it seems like we're poised, if we think narrowly down to specifically AppSec, like the whole thing's about to get turned over here. And what I mean by that is, what I'm going to now call classic tools, which are scan and dump the 10,000 results and things, are either going to have to adapt to become, to use LLMs so that they don't operate in the way that we got so used to them operating. And I'm talking about everything, right? SaaS, DaaS, IaaS, everything. **20:17 Chris Romeo:** Absolutely. **20:18 Robert Hurlbut:** in the classic sense has been all about let's just dump a large pile of stuff that a human has to deal with. And so the more I'm listening to you talk, the more I'm thinking about we're on the verge of a complete shift and incumbent companies with older technologies are either going to refactor that whole thing in such a way that they're able to be competitive and take out the grunt work of operating, dealing with the findings, or we're going to see a massive shift to startups and other newer companies that are— I mean, I never thought of it. I didn't come up with this. I'm sure somebody's already said there's got to be like an AI-first principle, almost. Like startups that were formed with AI at the core are going to have a unique advantage over those incumbents, as they've always had. We've always had incumbents. There's always been something new. This isn't a new challenge, but it seems like it just hit me listening to you talk. Like, We're on the verge of that right now, if it hasn't even, I don't think it's happened yet. I don't think we've reached the tipping point yet, but we gotta be close to that tipping point. **21:25 Brad Geesaman:** We are definitely close. I've had the privilege of talking with companies of all sizes and all maturities, and I see similar things about the amount of effort, the amount of time, the amount of triage, false positive rates, whatever, you know, whatever their pain is, it's pretty consistent. And I think you nailed it, that there's sort of a reimagining of the approach of maybe the prior approach where there's an add-on or an augmentation of AI versus a, let's rethink this AI-first or AI-native, as it were. And there's pros and cons to each, but I think over time, as models get better, the AI-first approach may be able to find things that are more deeper and more behavioral-based much more effectively. But there's a, you know, there's cost benefit of, if you can find or define a pattern for something, that is absolutely the fastest way to find it. That's not going to find the entire scope. And so there could be a hybrid approach at some point that might make sense. But fundamentally, like, let's talk about the shift. Like, what changed? It's, I have to know something in order to find it is the current approach. **22:41 Robert Hurlbut:** Mm-hmm. **22:42 Brad Geesaman:** The AI native first is, let me break the code down so I understand its data flows and then discover things I didn't know before. I didn't have to know a pattern to be able to find that thing in that code. And when you flip that, you think that's actually more akin to a pen tester or a code reviewer, white box, uh, where they're tracing data flows and they're looking for, okay, is that doing the, okay. Oh wow. That might be a thing. Let me go exploit that. validate that in my data versus hold on, let me write a pattern to find that and then look around. And so there's, like I said, there's pros and cons to both approaches, but that deep analysis, those behavioral things like business logic are actually able to be discovered. It's kind of crazy if you give it specific types of examples and specific criteria for those types of examples, you can surface those without having to know the exact thing that you're looking for precisely to make a pattern for. You're just like, I'm looking for this behavior. It looks like this, this, this. It does not look like this, this, or this. And surface me those potentials and run it down. Find me all the evidence. Show me where the data flows. Present that to me on a platter and you will find things. And that's sort of like the— that was our eureka moment when we started to like have success doing that a while ago. Oh, okay. **24:00 Robert Hurlbut:** Without naming any names, what are you seeing in the general AppSec market from those incumbents? Are people moving? My guess, I haven't really done any research into this, but I have a hypothesis that they're not moving as fast as the AI-native people and they're lagging and it's causing some market disruption only because that's been the pattern for the last— 40 or 50 years of new technology disruption. It, it's the, the incumbents always seem to be too slow to catch up, but I'm specifically in, you know, curious. **24:36 Chris Romeo:** Yeah. **24:37 Robert Hurlbut:** You're, you're looking at the AppSec market specifically, like what are you seeing out there? **24:41 Brad Geesaman:** Yeah, it, it's exactly what you're saying here is that, you know, folks that have an existing investment, an existing engine, an existing approach, the natural thing to do is say, you know, some of my outputs, I could do a little bit more. I could go hunt for more context, do more lookups, more validation. sort of put that onto the end of a pipeline and get improved results, natural progression. And then to your point about AI native that are looking for like, how do we find things that we didn't know about? How do we find pockets of finding classes that are just not covered by that approach? And so you have that sort of difference in viewpoint or starting point, we would say, but it's going to converge. There's going to be a hybrid of both. over time and whether that's an acquisition augmentation or a combination, I would say that that's probably going to be the most long-term effective approach for those incumbent vendors is to sort of take a little bit of time and go AI native and then blend it into your approach or, you know, M&A your way to get that capability. **25:46 Robert Hurlbut:** Yeah, I don't think they're going to do it just because it's, they've never, it's never really happened. I mean, you even think with cloud, like AWS and Azure and those types of things. Yeah, they came out of existing companies, but they didn't come out of technology incumbents, like especially on the AWS side. But you could say Microsoft was a— but they do things in such a separate way that it was almost like a new company was formed to do it. But it's Yeah, just, I'm just thinking about the comparisons with cloud and the companies that were, you know, you think about IBM used to own the personal computing market. They used to own the kind of consulting side of it. And then there was this market shift and they didn't move and new people came in. And then we have cloud providers. And yeah, does IBM do cloud? I think so. They used to. I don't know if they do anymore. You never hear about it, right? Because they're not a, they didn't, They didn't move fast enough because nobody ever does. You can't move an existing business fast enough to keep up with a market shift like this. Yeah, I think that's— **27:01 Brad Geesaman:** go ahead. **27:02 Robert Hurlbut:** No, I was going to say, I think that's where we are. I think we're in the midst of that right now. You're living it, so you know. **27:07 Brad Geesaman:** There's not enough incentive to take the big enough risk if you're in a large organization. There's too many things that you're going against in order to try that out. the risk of failure is higher. Whereas if you're in a startup, you know, your existence is predicated on building on successes. And so you either live or die by your capability. And if you were able to prove that to the market and bring that to the market, you can be successful. But like, who's going to go and say, yeah, you know, I'm working at this large company with this wonderful product, but let me kind of rebuild it from scratch and prove that it's not useful anymore. Like, how does that go over? **27:44 Robert Hurlbut:** Yeah. **27:45 Brad Geesaman:** Internally and all the marketing and all the— **27:47 Robert Hurlbut:** Yeah. Even though they should know better at this point, the powers that be, whoever they are, should, but nobody ever looks at history. Everybody thinks it's not going to happen to us in that way, like it's happened 7 other times in market shifts in the last 30 or 40 years. **28:04 Brad Geesaman:** I think we have about 6 to 12 months before convergence of consolidation, convergence, combination of approaches. I definitely think it's gonna be accelerated. It's not like there's a 2 or 3-year lead of any kind. You said the models get better every day. That is true. And being able to harness them as they get better is a skill, and it's not unlearnable. It's definitely learnable if you put your mind to it. **28:31 Robert Hurlbut:** Yeah. Well, let's change gears a little bit and talk about the day in the life of the AppSec engineer. And then specifically, maybe some common pitfalls, because a lot of folks that listen to the show are— they're in the deep end, the AppSec deep end. They are the people we were talking about at the beginning. They're toiling, they're struggling with too many findings and not enough hours in the day to move the needle. So, you know, what is— how does that— those folks, how does their life change by introducing LLM? **29:05 Brad Geesaman:** I like to, I like to sort of start by empathizing and saying, you know, what are, what are they dealing with? Like, get in their shoes. And I look at tools all the time that we compete against and run against our own code as a comparison. And you're sort of looking at this never-ending mountain. Once you get to the top of one mountain, there's just a taller peak of findings from multiple sources of variable quality, variable supporting evidence and transparency. And then you have like bug bounty reports. If you have a bug bounty program and maybe you have a pen test or 3 a year on a couple of your apps that get these kind of funds and you're, you're struggling to put it all together in a meaningful package. And that's what ASPMs typically are, are trying to help you with. Maybe you have one, maybe you don't. So you're looking at your ASPM console or your spreadsheet or what have you. And you're probably just sad. That's probably what most of your day is. And you're like, Hmm. How can I best make a dent in this bucket of sadness? And I like to think of like elevation of maturity of that motion of like, how do we get the time you're spending in a week to be leveraged work as opposed to undifferentiated or unleveraged work? So how can we minimize the amount of time it takes you to get the value out of the tool so that you can Up level, go up another level of altitude and go, wow, okay. If I have the ability to scan my repositories at some cadence that gives me recent results across all of them that are good, transparent with evidence that I can, you know, look at review and go, yeah, that's a good thing. And I can see the playing field. What a world. It's like fighting a fire, you know, like by hand with hatchets and you don't have a helicopter telling you where the wind's blowing and where it's going. Like, You're just down there in the forest. But when you have an LLM or you have that capability in the workflow, it's kind of like you have the helicopter and the plane and the plane just dumps the retardant all over and snuffs out that part of the fire because you have that visibility. So what the world should look like if these things are successful, these processes, is that you're looking at 50 to 100 as opposed to 5,000 that you care about. Yeah. That has business context, that has impact. You know where it's deployed. You can either go validate that by hand or with tools or Zap or Burp or, you know, other agentic proxies that we have open source. And you can say, all right, I'm confident in these 10. And now I have a remediation plan that you can focus your week on. So instead of thinking like, how do I make 50 go to 5 or 5,000 go to 50? Like, okay, I have my 50. How do I make the most effective use of the new machine? Because in theory, and then some folks that adopt this have in practice, I have 50 BOLAs in 12 applications. How do I get away from, or uplevel even one step higher and say, how do I squash BOLAs for good? Maybe there's a middleware library, or maybe there's another capability, or there's something that is unique to your coding style that you can have a conversation with the developers and say, folks, we have 50 BOLAs. How are we going to squash this out? I don't want to keep coming back to you every day, every time I see it. You don't want me coming. Like, how do we just design this out of the system and squash that with a leveraged fix as opposed to just doing it? And over time, you get the trust, you get the teamwork, the silos break down because, hey, whenever the security team comes to talk to me, it's actually a really good And they make it really easy for me to fix or communicate it well. That's where the time is spent, not fighting the tools, trying to get value out. It's driving communication, building trust, comms, strategic planning, strategic paydown, leverage fixes. That's where we want to shift it to the tenable, to the capable, to the status where, you know, it's level and you get one and then you go, oh, we're back up. That wasn't too bad. We're in a good spot. We know where our risk lies. **33:04 Robert Hurlbut:** And so far we've been talking about kind of more of a triage side of the value proposition of LLMs, but you did touch on code change in the last answer here. So now I'm curious on where you stand on LLMs creating fixes specifically for bugs. **33:28 Brad Geesaman:** So there's 2 schools of thought. Well, there's probably more than 2 schools, but 2 primary schools of thought, we'll say. There's the let the LLM fix everything, push ticket cannon of PRs. That will end as well as you think it will. However, I think there's a more kind way to allow the conversation to be a pull, and the code fix that's generated is mostly a suggestion. or even perhaps a prompt to the AI assistant that the coder, the developer is going to implement. What I mean by that is you're ultimately, when you have a finding and you're communicating it to a dev, you're having a conversation. You have to convince them that it's real. You have to show your evidence. You have to, you know, prove the value, prove the impact. And then the developer has to know what success is. And if you're not a developer, maybe you don't know exactly the lines of, maybe you do, but maybe you don't in every situation. Like, You need to change these 3 lines. If the developer goes, oh, I just have to change these 3 lines, that takes all of the confusion out of it. And so communicating the suggested code fix has the intent. And so the developer who owns the code, owns the behavior, owns the testing harness, owns all of that custom stuff that's very different per team, even within an organization, makes it so that if you put that package on their front doorstep, They can go, oh yeah, and then immediately slot that into their dev workflow. Then they own the fix, but they're given every part of the context of the entire finding, right? And they know exactly what success looks like. Like I said, the other way is like, I'm just gonna fire PRs over to you with code fixes, and devs are gonna be like, how do I test this? Like, what do I do? My local build requires all this custom tooling, and you're just jamming in one-line code fixes. We have a standard, all that stuff. So I feel like there's sort of 2 motions, and one is kind, and one is just get it off of my plate, the ticket cannon. And I think I prefer the former, of course, because ultimately it's a learning moment, hopefully. **35:36 Robert Hurlbut:** It'll be interesting to study the drift in applications. Some researcher out there, take a note here. And then send me the paper in a couple of years. But the drift that happens in applications as AI starts to make more code changes, like, do we see a drift in the model that you were describing as kind of the kind or the— I think of it as more of a positive approach to this. Once again, the human's still in the loop. They're getting the context, they're getting the code fix, but they're also able to verify it. And then make the change and then adjust any test requirements or harnesses or anything. It'll be interesting to see over time when people use the generate PRs and just let them rip mode of this, like, what is the destabilization that we see in applications over a period of time as that code, which is not always optimized, not always the cleanest way or the way that a senior developer would fix it, Will it work? It works in the moment. But what happens when you start stacking a bunch of those on top of each other in a codebase? Do you see a drift where it starts to become— where you reach a level of instability because there are just so many spaghetti calls that have been introduced over a period of time? **36:56 Brad Geesaman:** Yeah, I think you have something there. I think your intuition is exactly right, is that it's just patchwork, you know, PR, patch this, patch this, and there's duplicitous logic, and then you have, you know, 2 problems now. Yes, maybe the vulnerability is fixed, but now it's 2 libraries that are being used for basically the same thing but differently, and I have, you know, spider web of code flows. **37:20 Robert Hurlbut:** Yeah, I don't want to actually study it. I just want someone else to write the paper and send it to me. Maybe give me a call out somewhere in there that they heard it here or something. But, um, all right, so yeah, I mean, I'll be happy to show up at an award show. I can I do acceptance speeches. I can try some AI-generated comedy perhaps, which once in a while does get a chuckle. I found in the early days of LLMs, I would have them generate— have an LLM generate a joke at the start of a meeting or something when you're just kind of hanging around just to try on people. **37:52 Brad Geesaman:** It wasn't bad. **37:54 Robert Hurlbut:** I mean, if I only had— **37:55 Brad Geesaman:** If it fails, you know who to blame. Yeah. **37:58 Robert Hurlbut:** There was only one that I couldn't— it created one that I'm like, I can't read this. I can't say this out loud. This is a fireable offense in a lot of companies, and this LLM just recommended this joke to me. So, gonna skip that one. But to wrap up the conversation, Brad, I'm always— I wanna be careful with predictions in the world of AI because things are moving so fast. Like, we could blink and like the prediction could be made false or true in 48 hours from now when somebody releases something else or whatever. But If we narrow our scope here, our context down to specifically AppSec and the intersection of AppSec and LLMs, what do you, I mean, somebody who's studying this a lot, like what do you see, you know, a year from now, 2 years from now that we should expect? **38:49 Brad Geesaman:** What I wish for is a couple things. And I'm hoping that this is the case, right? I see sort of signs that might trend to this, but so far nothing really close to it, which is inside the dev loop security analysis that's meaningful and just the right amount of clippy in your taskbar. And what I mean by that is like, I'm developing with code assistance. I'm saying I'm a developer, I'm developing with code assistance, and my objective is to make feature, write the test, and all the models, all the training, all the examples, all the cursors, all the windsurfs, Claude codes are shipping a working prototype or a working app or making that feature work. And the extra tokens, the extra cost, the extra effort is just not the same incentive there. It's not easy enough. It's not, it's too high friction to go, okay, now that that works, let me start the entire process of modeling and looking for issues specific to this code base and doing all that without it meaningfully affecting my developer velocity. You could spend just as much time doing that security review as building the feature, in some cases more if it's a complex feature. And that, like, speed of that loop is just too slow. Hopefully, what our model providers are doing are either making, or we're getting smaller models that can do this, that are more focused and fine-tuned on this, that can run quickly enough. or the large model providers that are backing the code assistance have extra prompting, extra steps, and it's not overly expensive or overly frictionful to be able to get a security pass tighter in the dev loop. That still doesn't mean you can't go without scanning once it's in a pipeline to make sure, you know, it's validated going into production, et cetera. But like, just bringing that loop closer down so that we're just not generating the the insecure code. We have a pass on it. Well, inside that loop where it's the cheapest to fix. And so that's where I believe secure coding can and should be going to. And I hope that that is the case in the very near term. Currently, generated code is generated code. It is not reviewed or security-reviewed code. It doesn't have the threat model in its prompt unless you tell it to look at it adversarially. It won't do that. It'll try to please you and make you happy. Oh, look, I made the button. Like, button clicks an unauthenticated endpoint that you just made. Like, that's not in the prompt. So I just made you the button. There you go. So I would love that because our world depends on apps and the code that we are generating is not slowing down. It is getting more and more generated by AI. So the leveraged fix is to not generate the insecure code in the first place, to get as close as we can in that loop and make sure what comes out of it is at least unit tested or slightly integration tested, you know, something that's local to that behavior, at least has had some look at, some hammering on before it ships. **41:51 Chris Romeo:** Cool. **41:53 Robert Hurlbut:** Well, Brad, thank you for joining the show to share your deep insight and wisdom into how AI and LLMs are working in the world of AppSec. I know I've learned a couple of different things in, in this exchange, so I've really enjoyed our time together. So look forward to, uh, having you on the show again in the future where we can dive into however AI has changed in the last 7 minutes. **42:16 Brad Geesaman:** Really appreciate it. Thanks so much, Chris. --- Source: https://appsecpodcast.com/brad-geesaman-redefining-appsec-with-ai-shrinking-toil-expanding-impact-how-llms-are-able-to-reduce-toil-in-triage-heavy-appsec-workflows/