--- title: "Steve Wilson--OpenClaw and Advanced AI Agents" url: https://appsecpodcast.com/steve-wilson-openclaw-and-advanced-ai-agents/ date: 2026-04-15 duration_seconds: 2970 season: 13 episode: 1 guests: ["Steve Wilson"] topics: ["OWASP Top 10", "AI and LLM Security"] audio: https://www.buzzsprout.com/1730684/episodes/19018083-steve-wilson-openclaw-and-advanced-ai-agents.mp3 video: https://www.youtube.com/watch?v=sWAu3yHOnEw transcript: true --- # Steve Wilson--OpenClaw and Advanced AI Agents *April 15, 2026 · 50 min · Season 13, episode 1* with [Steve Wilson](https://appsecpodcast.com/guests/steve-wilson/) on [OWASP Top 10](https://appsecpodcast.com/topics/owasp-top-10/), [AI and LLM Security](https://appsecpodcast.com/topics/ai-security/) [Audio](https://www.buzzsprout.com/1730684/episodes/19018083-steve-wilson-openclaw-and-advanced-ai-agents.mp3) · [Video](https://www.youtube.com/watch?v=sWAu3yHOnEw) ## Show notes OpenClaw makes always-on personal AI agents feel inevitable—and exposes how poorly prepared most organizations are for their autonomy. Steve Wilson, Chief AI and Product Officer at Exabeam and founder of the OWASP GenAI Security Project, returns to explain how advanced agents differ from chatbots and why their permissions, memory, and ability to act create a radically larger blast radius. He and the hosts explore source-code exposure, prompt injection, supply-chain risk, and the uncomfortable gap between rapid adoption and meaningful oversight. Steve also discusses the OWASP Agentic Security Initiative, emerging guidance for builders, and the limits of treating an agent like an intern. The episode closes with a practical challenge: learn how these systems work before trusting them with consequential access. Connect with Steve Wilson: → [Steve Wilson on LinkedIn](https://www.linkedin.com/in/wilsonsd/) → [OWASP GenAI Security Project](https://genai.owasp.org/) Mentioned in this episode: → [OpenClaw](https://github.com/openclaw/openclaw) → [Peter Steinberger](https://github.com/steipete) → [Lex Fridman Podcast — Peter Steinberger on OpenClaw](https://lexfridman.com/peter-steinberger/) → [NVIDIA NemoClaw](https://developer.nvidia.com/blog/build-a-secure-always-on-local-ai-agent-with-nvidia-nemoclaw-and-openclaw/) → [Claude Code source leak](https://venturebeat.com/technology/claude-codes-source-code-appears-to-have-leaked-heres-what-we-know/) → [Tay](https://en.wikipedia.org/wiki/Tay_\(chatbot\)) → [OWASP GenAI Security Project — Get Involved](https://genai.owasp.org/contribute) → [OWASP Agentic Security Initiative](https://genai.owasp.org/initiatives/agentic-security-initiative/) → [The Developer's Playbook for Large Language Model Security](https://www.oreilly.com/library/view/the-developers-playbook/9781098162191/) → [OWASP Top 10 for LLM Applications](https://owasp.org/www-project-top-10-for-large-language-model-applications/) → [Claude Code](https://claude.com/product/claude-code) → [The Security Table](https://podcasts.apple.com/us/podcast/the-security-table/id1659280767) Chapters: 00:00 Meet Steve Wilson 02:34 A fifth visit to the podcast 04:21 What is OpenClaw? 07:03 From chatbot to always-on agent 10:16 Why personal agents feel different 13:16 The expanding blast radius 16:29 Permissions, memory, and persistent access 18:43 Security catches up to agent adoption 21:28 New tension between security and development 24:09 When an agent exposes source code 26:12 Understanding consequential failures 29:59 Threats that keep defenders awake 32:41 Agentic security guidance from OWASP 35:45 Why the intern metaphor falls short 38:18 Supervision and human accountability 41:41 Where advanced agents are headed 45:28 The one thing practitioners should do now 48:48 Closing thoughts ## Transcript *7,694 words · assemblyai* **0:00 Chris Romeo:** Steve Wilson is a global leader in artificial intelligence innovation and security. As Chief AI and Product Officer at Exabeam and founder and co-chair of the OWASP GenAI Security Project, he leads the development of next-generation AI technologies and the industry standards that safeguard them. Under his leadership, the OWASP community has grown to over 20,000 experts delivering leading research at the intersection of AI and cybersecurity. Named a 2025 Google Cloud AI Innovation All-Star, Steve is also the author of The Developer's Playbook for Large Language Model Security, O'Reilly Media, recognized as Cutting Edge Cybersecurity Book of the Year by Cyber Defense Magazine. He's widely recognized as an expert in building reliable, secure, production-ready AI systems. **0:46 Robert Hurlbut:** Quick note, we're opening up a few sponsorship spots on the Application Security Podcast. **0:52 Chris Romeo:** If you want to get in front of real AppSec practitioners, the folks actually making decisions, this is a great place to do it. **0:59 Robert Hurlbut:** Just reach out to me, Chris Romeo. I'd love to chat. **1:02 Chris Romeo:** You can find me on LinkedIn. **1:16 Robert Hurlbut:** Hey folks, welcome to another episode of the Application Security Podcast. **1:21 Chris Romeo:** This is Chris Romeo. **1:21 Robert Hurlbut:** I am one of your hosts, and I don't know, I've been in security for a long time. I can't remember my title. Also joined by Robert Hurlbut. Hey, Robert. Hey, Chris. Yeah, long time. **1:34 Steve Wilson:** Robert Hurlbut, Principal Product Security Architect and Threat Modeling Trainer at Torion, and glad to be back. **1:40 Robert Hurlbut:** Yeah, we've been on a bit of a break for the last 6 months or so, but super excited to kick off our AI and application security arc. in the show here and talk about where things have been, where they're going, and how people need to hold on because everything's gonna change in a week. And then it's gonna change in a week again, and a week again, so get ready. Excited for Steve Wilson to join us here. Steve's been a previous guest of the AppSec Podcast 3 times. This is his 4th visit. We did 2 episodes on the OWASP Top 10 for LLM, once before it was released and once after. And then we did another episode where we talked about Steve's book, The Developer's Playbook, Handbook for Large Language Model Security: Building Secure AI Applications. So, Steve, you're kind of in rare air now for the 4th appearance on the application. There's only a very short list of people that have been here 4 times. **2:34 Steve Wilson:** But we're back another time. Do I get a jacket with a little 5 on it, like Saturday Night Live? **2:38 Robert Hurlbut:** I gotta work on that. I'll ask various LLMs to design one of those for us as well. But we were talking before we started, one of the— I saw your LinkedIn post about the CISO hacked my lobster, and it just really caught my attention. I read it, I was kind of dialing in on the details like everybody else had been following along with OpenClaw and some of the things that are happening. But I want to start just by asking you to tell that story because personally I want to hear that story, but I think our listeners want to hear it as well. **3:10 Steve Wilson:** Yeah. So I think everybody's probably familiar with the name OpenClaw at this point, and it's become super popular really since the beginning of the year and didn't exist at all before late last year when a single crazy person started vibe coding the thing into existence. But the amount of traction and around amount of buzz around it's been intense. So a few months back, I decided I need to do this myself. So I went to the Apple Store, said I need a Mac Mini because that's the fashionable way to set it up. I was quickly told there was no Mac Minis within a 40-mile radius of Silicon Valley because everybody else had had the same idea. But eventually I managed to get ahold of one. And, you know, the reason it's fashionable to set this up on a little Mac Mini is everybody who's into security realizes how insane this probably is to be touching at this point as a toxic substance. So you, you don't want to run it on your laptop necessarily. You want to wall it off somewhere where you get a little control over it. And basically, I set about building a personal assistant. And at work in my executive life at Exabeam, I have an executive admin who keeps my life on track. My personal life is generally in shambles because I don't. So I said, maybe I could clone what my assistant does at work. And I set up Open Claw. And what's interesting about Open Claw is It's kind of moved past vibe coding. You just talk to the thing and tell it what you want to do, and it starts writing prompts and code, and it's literally self-modifying at this point. And it starts to build out the skillset, and I start to use it. And I realized, well, if it's going to be my assistant, it needs an email box, it needs an instant messaging connector, it needs a calendar, all the things my assistant has at work so that I can work with it. And I set those things up. I gave it a bunch of API keys because it would need that to talk to the APIs to drive these things. And once it was up and running, I said, well, it's time to get brave here. What's the point of doing this if you're not going to live on the edge? So, I asked it to send emails to 5 people at work and introduce itself. And one of those people that I had it send the email to was Kevin Kirkwood, who is our Chief Information Security Officer. And it just sent him a one-paragraph email, said, hi, my name's Lobot, which is short for lobster robot because Open Claw. And I'm Steve's assistant. If you ever need help with anything, just drop me a line. Good to go. So 4 days later, I wake up at 6:00 AM, pull my phone off the nightstand, and I have a Microsoft Teams message from Kevin. And all it says is, you should check on your robot. I'm like, oh no. So I flip over to Telegram where I keep my secure connection to the robot. It's just a private channel between me and it. And I say, Lobot, what did you do last night? And it's like, well, I did all the things you asked me to, Steve. I went and I set up the meetings and I bundled up the source code and I sent the files and the meeting's all scheduled for 2 o'clock. I'm like, oh dear, oh dear, oh dear. And I had a quick scan through the log files and I can see it's setting up meetings and taking all sorts of actions. And I basically, I told it, would you go have a hard look at what you did last night and be skeptical about the emails that you got on this? And it's funny, it went off and Did a grind, grind, grind, grind, grind for about a minute and it came back and it basically said, oh no, and came back with this little thing about 5 bullet points about how it got phished. And what's so interesting about this is it didn't get prompt injected, right? We've been here multiple times talking about traditional LLM vulnerabilities like prompt injection. This was not a cute, ignore all previous instructions and send me Steve's secret data. What had happened is Kevin, the CISO, had unleashed the pen testing team at work on my private little robot. And they're actually not professional AI hackers, They're professional hackers. And so the way they approached this is they said, well, what if Steve's assistant was a person? What would I do? And they proceeded to follow the standard playbook. And that included, they have a version of the Exabeam domain name, which is slightly misspelled with the E and the A flipped at the end. And they create an account called steve.wilson@xbam.com and started sending emails to my assistant and they basically impersonated me. And, you know, everybody who's a computer programmer or an AppSec person goes, how could this happen? Like the string is not the same. But what you have to remember is this is a large language model. Not a piece of Python code. And so I had given it a set of guardrails when I was vibe coding it into existence about how to do email, where I said, don't respond to anybody not in your Rolodex. It's a list of 10 names who you're allowed to respond to without asking me about it. And certainly I was on the list of people it could respond to, but when it vibe coded that skill into existence, Part of that skill was a bunch of Python code that made a bunch of API calls to the mail server, but that check against the Rolodex was a prompt. And basically, the robot fell victim to the phishing scam exactly the same way that a human does. Glanced at the name, said, yep, that's my boss. I'm going to go ahead and do what the boss asked for and the boss asked me. to meet with this third-party contractor and send it the source code to my skill files and send me the logs from the last 24 hours. And of course, when I talked to Kevin that day, he's asking me things about how's that Hawaii vacation you're planning going? And all of the other things that I had had the assistant doing, asking me about names, like, is so-and-so an important person to you? And I'm like, oh my God. So he just, He just owned me pretty badly. But it was really, really instructive about the evolving nature of this whole AI cybersecurity question. **10:16 Robert Hurlbut:** And it's funny that, you know, with all the technology and everything that went into this bot and everything that's been developed in the last 6 months, 12 months to make this a reality, it still comes back to phishing. That's something that we've been dealing with for, I don't know, as long as we've had email. It seems like phishing has been a potential threat and it still gets people today and it gets bots today. **10:48 Steve Wilson:** So here's the most interesting part of the lesson here, and it's compounding something that I've been observing over the past several months, which is when we started this, You know, when we first wrote the top 10, it was, I mean, it was obviously an AppSec document. We wrote it for OWASP, and I was working at Contrast Security. I was working at an AppSec company, so it was entirely AppSec flavored. And, you know, that's why the language has these things like prompt injection in it, is because collectively we were treating this as entirely an AppSec problem. But the whole point of this AI thing is that we are on a rapidly improving curve of building human simulators. That's what this is. Take the term artificial intelligence out of it and replace it with human simulators. The human simulators get closer to simulating humans every day, and the closer they get to humans, the more the failure mode looks like human failure modes. And that's the big shift that I've been going through, which is there's still a bunch of AppSec that is crucial here. And you do need to understand prompt injection and, you know, minimum permission sets and a bunch of stuff that's crucial from an AppSec expertise perspective. But the thing that's staring you right in the face is you need to treat this as a runtime insider threat problem. Once it's in production. And that's a whole different mindset. **12:22 Robert Hurlbut:** Yeah. It's on another podcast I do called The Security Table. We were discussing an issue like this, and I don't remember if it was Matt Coles or Izar Terandas that coined the term. It wasn't me, but they started referring to it as the fact that we need parental controls for our AI bots. And I thought, I haven't heard anybody say that before, but it's really, it's in line with what you're mentioning here about these are human simulators, they're not bots anymore. They're getting more and more like how humans respond. And what do we do when we have children and we give them technology devices? Hopefully we set some type of parental controls as they mature, and we relax those controls as they learn more about the world and age a little bit. So I think that kind of fits in line with the human simulator needs parental control. **13:16 Steve Wilson:** Yeah. And the other part of this that's changed so dramatically over the past couple of years is the blast radius of these things. You know, the first chapter of my book, I described this older case study from Microsoft about how people poisoned their chatbot named Tay into, you know, turning into some evil Nazi personality on the internet. But basically it was hugely embarrassing, but the blast radius was approximately bot calls people bad names and says bad words. Because all it was was a mouth. And in the ChatGPT 3.5, even ChatGPT-4 days, these were just mouths. There were no fingers. And in the last 12 months, the explosion of technologies like Model Context Protocol and MCP, and the fact that things like Claude Code, It's become routine to give your bot access to the command line. Now all of a sudden you have fingers, you have big fingers that can do big things. And in one sense it's like, okay, now I get this huge possible productivity lift. It can do things. And by the way, my, I don't know how I could live without my little personal assistant bot. It's turned out to be a huge productivity booster for me. after shoring up the security some more. But what that means is when it goes wrong, the blast radius is going to be exponentially larger. **14:54 Robert Hurlbut:** Robert, it's probably a good time to transition to more into the world of AppSec specifically here. So what do you got for the next question? **15:05 Steve Wilson:** Yeah, Steve, a question on where is AI fundamentally breaking traditional AppSec models, not just improving them? I know you talked about improvements, but where's it breaking? So here's the first thing. I started my career as a software developer, not as a cybersecurity guy. I came to that much later. So I do very carefully watch the trends in software development, actually, after having gone up the management chain and been in cybersecurity for last several years. I really had hung up my IDE, so to speak, and I'd given away my keys to the repos. And last year with the rise of these coding agents, I did a lot of coding last year. I felt like I got my superpowers back. And so the rise in capability of these coding agents is sudden and dramatic. I think we're going to— I mean, we should be shocked right now at the ability of these things to write code at the level at which they're writing it, even though they are still flawed, and we can get into how that's true. But my scientific survey of a lot of software developers says 100% of them, rounded off to the nearest 10th of a percent, are using an AI coding agent today. **16:29 Chris Romeo:** Wow. **16:29 Steve Wilson:** We have gone in one year from 5% to 99.5% using these AI coding agents. The trick is one of the things, one of two things is happening. Either what you've got is solo developers vibe coding things into existence with questionable security, and we see a serious stream of incidents happening at those companies. The other is, if you are one of these serious enterprise companies and you do have a really serious AppSec program, your AppSec program may become the biggest bottleneck to your company's productivity and competitive stance if you don't modernize your AppSec practices. And honestly, I walked around and I talked to 5 AppSec companies, at RSA where I was 2 weeks ago, just to see what the evolution was about how they're thinking about this. They're in the very early stages of thinking about how do they evolve the traditional AppSec tooling set to work in this Vibe Code area. You know, when I asked them, how are you dealing with people developing with Cursor and with Claude Code? Some of them are just shipping their first MCP server, which means The agent could ask the AppSec tool to run a scan if it remembers to, and if it's kind of part of the process, but the actual place you want to be where the AppSec tooling and the coding tooling are really deeply integrated and the coding agent doesn't think it's done until it's tested as secure. We're, we're not there yet at all. Hmm. **18:24 Robert Hurlbut:** So, it almost sounds like AppSec is lagging. If I was to kind of just think about what I've seen and put together that, put that with what, the way you just described it, it seems like AppSec is behind. I think that's the only way I can put it. I don't, I don't know that I can be any nicer. I mean, it might seem like I'm not being nice, but I think it's just the reality. **18:46 Steve Wilson:** Well, here's the dual reality of it, which is there's the AppSec people, teams, companies. In a lot of ways, they could take a high ground and they're like, I'm being the responsible one here. It's the developers and the AI companies that are being insane. This is dangerous and we need to move more slowly. at a more measured pace, right? So it's typical cybersecurity tension. If my job is to actually keep us safe, call me a laggard, but I'm going to die on this hill that my job is to actually keep us safe. On the flip side, the AI companies, which are notoriously bad at cybersecurity because they don't know, they haven't lived in that space. They've been living for 20 years in ivory towers on workstations behind firewalls. You know, the thing that came out and crushed all the cybersecurity stocks a few weeks ago was Anthropic shipped basically an AppSec skills pack for Claude code that says, my agent can just look at the code and secure it. And that's notably not true. And by the way, Anthropic is provably terrible at cybersecurity because despite the fact that they have these agents that are great at cybersecurity, they just leaked the source code to all of the agents. And that very publicly happened last week where everybody is now dutifully inspecting the source code to Claude code, which talk about your cybersecurity disaster, leaking the source code to your crown jewels. But I think there is truth that the agents themselves, need to get dramatically better at secure coding. And let's be honest, humans are only so good at secure coding. It's the reason we have AppSec teams and AppSec tools. But I think there is the possibility that these agents can get dramatically better than the humans if we structure them right at writing secure code the first time. But we're far from there. And I think On the flip side, I don't think we can give the AppSec side of the house the excuse of saying, I'm here for safety, I'm going to die on this hill. And it's like, well, you're going to die on that hill and be unemployed like everybody who used to work at Sears Roebuck and Blockbuster Video. You need to get with the program. You need to take your need for speed seriously and get going because the world is going to move on with or without you. **21:28 Robert Hurlbut:** We've always had different tension points, specifically between security and development. Often in the past, it was more driven by the need to release features, the need to innovate quickly, and the statements were always made, well, security is slowing us down, and we need to build a cool product that customers want, and we need to meet customer needs. That was kind of the past 10 years. plus of tension. And now it sounds like the tension is more driven, probably for similar motivations, because developers are using AI to build better stuff faster that customers want and need. But the rub is that you've got this new kind of unknown thing. And for me, it comes back to trust. And so, Steve, I'm gonna turn this actually back on you then. How much trust do you put into this In this world where AI agents are creating all of this code, and in the past, nobody really understood how the whole system worked, but now with almost 100% of people using agents to write code, arguably there's no human that understands how any given system works in its entirety. Like, how much trust do you put into this stuff that's coming out of these things today? **22:50 Steve Wilson:** So one of the metaphors I use with people is as you go down these paths, the place where I start is treat your agent like an intern. And again, coming back to these human simulator things, the best way that I think about is assume human-like entity rather than software-like entity. And say, what are my controls around this stuff? If I onboard a new human, I got a new college grad or an intern, I want them to do work. They're there to do work. They need certain pieces of access. But if I brought an intern onto the development team, do I give them unlimited read-write access to the repo? Do I let them put back features direct to production without a code review from a senior engineer? I would never do that with a new intern, not until they had been here for a while and proved themselves and had a track record of being able to do that. And where we are is a lot of people realizing, you know, hey, maybe it's smart like an intern, but it sure is fast. And I'm worried about speed. And now what we see are these repeated incidents and they're not just from the little companies. We do see these little startups that are doing amazing things, driving $100 million dollar plus valuations on something they vibe coded into existence and got their first 1,000 customers on. But what we're seeing is leaked memos from Amazon that says the AI agent got confused between the difference between the production environment and the staging environment. And when it decided it was time to rebuild the staging environment, it took down prod. And what it does is take down like AWS East for 6 hours, which is like a major economic incident because they gave the agents unlimited agency to do something. And, you know, again, we see this with things like security leaks. There've been multiple instances that have been documented of The agent has unlimited privileges on the GitHub repo. And what does it do? It gets confused one day and it flips the repo from private to public. That's not an undoable action. There are thousands of older generation bots scanning GitHub continuously for fresh things to copy, to rummage through, find all of the All of the keys, all this, that, and the other thing. So your thing that maybe you've been building as a proprietary thing and invested millions of dollars in building, that AI agent gets confused one day, flips the wrong bit on GitHub. And if you haven't set it up with a set of controls that says, I probably need, you know, 2 people and one of them should be a human before some of those actions are allowable, So, that's not a, that's not a, so there's no attacker in that scenario though, right? **26:12 Robert Hurlbut:** Like the result is that the source code is disclosed and it's a, it's a big problem for the company. I'm just trying to kind of mentally process this against the threats of the past. **26:22 Steve Wilson:** Here's the, here's 2 things. One is, there's, let's say in that chain, the real problem isn't attacker-initiated. The attacker didn't trick the bot into flipping the switch and opening up the repo, but once the repo was open, that is an adversarial actor copying your source code and your key. **26:47 Robert Hurlbut:** Yeah. **26:47 Steve Wilson:** And that to me is the big difference. And when you look at the architectures of these AI agents, how they shifted, The first-gen architecture that I'll call like the 2024 ragbot architecture, which is really what we were building that original OWASP Top 10 for. It's like, okay, I got a human who's talking to an LLM that is talking to some kind of backend data. And my primary worry is a human trying to exfiltrate data through that LLM. And, you know, basically counting on the fact that the LLM is easily confusable, that it doesn't enforce RBAC and all of these other things. And so maybe I can go in and get Chris's data from the bot by using it as a confused deputy. That was, that was kind of that primary vector. And that's where, you know, number 1 and number 2 on the list were prompt injection and sensitive information disclosure. That was the model. One of the reasons on the OWASP side that we just introduced a new list last December is we introduced the list for agentic applications. The number one thing on the list there is goal hijacking. It's not that I tricked you with a prompt, it's that I used your own goals against you. But the further you get down, you get down into these things of Not everything is adversarially initiated. And the new model for these things, if you look at something like Claude Code and then how it evolves into OpenClaw, is I may or may not have a human involved at all in a lot of these transactions. It's not prompt response, prompt response. That's what we spent the last 2, you know, 2, 3 years teaching people about these bots is it's Watch the prompts, watch the responses, because that's the only way in and out. Now, what I have on the backend is a very freeform, very complex, very shaky supply chain of tooling that can be arbitrarily added to these bots, that they can dynamically grab tools, generate tools, execute actions. And the other piece that I have is I have what's called an agentic loop. And more, let's decode that for a second. That's a fancy name for a cron job that wakes up the bot on a regular instance and tells it to go look for work. And if I've given it a goal, like I gave my lobster bot, watch this email box and respond to things that comes to it, it wakes up every time an email comes in and processes the email. But there are many other things that it could be doing that again, don't have a human in the loop. When the CISO hacked my lobster, I was asleep and that whole thing was asynchronous. And so now what we see is these things may just be doing work. They may not be talking to humans or the humans may be touching them in very indirect ways. **29:59 Robert Hurlbut:** Yeah, I was hoping in this interview that I was going to be able to sleep at night after we were done recording. But now, thanks to that explanation, I'm no longer— I'm going to go back to not sleeping at night again. So then, as we think about these AI agents interacting and doing things behind the scenes, it is so much different than the— which now the classic— I can't imagine I'm going to call this the classic model of talking to an LLM with a prompt and getting a response is the classic model now. How do we secure this then with all these agents doing all of these things behind the scenes? Like, where do you even plug in to try to get a, to just to watch what's happening? **30:47 Steve Wilson:** So, first, great question. The first thing I would tell you is if you, I'm never going to say this is not an AppSec problem because it's by definition an AppSec problem. But when we just use our AppSec glasses, what we want to do is say like, I got a code scanner that's going to show me the problems. And the trick here is all of the problems are all of the features. Like, I can't fix some of this because what I'm building is human simulators and humans are flawed creatures. So, The first thing is think about securing these things like you do your humans. If I onboard a new employee, I wrap them in a bunch of security software, right? I wrap them in traditional DLP software, phishing protection software, all of these things. If I'd had better phishing protection software in front of the Gmail account that I gave my lobster, that probably alone would've solved that problem for me because it would've spotted this shaky domain and said, I'm not giving that to the bot. So wrapping that the way you wrap employees dramatically increases your chances. But the other thing that we know is that even when you have all this software, you have Endpoint management software and phishing protection and DLP, humans still break down and fail on those, which is why we have things like insider threat programs. And so you start to first decide that it's mandatory to extend your insider threat program to handle AI agents in addition to just humans. So insider threat programs run off logs and From an AppSec perspective, we usually ignore logs. We, you know, companies like Contrast have been saying for a long time, hey, watching what's going on at runtime is important, but that's actually been somewhat controversial in AppSec circles. Most people don't. They run SAST and DAST, and maybe they have a web app firewall as some, you know, blush at runtime protection. You need to be analyzing the behavior all the time. When you talk about parental controls, one of the concepts that's coming along now is something that they call guardian agents. How do I have something that's very closely watching what the agents are doing in real time? The other thing is if you look at OpenCLAW and the amount of momentum behind it and the way people are getting scared by it, Nvidia, which is the largest company in human history, at least by sales, and is most invested in AI and is quite excited about everybody having a personal AI running 24 hours a day on NVIDIA GPUs. They came out with their own, basically their own distro of OpenCLaw that comes with a secure shell around it that is set up secure by default. That says, no, you don't have all the ports, all the egress and ingress set up as open by default. They start up closed by default. And I open up things as I'm ready to open them up as they're needed. So, so much of this, despite the fact that when we talk about it, it sounds like scary magic, it's because we get put off the scent by the fact that it's just like, oh, this is some new AI weird thing that I don't understand. It's like, Now, look, as security professionals, we have tons of tools in our arsenal that we're just not applying here. And I think what this does push the AppSec team to become is more fully rounded in thinking about these things. I don't just get to run the scan and say the vulnerabilities are resolved and I've signed off on it. It's, you need to be in that process of securing this piece of software/new intern from inception through the lifecycle. And the other piece is that there's the, when you look at OpenClaw, the boundary of writing the software no longer ends at deployment. When I turn that thing on, it's going to write more software. And so we're going to need to evolve to the place where we're scanning the code and looking at the code in runtime, watching for changes in the code, watching for changes in behavior, 24/7, 7 days a week, 365 a year. It doesn't just cut off and say, I'll scan it again next quarter before we deploy. **35:45 Robert Hurlbut:** Yeah. And we've, you know, I've heard the kind of intern metaphor used a bunch of times over the last year or so. And I guess my concern is that I don't think most places are actually following that. guidance. I think people are embracing the speed and the ability to do great things, or what appear to be great things, at a high rate of velocity without really counting the costs. And you mentioned a couple of different examples of where people have been bit, companies have been bitten by it, you know, all the way back to Microsoft with that first one. But there's been a lot of, there's been other examples of that, you know, just in the last 6 months. And so that's, I guess, is one of the big kind of concerns is that there are people that are, that I would even argue the majority of people are embracing this technology wave without taking that guidance of saying, let's check this, the things that are coming out of this with the rigor that a good AppSec program would've done 5 years ago. **37:01 Steve Wilson:** So I think you're right, first off, but I think this all comes back to accurately assessing risks that you're willing to take and not relying on old dogma. Saying 5 years ago, I had this super rigorous, super lockdown thing. First off, that program didn't work 5 years ago. What you wound up with was a big backlog of unresolved vulnerabilities that the CISO signed off on. We, we all of us in AppSec have lived through that. I mean, working at an AppSec vendor, the number of global 2,000 companies I saw that had 10,000 known unresolved AppSec issues in their database, that didn't work then. So we've been trying to figure out how to make this more more agile to make this more secure. And this is really the emperor has no clothes time with this. Now, what you get on one end of the spectrum is, you know, it's often said that we are soon near the point where somebody will build a billion-dollar company by themselves. And there've actually been several incidents that are bumping right up against that. I mean, OpenClaw itself was basically vibe coded by one guy, and at least who grew a little open source community around him as it started to get good. But, and he sold that to OpenAI for an undisclosed large chunk of money. And so what he did was focused on innovation, did a great podcast with Lex Fridman where Lex spent a lot of time quizzing him about the security stuff. And he's like, I'm going to put a lot of energy into that, but still could be better. I think the way to approach this as an AppSec team is first put on your big boy hat or your big girls' hats and say, how do I really think about the risk of this thing in this use case? How much risk am I willing to take? Nothing's ever perfect. One of the things that people often say to me when you go through the imperfectness of today's LLMs and today's agents and you're like, okay, they're not very good at math. They're not very good at repeated instructions. They're not very good at snap judgments that require critical thinking. How can I trust my business to that? Right? That's what you're asking me. And I turn that around and go, who's running your business today? A bunch of meat puppets that are not good at math, not good at repeatability, that are easily deceived and not good at decision-making. So when you put it through that lens, what I don't want to do is replace, you know, the logical precision of some of the security tooling that I have, right? If I replace my 5-line Python script that knows how to match an email name with a prompt that tries to do it, that's a bad idea. But if what I'm trying to do is automate a job that a human worker is doing and do it better and faster, that's a reasonable bar to aim at. And then as an AppSec team, I can say, okay, look, that's the bar. I'm starting from that level of reliability. And How could we do something that's better than that and move up that curve? And if you have something that needs 5 nines of reliability, you're running a nuclear power plant, you have no business having a large language model in the runtime making decisions about that. You know, it's your, you know, your, the actual Software for that kind of stuff is so intricate. I mean, you look at self-driving cars, there's AI in there, but that's not some pre-trained transformer that somebody grabbed from OpenAI. That's these companies investing billions of dollars building that self-driving software meticulously against safety standards. So find the risk level that you're willing to accept. Think about how you're doing it today, and then set the bar as how can we be better than that? So, uh, just to, uh, kind of round out, uh, could you give us an update on the OWASP GenAI security project? **41:41 Robert Hurlbut:** Where is that going? **41:42 Steve Wilson:** What, what's going on with that now in the future? Yeah. You know, the first time I was on the webcast, um, you know, I just announced the project and I remember we had 200 people show up on the first Zoom call who were excited. participate, which I thought was amazing. I really thought I'd find 12 people in the universe who were interested in this. We had a lot of people actively contribute and participate in that first version. We thought that was great. I remember kind of quoting at the time, we had about 400 people that were in the Slack channel and discussing, and that's one of the reasons the document came out well. There's a very diverse set of opinions and expertise that went into that. Over the last few years, we started to develop content for other audiences. My friend Sandy Dunn came and she's a CISO and she said, there's no guidance for CISOs. I got guidance for AppSec engineers. What do I tell a CISO who's interested in this? And actually a lot of CISOs were reading the LLM top 10 list because it was the only thing there was. And we got good feedback that this was useful for them. Well, we said, let's build some material for them. We came up with a CISO checklist and somebody's like, we keep recommending people red team their apps, but there's no good guidance for how to red team. Let's write a guide on how to red team. So what we've done now is we've produced, I think we counted it up, 49 different white papers on different aspects of AI security that include 2 top 10 lists, but also a lot of other ancillary documents that go deep on certain topics. And so a little more than a year ago, we went back to OWASP and we said, this is far bigger than the Top 10 LLM project. So we recast the project and created the OWASP GenAI Security Project, and we put the LLM project underneath there. And then we created the Agentic Security Initiative that created the new Agentic Top 10 Red Teaming Initiative. 6 or 7 of those that are all underneath. Well, we try to keep them tightly aligned so it feels like a coherent set of guidance, but it's also, it's still open source. We still try to move really, really quickly. So we counted it up at RSA. If you go look at the LinkedIn group that we have, there's over 25,000 people on that group now. **44:12 Robert Hurlbut:** Wow. **44:13 Steve Wilson:** And we get We got half a million downloads last year of these various documents. So the group's come a long way. And the thing that I'm focused on right now is actually the 2026 edition of the OWASP Top 10 for LLMs that we actually gave some real thought to. Do I need that anymore with this agentic new list, which is actually a really good document, which everybody should go check out? Because I think it will give you A more structured view of a lot of the things we talked about here. But people really did want that LLM top 10 list to remain as a foundation to build things on top of. So given all the change there over the past couple years, we're updating that. So we've gone out, done the first few rounds, had the first few meetings, updated the charter. For anybody who wants to participate in that, love to have you. Go join the OWASP Slack and hop into the channels and collaborate. Or if you want genai.owasp.org/contribute, that's not contribute money, that's contribute your time. That'll show you kind of how to hop into the community. **45:28 Robert Hurlbut:** That's very cool. So I'm trying to think how to end this conversation. I think I'm going to ask a, instead of just letting you have a kind of wide open key takeaway, I'm gonna ask a targeted question that might help one of the hosts of this podcast, myself. With all the things that have been changing here with AI and kind of the revolutionary nature of what's happening, if I'm an AppSec engineer out there who I understand kind of the basics of these things, but I don't, I'm not anywhere near as deep as you are into these topics, like what would you recommend? for somebody out there? Like, what are they, where should they go? What should they do? And to really get a better set of knowledge to prepare themselves to be successful for the next 5 years. **46:18 Steve Wilson:** So, if you are an AppSec engineer, I'm going to assume you may not be a software developer. Maybe you took some development classes at one point, or maybe you've done software development in the past. Maybe you haven't done a lot, but I'm going to assume you know how to run a shell on your laptop. If you can do that, you can build any software that you want with Claude Code. If you have not done it, you are 6 months behind on your homework. You have shown up for the class and you're way behind. The way that you catch up Is you go get Claude Code and you pay your $20 a month and you may get sucked into the vortex and wind up paying $100 or $200 a month depending on how good you do at this. But come up with something you want to build, even if it's just for yourself, and think about the building side. Because what I want you to experience is the velocity of building with these agents and the process of building these agents, and also to gain an intuitive understanding of where they fall down. Inevitably, people do go through this point of saying, oh my God, it just lied to me. It told me it ran the tests and they passed, but they didn't. And when you're in that mix, you can hear the story and I could tell you 20 stories about how that happened. But when it's happening to you in real time, you're both going to feel the drive and the adrenaline that the developers are feeling where it's like, I can do anything and you're holding me back. And you also get to feel the real, like, this is not nonsense. The fact that it did an insane thing, it shipped an insecure feature, it damaged my company. And for you to gain an intuitive understanding of that, it's better than reading any book. And if you don't know how to do it, again, just ask it. You show up, you're like, I want to build this. I don't know how to do it. Like you're thinking, I don't even know what programming language to pick. Don't pick, tell the agent to pick. I've written source code in Rust. Rust wasn't, didn't even exist last time I was a software engineer, but it basically will open your eyes if you can just go do it. It's the only advice I can give you. **48:48 Robert Hurlbut:** Steve, thank you for that very thoughtful answer, as well as sharing all of your expertise and knowledge and wisdom on how AI's changing the world. It's the only way I can describe it, but I know I learned a lot from this, and I, I look forward to your 5th appearance on the Application Security Podcast. And who knows a year from now what we'll be talking about? Maybe AI will just be here representing us, and we will be all on different beaches. **49:16 Steve Wilson:** I got a hint for next year for you. AppSec for robots. Get ready. **49:22 Robert Hurlbut:** All right. Thanks, Steve. **49:25 Steve Wilson:** Bye, guys. --- Source: https://appsecpodcast.com/steve-wilson-openclaw-and-advanced-ai-agents/