Skip to content
AppSec PodcastThe Application Security Podcast — home
43 min

Rob van der Veer -- OWASP AI Security & Privacy Guide

with Rob van der Veer

on AI and LLM Security and Privacy and Compliance

Audio hosted by Buzzsprout. Nothing loads until you press play.

Rob van der Veer has a 30-year background in software engineering, building AI businesses, creating software, and assessing software. He is a senior director at the Software Improvement Group, where he established practices for AI, security, and privacy. Rob is involved in several standardization initiatives like OWASP SAMM, ENISA, CIP, and AI security & privacy guide. He leads the writing group for the new ISO standard on AI engineering: 5338. Rob co-leads the OWASP integration project, with as a key result, aiming to create alignment in the standards landscape. Rob joins us to introduce the OWASP AI Security and Privacy Guide. We cover Rob’s observations on how AI engineering differs from regular software engineering, typical software engineering pitfalls for AI engineers, the new guide’s scope, threats introduced with AI, and mitigations that orgs and teams can use to build a secure AI system. We hope you enjoy this conversation with…Rob van der Veer.

Show Notes:

  • Visit the OWASP Security & Privacy Guide here –>

Mentioned in this episode

Enjoyed this one? Get Reasonable AppSec, the newsletter with new episodes and picks from the archive.

Transcript

6,478 words · assemblyai

0:00Chris RomeoRob van der Veer has a 30-year background in software engineering, building AI businesses, creating software, and assessing software. He's a senior director at the Software Improvement Group, where he established the practices for AI security and privacy. Rob's involved in several standardization initiatives like OWASP SAM, ENISA SIP, and the AI Security and Privacy Guide, and he leads the writing group for the new ISO standard on AI engineering, 5338. Rob co-leads the OWASP Integration Project with OpenCRE.org as a key result, aiming to create alignment in the landscape of standards. Rob joins us to introduce the OWASP AI Security and Privacy Guide. We cover Rob's observations on how AI engineering differs from regular software engineering, typical software engineering pitfalls for AI engineers, the scope of the new guide, threats that were introduced, with artificial intelligence and mitigations that organizations and teams can use to build a secure AI system. We hope you enjoy this conversation with Rob Vanderveer.

1:04None of the top 50 university programs teach secure coding in their curriculum. At Security Journey, we help enterprises reduce vulnerabilities through application security education for developers and everyone in the SDLC. With over 400 up-to-date lessons created by industry-leading security experts, and a programmatic approach that creates security champions, our program has increased AppSec knowledge as much as 85%. Visit securityjourney.com to try our training today.

1:31Chris RomeoHey folks, welcome to another episode of the Application Security Podcast. This is Chris Romeo, CEO of Curve Ventures, and also co-host of, you know, said podcast. Super excited to have Robert once again with me talking about something related to AI. But first, Robert, who are you? Yeah, Robert Hurlbut, threat modeling architect and lead. And yeah, AI has definitely been an interesting topic for me over a number of years. I've looked at threat models around AI as well. So really excited to dive into this topic today. Yeah, I think from my perspective, I don't even know enough to be dangerous at this point. Like, I'm whatever level is below knows enough to be dangerous about a topic. So I'm super excited to learn. But let's kick off with Rob Vanderveer, who's here to talk about a lot of different things related to AI. But first, Rob, let's hear your security origin story. So how did you get into this world of security? But also, I'm going to put a little tweak on it. How about AI as well? So talk to us about how you got into security, but also give us some flavor on AI, how you got into that.

2:42Rob van der VeerSure, Chris. Robert, Chris, thank you very much for having me here. Uh, how did it start? It started in 1986, somewhere in the previous century, where I won a programming competition organized by the Dutch government, and I won a trip to the UK with a couple of the other contestants, and we went to multiple sites. And one of them was a base from the Royal Air Force. And they asked us, hey, why don't you try to hack us? And we tried and we did. And I liked it. It wasn't really a big deal, but the newspapers picked up on it. And in every terminal on that site, we managed to display the text, the Russkies are coming. And it created quite a bit of excitement. Quite a fuzz. And that's actually how, how it all got started. I went to university, did my master's in computer science with the specialization of AI. So that's where I really started to get interested in AI and whole philosophy, but also how neural networks work, etc. And I started in the AI business in 1992 as a data scientist, as a programmer, in a company called Scented Machine Research. And, uh, I grew as a development manager. Uh, I became CTO, and we did all kinds of AI stuff. We did face recognition, texts, uh, forensics, organizing blind dates— it's a very hard optimization problem— predictive crime analysis for the police. And because we dealt with so much data, we were dealing also with a lot of security and privacy issues. The privacy issues in particular with the police applications back then. So, I needed to learn a lot about it. But when I learned the most about security is when I had a gaming site. We had clanbase.com, which was a gaming site with about a million active gamers all around the world. It was the first of its kind. It was really nice to build it. We built it in PHP and MySQL like you do. And we were perhaps, I don't know, but we could have been the first case of cross-site scripting. And when that happened, we really didn't believe what we saw. And I guess that the term XSS wasn't really coined yet back then because we're talking 1998. Anyway, we had to learn a lot about security because all these gamers were trying to hack us to, you know, because that's what you do as a gamer. You try to influence the scores of your clan. To win the competitions. In the AI industry, I stayed a very long time building businesses. Also was a CEO for nine years for doing all kinds of AI stuff, particularly predictive policing. And in after eighteen years, I felt the company was you know solid enough. You know, this was really avant la lettre. Nobody wanted to hear the term artificial intelligence nowadays. It's a term that sells. Back then it was a term that really terrified people. And it seems that we could go back to that, but I don't know, hard to predict. But anyway, it was a different world. And I went on to a company called Media Lab where we also did some information retrieval applications, visualization of information retrieval terms. And after that, I went on to Software Improvement Group. That was in 2012, and that's where I still work, and it's the best professional time of my life. I got to establish the security and privacy practice, the AI practice. We're doing evaluations of systems, really cool, big, important systems all over the world. We have a platform that helps our clients analyze these systems, AI systems, but also traditional traditional systems. And maybe it's a bit of a long story, but I think you got the most important stuff on how it all works.

6:55Chris RomeoYeah, no, it's helpful to understand because you're coming to the interview, to this conversation with a really different background than a lot of people that we talk to. And so, a lot of the people that we talk to have made some path into security at some point from a lot of different backgrounds. And so, we didn't mention this from the top, but you are one of the project leads or the project lead for the OWASP AI Security and Privacy Guide. So, before we even go deeper into AI and some of the other things we want to talk about, from a background perspective, so, as someone who's not coming from a security perspective, how did you learn all the things about security to know what would go into this guide?

7:39Rob van der VeerWell, I had to learn stuff. And you, Yeah, you read books, you Google, you go to conferences. And when it comes to AI, this is all pretty bleeding edge, so to speak. I think it's a very current topic and many, many organizations are jumping on it as OWASP is jumping on it. But there's also the work of Gary McGraw, for example, with Berryville Institute for Machine Learning. who's done a lot for software engineering, best practices for machine learning. As an industry, we're still learning on how to do proper software engineering for machine learning. The ISO 5338 standard on which I worked very hard is about to be released in August. So, only now we're going to be able to work with an ISO standard on software engineering. So, we had to all figure it out in this industry on the types of attacks that we stumbled on and tried to document them. And be a field of research actually and share these things. And that's also why I started the AI Security and Privacy Guide at OWASP, because I work for a company and I could have published a white paper from the company, but I know that I don't know everything. And by using an OWASP project to get input from everybody who wants to provide input, that helps pushing that cutting edge and getting this overview of the state of the art. So it's been a search.

9:21Chris RomeoSo when we're thinking about, you know, in your experience and observations on AI engineering and, you know, in terms of this guide and other guidance that's out there, what do you see as different between AI engineering and regular software engineering?

9:40Rob van der VeerYes, every time I get to talk about this, I always start by saying that yes, it is different. But it's also much more, uh, the same as software engineering is. And by viewing it as regular software engineering, we can use the practices that we're accustomed with— versioning, unit testing, documentation— and apply those to AI. In that sense, it's nothing really super eccentric, not requiring anything really different from a lifecycle perspective. So the 5338 standard is discussing AI engineering from the perspective of the system lifecycle, and then just goes over a few things that are different. But the main things that are different is that there's also data engineering involved and model engineering involved. So you need to collect a lot of data because in machine learning, the data dictates the behavior of your model. And the model engineering is there to experiment with different types of algorithms, to— with different types of features, uh, with different types of approaches and combinations of things. And this whole experimentation, that's also very typical of AI. Uh, typically if you build a software system, you build a certain volume of code, and with, with a few rules you can sort of determine how much effort went into it based on the productivity and the number of, uh, a number of lines. But with AI, they could have spent a very, very long time trying different approaches. So this experimental aspect is also really what sets AI apart. And because you have this data engineering, you also have an extra attack surface, which is a security problem. All the things that you're doing with collecting the data, processing the data, storing it in a data warehouse, et cetera, That takes place outside of your application. So, where normally with application security, you focus on the data that's inside of your application, you now have to deal with a whole development lifecycle that involves data that needs to be protected. That's what makes it different. And what makes it especially different is that this data is real data. It's not test data that you use as a development team to test your system. It's not anonymized or, or, or, or synthesize stuff. No, you need real data by definition because you need to build it— a model based on reality needs to work in the real world. That is why engineers, AI engineers, data scientists have access to it. And this requires a certain hardening when it comes to access to this data. Not everybody in the team needs access to it. Maybe only people who transform it, maybe not even the people who build the models. Let them use a setup where they can run the model based on the data, but not per se provide them access to that data. So, these are examples of countermeasures, how to deal with this difficult attack surface that is data engineering. Go ahead.

12:49Chris RomeoI've heard somebody say, I've heard people say, like, at a very super high-level generic about AI, like, an AI is only as good as its model. It's only as good as it's— as the data that goes into it. Is that a true statement or is that too much of a generalization or how do you— what's your reaction to that?

13:08Rob van der VeerNo, this is totally true. The data determines the behavior, but the results can be amazing. I mean, look at ChatGPT. It's almost non-intuitive what can, you know, emergent behavior can come out of finding patterns in data in such a way. The machine learning is very powerful, but there's an example, and I use it in my presentation at the OWASP event in Dublin as well, where there's a couple of pictures of wolves and dogs, and they are used as training data for this machine learning model. And the machine learning model does it really, really well. But when they analyzed it, it turned out it recognized the wolves because all the wolf pictures had snow in them. They were taken, you know, in the wild. And in this case, it was all in the snow. Really smart from the model. But when you put another picture of a wolf not in the snow, it wouldn't recognize it. So indeed, the training data determines everything. And the training data is also one of the sources of manipulation when you talk about security and people get access to this training data. They can change it. This was one of the threads I highlighted, uh, 2 weeks ago when I was at the military AI conference in, in the Netherlands, that if you have an AI system that has learned how to recognize the enemy, for example, but you managed to manipulate that training dataset and you managed to put in examples of certain visuals and said this is the enemy And you could represent that visual in a sticker, and you would put the sticker on a building with civilians or whatever. The autonomous weapon system with that AI would recognize that building as the enemy and perhaps engage. Of course, there's not yet a clear view on whether you should be able to do this at all, but it illustrates how important indeed the data is. how big of a threat it is because it's hard to detect these kind of things.

15:19Chris RomeoSo, you know, excuse my naivety here, but from when we're talking about a training set or training data that we're using to train a model, is this primarily image files that are being fed into to train this model, or what other data sources go into training a model?

15:40Rob van der VeerI use image, uh, images as examples in my slides and in the guide because they look nicer, uh, but, um, I think in most situations it's, it's not images. Of course, there's, there's facial recognition and there's, uh, self-driving cars with cameras. I mean, that's, that's a whole industry. But think of ChatGPT based on, on, on natural language processing, so it works with a lot of text. Think of structured data. People who do loan applications or people who apply for a job, you would then need a dataset with job applications that were either successful or unsuccessful and require structured data. So these are the main categories, I would say, of, of input data, perhaps equally distributed in, in the instances of, of AI in the world. There are exceptions. There's also DNA has a different data structure. It's almost like an image. Sound, obviously. But yeah, those are the main categories.

16:48Chris RomeoThat's helpful. So, I had another question. As you were talking about the AI engineering and how that differs from software engineering, you mentioned the ISO standard. At the— with ISO or with anything that you've seen, has there been any greater emphasis on safety placed into the idea of AI engineering? Because, and, you know, when you think about, like, airplanes, you know, here in the United States, we have the FAA, Federal Aviation Administration. They focus, they have rules and regulations and all of these things, like, so that when I go up in an airplane, I'm almost 100% sure that I'm going to land at my destination and there's not going to be any big problems because safety has become the number one priority. Have you seen anything in the world of your travels of AI engineering? Is safety being prioritized anywhere or is it just not even being considered?

17:41Rob van der VeerOh yes, it's extra important for AI because as you know, AI, well, relatively speaking, integrates more with the real world than other technology because of its ability to have sensory capabilities, uh, to, to detect things from, from sensors and make a conclusion about it. Also the speed of its decisions, which makes it interesting to, you know, to use it in, in self-driving, uh, self-driving cars. And that's where safety, uh, comes in. So the trustworthiness of AI systems because of this is a very big theme in, in, uh, uh, well, in research, but also at, uh, at ISO because of this influence. For application security specialists, of course, this ties into how you deal with the risks. You notice a certain vulnerability in another software system, a typical software system, it wouldn't be that bad. But in case of an AI system that integrates with the real world, Yeah, that risk could be a lot bigger and that vulnerability could be much more important. Then again, safety in itself is a whole field. Of course, if you are a software engineer or application security specialist working on a system that integrates with the real world, you need to understand some of it, but enough to be able to So, what are some of the typical software engineering pitfalls then for AI engineers? What we notice is that AI engineers, or let's make it specific, data scientists, people who, you know, mold the data and collect the data in all kinds of forms and use the platforms that we have today, to create models is that they're really coding for the now. They have been educated and they, they are incentivized to build models that work, whereas you could say that software engineers in general are more focused, relatively speaking, on building systems that are maintainable, uh, for the future and make sure that you're still able to read it and that other people are able to read it and understand it, that there's unit testing etc. Those best practices we often see lacking in the code of the data scientists produce. Now, data scientists, I, I admire them. I, I was, uh, I'm still one, uh, myself. There's each role in the, in, in a team has its profile. That's why it's important to mix these profiles in the team, to have experienced software engineers together with data scientists, to educate data scientists on writing maintainable code, to use tools that provide measurements and findings with regards to this maintainability, help them to, to create units, unit tests. Even the notion of abstractions, or creating a function that contains a piece of logic that you reuse, is sometimes, uh, unfamiliar to, uh, to data scientists that we work So that's, that's a really big, uh, big pitfall there. Um, so not coding for the future, you could call it. Uh, other engineering best practices we see lacking sometimes, like versioning, documentation. What we almost always see when we do due diligence is the lack of documentation of experiments. These teams have been working for, for many, many years trying all kinds of things And there's a system there, but how it came into existence is not really clear. Everything they tried, everything that failed is typically not very well documented. And when another team needs to take over or somebody goes for another job, you have a problem there. And this is a typical problem in software engineering, as you may recognize. Especially with data science. Now, talking about application security, this is often overlooked. The specific risks that there are, like I said, the increased data attack surface that there is, the fact that people are working with real data, the privacy risks of being really enthusiastic with all that data, Not being aware that, well, certain data should be removed after some time. Certain things that you're doing are for a different purpose, which is a problem. So the lack of insight into these privacy boundaries is a problem in development teams. And then finally, the typical model attacks that can take place, like poisoning the dataset, like reverse engineering a model to see if you can find any personal data that was in the training set by using the model. There are all kinds of tricks for that. Data scientists are not always aware with those, which is also one of the reasons why we wrote the guide.

23:05Chris RomeoYeah, I took a journey, I guess, into the world of data science and created some content for Security Journey about how to, you know, use data science from an application security perspective, the R programming language, all these things. And, you know, what I realized is there is a— for me, my eyes were opened to the fact that there's a giant threat landscape that exists just in the data science world, just in the amount of data that data scientists have access to, like that includes personal data. Because, you know, I know you mentioned something about anonymization in the guide, but, you know, there's just, for a model to be good and put out the right results, sometimes you need perspective of real people, you know, going into the dataset. It just opened my eyes to the fact that there's, the threat landscape is so large. And a lot of this is things we haven't really thought about as AppSec people.

24:08Rob van der VeerYes, indeed. And because I've been in the AI industry so long, I've seen a shift where initially in the early years, we were just playing along with data as much as we could. And then we started to see some of the negative side effects. And we witnessed a growing realization until where we are today, where we're increasingly careful with what we do with data. And back then, the phrase was coined, data is the new oil. But we came a little bit back from that, didn't we? And of course, maybe to illustrate it, I often say to developers, you should regard personal data as radioactive gold. So, yes, it's very valuable.

24:59Chris RomeoMm-hmm.

24:59Rob van der VeerBut once you start spreading it around and leaving it all over the place, you need to be careful with it. So indeed, going to do data science as you describe is a slippery slope.

25:13Chris RomeoI'm gonna borrow that, data is radioactive gold. I'll give you credit for it the first 3 times I use it, but after that, you know, forget about it.

25:22Rob van der VeerBut yeah, so this is about personal data obviously, but confidential data in general, yes.

25:29Chris RomeoSo looking back at the guide, the scope of the guide, how does it relate to AppSec? How would somebody who is looking for AppSec help, would the guide help them?

25:41Rob van der VeerYes, the guide focuses on application security of AI systems, but indeed you're right, AI has different roles when it comes to application security. For example, if you're building a system, you can use AI to help you write code like ChatGPT does or Copilot does. We don't cover that. We can— I guess we can do a whole episode on that. It has its own risks and advantages, and it's a very current theme, a very exciting theme as well. Then there is the verification. So using AI to verify the security of a system, you can do it in multiple ways. You see increasingly that vendors use AI or say they use AI, doesn't really matter. It's about using smart software to identify patterns, to, for example, reduce the number of false positives or to correlate static or dynamic analysis findings. That's an example of using AI for verification. Then there is using AI for defense, typically trying to identify certain patterns in behavior, typically network behavior, that are either outliers or really specific for certain types of attacks, doing pattern recognition. And it can be a neural network, but can be a set of rules, can be regression. I would say it's all AI. It's all smart algorithms trying to detect this dangerous, potentially dangerous behavior. And then last but not least, using AI for attacks. There are several ways to do this. Some publications are about using reinforcement learning to try all types of attacks to see which approaches are most successful. That's a relatively new field. And there's also examples of AI writing phishing emails and personalizing them with ChatGPT and similar language models. It's become easier to make those types of, let's say, red teaming efforts much easier and more effective, I'd say. But we focus, as said, on how do you design, build, test, and procure AI systems so that they are secure and privacy-preserving.

28:15Chris RomeoSo let's jump into the guide now. And we've kind of talked about one of the threats that you describe in the guide. It's kind of been interwoven into our conversation just 'cause I was asking questions about it, trying to understand it. So I'd love to pull this thread all the way through So this is the— under the AI model attacks in the OWASP AI Security and Privacy Guide, you have an attack called data poisoning adversarial attack or changing training data. Behavior of the model can be manipulated. I'd love to just get a little deeper understanding for how that actually happens because, as I said, this is one of the things in this guide that fascinated me the most that an attacker could somehow manipulate data into a model that would get reflected into another system somewhere downstream, introducing a vulnerability. Like, that type of transitive vulnerability is fascinating to me. So, I'd love to— if you could just give us a short lecture on that particular attack, I'd love to understand it better.

29:18Rob van der VeerI share that fascination because the behavior of the model can indeed be malicious because of an attacker changing some of the content of the training set. And it would be hard to detect this behavior. Of course, that's part of the whole way the manipulation takes place. And I'll give you an example shortly how that works. It's hard to detect also from the logic, because if you look at a neural network, for example, it's just a bunch of weights. You can't read how it behaves in different situations. You have to test, which is why testing is so super important. But also really protecting that data pipeline, make sure that this data poisoning cannot happen. The example we have in the guide is about a self-driving car being able to detect traffic signs. This is important because at a stop sign, you stop, and if it's a 35 miles an hour sign, you drive, but you don't go any faster than 35 miles. Being able to recognize them is important. If you have a training set of pictures of traffic signs with labels, so this is a stop sign, this is a 35-mile sign, et cetera, like thousands of examples, you present that to a neural network or any other image recognition machine learning model, and it learns the particular visual cues, and it will be able to identify them in the real world based on camera images of the self-driving car. If an attacker manages to hack into that training set and adds a number of pictures of stop signs with little yellow stickers on it saying that these are 35-mile-an-hour signs, you have now implanted actually sort of a worm or Trojan horse into the behavior of the system. Because if you go outside and this this thing is in, uh, in, in the wild, is in self-driving cars. You can put a yellow sticker on the stop sign, nobody will sort of find that really weird, and just, you know, yellow sticker, I've seen worse things. And testing, uh, this, this whole algorithm, uh, by letting the car drive through a regular neighborhood, it will work. It will recognize the stop signs because they were still there in the training set. But because of these poisoned examples, when it goes to approach the crossroad with the stop sign with the yellow sticker, it will keep on, keep on driving. So that's an example of, you know, the dangers of AI interfacing with the real world, the importance of human oversight, the importance of other types of oversight, algorithmic oversight. So a secondary system that does a couple of extra checks. self-driving cars, they have those, but it's indeed an intriguing way to attack a machine learning model.

32:19Chris RomeoYeah, especially when you're thinking, you know, you mentioned earlier that the, you know, we're not going to pick on any particular self-driving car company. Let's just speak of them generically. But the fact that these cars are taking in images, and my guess is those images are being fed back into the model at some point. Maybe they're going through some type of a filter, they're going through some type of a check, but doesn't the self-driving car collect all the data from all the self-driving cars from that particular company that are on the road? They're collecting all that data and siphoning that back towards the model, right? And so there's an— so, because like when you think about what's the interface for an attacker to poison a model? Well, they might have to break into the the building where the model is being processed. They might have to gain access over the internet. But, but can you— could you manipulate the model through the real world, through the valid sensors that the cars are having and maybe feeding back?

33:18Rob van der VeerYes, you could, but you would also need to influence the labels. So the training set contains images but also labels, and that's, that's relatively hard. Um, however, if the training set is based on things on the internet and there's no labeling required, then things can go really haywire. I have an example in my talk about Microsoft chatbot Tay many years ago. Many people know this story where just overnight Tay became super racist based on learning from conversations with people on Twitter. I don't think it's good advice to base all your learnings on what people say on Twitter. And so the same goes, I guess, for a machine learning model. So the way that training data travels into your training set, that's part of your pipeline. And if it's from the internet, it requires really, yeah, smart curation. And perhaps you don't want to do it at all and build your own training set and label them yourself.

34:29Chris RomeoWhich is why I think, you know, some examples I've seen lately that that's where they're doing the dataset is closed. So it's not something that can be influenced from the outside because they want to have some control over that. But then the byproduct, I guess, is then they say, well, it's still missing things. You know, it's still not quite quote unquote smart enough, you know, yet. And so there's some trade-offs, right? Always gonna be trade-offs in trying to get the right data. Where do you get it? And making sure it's, solid for what you're trying to do.

35:02Rob van der VeerIndeed.

35:04Chris RomeoSo thinking about organizations who are looking to build an AI system and the development teams that are doing that as well, how would they safeguard security for those systems?

35:19Rob van der VeerThey already have a security program. Hopefully. I mean, that's the situation nowadays. And also something that's different between now and back when I started, which is good, I guess. We're making progress. The key is to involve AI into that. So your code reviews, your awareness training, your technical training, requirements, risk analysis, static analysis, pen testing, make your AI part of that. In many situations, it is not. Because it has a certain place in a sort of a lab. This is all fun, but the moment the lab produces something that the business requires worthwhile, it needs to go live real, real quick, and then you're in trouble. So start early with involving all these best practices already into AI, and not just security, as, as mentioned earlier, also the other software engineering best practices like, like unit testing and, and, and documentation. Then of course, really protect that data pipeline and make sure that access control and encryption and hardening is in place for that, for that data. If you're worried about data poisoning, poisoning, you can also include data audits as part of this. So on a regular basis, that you'll do visual or algorithmic checks on that data to see if it's still uncompromised. Then people need to learn to get to know AI model attacks like data poisoning, input manipulation, membership inference. Some of these things are really surprising and important for engineers to know that they exist so that they can take some of the countermeasures, like for example, have some rate limiting on an API on your model. If you provide your model and people can play with it and there's no real limit, they can present many, many examples to it and perhaps learn how to abuse it or reconstruct some of the training data that might contain confidential data. Understanding these sometimes surprising attacks is really important for organizations. So this was the security part, and then there's the privacy part, and that has a role already, uh, before you start. So if you want to do something with AI, um, really think, do I want to do this because I want to achieve something, or do I, I, I, I just want to do AI, period, or I want to gain experience? Then look at What's the purpose that you're going to apply this for? And is this legal? The data that you collected it for, does it have a similar purpose as what you're going to use it for in the model? Do the users know what are the agreements with— if it's personal data— with the people that actually own the data, which is the individuals themselves? And there's a whole range of privacy aspects that we've discussed in the guide. From fairness, transparency, purposeful data limitation that you need to take into account. And this is gaining, thankfully, increasing attention because there's so many examples of initiatives, also government initiatives with good intentions that needed to be canceled because these things were not arranged properly. And I think that about sums it up.

38:59Chris RomeoYeah, excellent. So, I certainly learned a lot through this short process already. I'm going to go back and really scour the guide and read this a lot closer. I kind of read it at a high level, but I feel like I understand some things better now to prepare me to read through in more depth. So, Rob, when you think about a call to action, for our audience? What, what do you want to, what do you want to call them to do here as a result of our conversation?

39:30Rob van der VeerWell, the great thing about your audience is that people have a lot of knowledge and insights into these aspects, and I would really welcome the input, even if it's a question or pointing out something that's not clear. I would welcome people to do it. I mean, it's currently the Spotlight project. So if you go to owasp.org, you'll find the project. There's a repo there. You can submit issues, you can submit pull requests. I've had great help from Engin Bozdağ, who is the lead privacy architect at Uber. And Uber, as you may guess and know, is dealing a lot with privacy matters around algorithms. They're really professional about these things.

40:17Chris RomeoYeah.

40:17Rob van der VeerAnd Engin is a thought leader in his field, and he joined in and did a lot of good work on, on the privacy part. And we need that because this field is evolving. There are many ways to look at it. Some, some information is probably missing. We like to hear about it. So that's my call to the community to join in. And the least thing you can do is, is write me an email with, with some of your observations that could perhaps improve the guide. And next is go have fun with AI. Find some tutorials out there that, you know, provide you with an environment in which you can build a model just to get acquainted with it, just to, you know, have done it once and see how expressive the frameworks and the platforms of today are. Because doing hands-on work really helps you understand what this is all about. And then, of course, read the guide, spread the word if you like it, share it on Twitter or LinkedIn or to your friends, and don't be overwhelmed. I think that's the key here. I think the same goes for application security in general. The key is to start small. And not too big as an organization. So do it in the lab, but the moment that you start doing it, start doing it as if you were doing it in, uh, in production. Start treating it as professional software engineering. Learn about the particularities. Use the guides. Have a look at ISO 5338 and the other ISO standards that are coming up, and some of the links that we have in the guide. Work from Etsy. Work from Enisa, that can be really helpful, but you don't need to understand everything just to get started. So don't be overwhelmed, make it small and let it grow, but take it seriously.

42:16Chris RomeoVery good. So, Rob, thank you for sharing this knowledge, for leading this project. I know you put a lot of effort into the project itself, and thank you for sharing that with our community. This is certainly a cutting-edge, if not bleeding-edge, issue that we as application security people need to understand and be ready to help our teams and our organizations deal with this properly and, and have a secure artificial intelligence future. So thanks for being a guest on the show. We look forward to a future conversation where you can teach me more about artificial intelligence.

42:50Rob van der VeerIt was an honor. Thank you both.

More on AI and LLM Security