--- title: "Kristen Tan and Vaibhav Garg -- Machine Assisted Threat Modeling" url: https://appsecpodcast.com/kristen-tan-and-vaibhav-garg-machine-assisted-threat-modeling/ date: 2022-05-10 duration_seconds: 2777 guests: ["Kristen Tan", "Vaibhav Garg"] topics: ["Threat Modeling", "Privacy and Compliance"] audio: https://www.buzzsprout.com/1730684/episodes/10593928-kristen-tan-and-vaibhav-garg-machine-assisted-threat-modeling.mp3 video: https://www.youtube.com/watch?v=ji_WOuV-iEM transcript: true --- # Kristen Tan and Vaibhav Garg -- Machine Assisted Threat Modeling *May 10, 2022 · 46 min* with [Kristen Tan](https://appsecpodcast.com/guests/kristen-tan/), [Vaibhav Garg](https://appsecpodcast.com/guests/vaibhav-garg/) on [Threat Modeling](https://appsecpodcast.com/topics/threat-modeling/), [Privacy and Compliance](https://appsecpodcast.com/topics/privacy-and-compliance/) [Audio](https://www.buzzsprout.com/1730684/episodes/10593928-kristen-tan-and-vaibhav-garg-machine-assisted-threat-modeling.mp3) · [Video](https://www.youtube.com/watch?v=ji_WOuV-iEM) ## Show notes Can machines make threat modeling faster without stripping away the judgment that makes it useful? Kristen Tan and Vaibhav Garg join Chris and Robert to discuss their analysis of open source automated threat modeling tools and what it reveals about automation, extensibility, security, and privacy. They explain why they studied the available tools, how they evaluated them, and where machine assistance can support rather than replace human reasoning. The conversation covers developer-friendly formats, privacy requirements, organizational fit, and the gap between generating threats and helping real working teams make better design decisions throughout the modern software development lifecycle in practice. The Application Security Podcast is brought to you by [Security Journey](https://www.securityjourney.com/). About Security Journey Security Journey provides application security education for developers and everyone in the software development lifecycle. → [Learn more about Security Journey](https://www.securityjourney.com/) Connect with Kristen Tan and Vaibhav Garg: → [Kristen Tan on LinkedIn](https://www.linkedin.com/in/kristentan1) → [Vaibhav Garg on LinkedIn](https://www.linkedin.com/in/gargvaibhav) Mentioned in this episode: → [Analysis of open source automated threat modeling tools](https://www.usenix.org/publications/loginonline/analysis-open-source-automated-threat-modeling-tools-and-their) → [OWASP pytm](https://github.com/OWASP/pytm) Chapters: 00:00 Machine-assisted threat modeling 01:44 Kristen Tan’s path to threat modeling research 03:29 Vaibhav Garg’s security and privacy background 06:49 Why analyze automated threat modeling tools 09:18 Security and privacy as connected disciplines 12:17 What motivated the research 14:37 Can machines automate threat modeling? 20:45 How the tools were evaluated 24:24 Open source tools and extensibility 28:59 Evaluation criteria and research results 35:14 Choosing a tool that fits the organization 37:00 YAML, developers, and usable workflows 39:00 Keeping threat modeling user-centered 45:00 Final takeaways ## Transcript *8,035 words · assemblyai* **0:00 Chris Romeo:** On this episode of the Application Security Podcast, we talked to Kristen Tan and VG from Comcast. They wrote a paper called An Analysis of Open Source Automated Threat Modeling Tools and Their Extensibility from Security into Privacy. When I first found this paper browsing online, I stopped reading and said, I want to interview these people and, and hear their direct story for what they were doing, why they did what they did. And so we hope you enjoy this conversation with Kristen Tan and VG. **0:26** Hi, everyone. **0:27 Chris Romeo:** Tristan Tan and VG. **0:29** You're about to listen to AppSec Podcast. When you're done with this, be sure to check out our other show, High Five. **0:36 Chris Romeo:** Hey folks, welcome to another episode of the Application Security Podcast. This is Chris Romeo, CEO of Security Journey and co-host of the podcast. I'm also joined by my other co-host, Robert Hurlbut. Hey Robert. **0:51** Hey Chris. Yeah, good to be here. Uh, Robert Hurlbut. I'm a principal application security architect at Acquia. and focused on threat modeling. And I am really excited that today is that topic on threat modeling. **1:04 Chris Romeo:** Yeah, because I was scrolling, I don't know, somewhere, I couldn't even remember where. Maybe it was the OWASP threat modeling channel, maybe it was somewhere else. And I happened upon this paper that was talking about, of all things, threat modeling, the number one topic on the Application Security Podcast. But Before we get there, I do want to hear the security origin stories of our guests. And so, Kristen Tan, we're going to start with you. If you can just tell us your security origin story. If there was a comic book and there was an episode 1 or, you know, series 1 of your comic book, what would the origin story for security sound like? **1:44** Sure, I can try and do that. I don't think it'll be as entertaining as any comic book, but I got my start in just the tech world in high school. So I had a couple computer science classes, but then I went to college for software engineering. So definitely was exposed to secure design principles and had those ideas in mind as I was, you know, writing code and starting to develop my first applications. And then I got my master's in computer science. So I still hadn't really figured out what the entire breadth of security was. And then I started working at Comcast NBCUniversal. So I'm in a program called the Core Tech Program, and it's a rotational program that allows you to jump from one area of the business to another over a course of 2 years to get exposed to different areas of tech so that you can try and figure out what it is you're interested in. So I got my start there in production infrastructure, and then only after a year of working full-time did I finally sort of stumble upon security. So that's when I had my rotation on VG's team, which was focused on cyber R&D, and in particular, My project focused on privacy threat modeling. So I sort of got exposed to privacy engineering at the start, and then from there found my way into threat modeling and then put the two together. So that's kind of how I got to where I am now. **3:02 Chris Romeo:** Ah, very cool. **3:04** Very cool. **3:04 Chris Romeo:** So that might be the first origin story that came directly into threat modeling that we've ever heard on the Application Security Podcast after, you know, almost 200 episodes. So you have that little, That asterisk next to your story there. **3:18** I'm a little old-fashioned, I guess, right? **3:20 Chris Romeo:** Yeah, definitely. Very cool. Very cool. Well, VG, Kristen kind of gave a little introduction. She mentioned you in her intro, but can you share your security origin story with us? **3:29** Yeah. So I did have an answer planned, and then you told me about your favorite security origin story. So I'm thinking of going as far back as, you know, when I was 4, but I'll... What I'd say is, you know, I've sort of always been interested in security. You know, when I used to read a lot when I was a child, and there's this one particular book which had a lot of sort of spy mysteries in Hindi, and like all of the spies always had to solve, you know, some kind of an encrypted message. And so that was always sort of, you know, in the back of my mind. And then when I went for undergrad, There was a course called Security Protocols. It was taught by a professor who was notorious for giving bad grades. So famously, he gave like one A in a class of 240 students. So I wasn't particularly tempted to take this class, but then turned out my grades would only allow me to take his class. So I signed up and As it turned out, I was, you know, while I was not always great at everything academic, that one particular class I was very, very good at. And then I ended up doing a summer research internship with him where— so I got a paper out of that. And then I was like, okay, this is good. I like this. I enjoy it. This is what I want to do for my career. And so got advice from him and And then ended up doing my master's and PhD in the same domain. And that's kind of how it went. And then after my PhD, I did some doctoral work at Drexel, and then I moved to the industry and I worked for Visa and did a bit of work in how— and so my sort of expertise is in How do you communicate complex risks to non-experts? So in particular, my doctoral work was on how do you communicate security risks to people over 65? So that sort of translated into a career into security awareness at Visa. And then from there, I came to Comcast where I was originally hired to do cybersecurity policy, cybersecurity public policy work. And then slowly as I've been here, I've been doing more cybersecurity R&D work, and so building tooling for internal customers. And last year, one of the things that we were asked to do was to look into threat modeling, in particular, how can we do privacy threat modeling? And that was actually my introduction to your podcast because I came across the Lindem framework And I was like, okay, I have to— so I heard a fair few talks by Dr. Kim. And then I was like, okay, let's see what else is available. Where else has she spoken? And your podcast was one of them. And I started going through the list. I was like, oh, Adam Shostak has been a guest as well. So I went through a few of those podcasts in sequence. But yeah, that's sort of my story into this. into this journey. **6:49 Chris Romeo:** Very cool. Yeah, Dr. Vutz is a good friend of the podcast. And we, Robert and I worked with her on the Threat Modeling Manifesto as well. So yeah, she's just, she's, she's awesome. That's, that's how we'll describe, describe her and your, your, your graduate, what your work on kind of the explaining risk to people over the age of 65. I'm fascinated. And I'm tempted to ask a bunch of questions about that. But I'll save that for a future conversation someday. Because that's, that's something that just where you were doing that kind of from Visa's perspective must have been fascinating to say, how can we help make people more secure? But we're going to talk about threat modeling though, right, Robert? **7:28** Absolutely. Okay, let's just jump into the project, the paper that was written. So the title that I see is An Analysis of Open-Source Automated Threat Modeling Tools and Their Extensibility from Security into Privacy. Great title. Looks like it came out February 17th, 2022. So fairly recently this year. So one great question for you, Kristen, I wanted to ask is, please explain the project, but also who was involved and what was the goal in writing this paper? **8:03** Yeah, absolutely. So when I joined Vijay's team, we started to look into expanding our existing threat modeling process. So there was a threat modeling process in place, but we were looking to specifically add privacy engineering principles to that. And as a result of that, the threat modeling architects would be essentially doubling their workload, but that was not something that we were looking to ask them to do. So what this project was, was to build out an end-to-end machine-assisted threat modeling solution such that we wouldn't be introducing essentially a 200% workload. So my subpiece of the project was to take a look at existing machine-assisted threat modeling solutions to try to understand if there was anything out there that we could use rather than having to build something from scratch. So I took the lead on this project with guidance from VG, and our goal was to find a solution that would be easy to integrate with the existing threat modeling process and something that would be extensible for privacy and that would be fairly easy for users to get onboarded to such that we weren't putting them through an extensive training process. So I guess that kind of summarizes where we were going with it and why we wanted to do it in the first place. I don't know if there's anything you want more detail on, but I think that's sort of the overall picture. **9:18 Chris Romeo:** So you separated security from— so this is specifically privacy? **9:24** The tool we wanted should have the capability to do machine-assisted threat modeling for both security and privacy. So we have existing security threat modeling in place, process around it. And we were looking to add on privacy threat modeling to that. So we wanted a tool that would support both of those functionalities. **9:42 Chris Romeo:** Okay. **9:43** And if I can add just a couple of pieces to that. So as part of our security threat modeling process, we do ask questions about privacy as well, about privacy threats as well. But what we wanted to do was have something that would be a little bit more formalized. So, for example, right, like you have STRIDE for security, right? Like it's a relatively formal sort of framework and structure. And we were like, okay, well, what can we do for privacy that would be sort of similarly structured? And so, you can say that, you know, I have sort of comprehensive coverage across a class of threats, right? So, that's where we started looking into, and that's where we sort of came across Linden. as the framework that we could try and adopt. But then the question was, okay, well, it's taking us 8 hours to do a security threat model. Now we add another formal framework, that's another 8 hours of work. And it's not just work for the threat modelers, it's also work for the developers. And I think that was the bigger challenge for us because the number of asks from the development teams from security and privacy are increasing every day. And legitimately, they feel like they have to sort of redo a lot of stuff a lot of times. So one of the things we wanted to see is like, if we can use something that would be machine-assisted, where we can pull data from things that we have already done, because we've already collected a ton of information from from these development teams about their technology stacks and things like that, why don't we just use that existing information and try to build something that is machine-assisted? So one of the reasons we wanted to look at this tooling is we wanted to see, oh, is there some tool that already implements Linden? And then, and that That was sort of like our first, you know, one of the first things that we started looking into. But as we did that, we were like, okay, well, we actually need to come up with, you know, some sort of way of assessing these tools so that we can bake off each of those. So yeah, I mean, Kristen very much led that effort and that we figured it would make sense to, you know, share our findings and that way people don't have to replicate what we have done. **12:20 Chris Romeo:** So when we think about motivation then for going down this road and considering this effort, I'm gonna— I'm imagining that some of the motivation is to save time because I've already heard you've already described that a little bit. And I could see how if you separated security and privacy threat modeling, and like you mentioned, Vijay, that like, you know, if it's 8 hours for one, 8 hours for the other, now you're asking your threat modeling architects to do twice as much work. You're asking developers to pour through more findings and And they may even have some duplication of effort happening across those 2 things. So I'm guessing motivation was a little bit from a resource perspective, but what else would you say were the kind of the most important things from the motivation perspective? **13:07** So there were a few different things. One was, you know, when we do sort of manual threat modeling, right? There is a certain, inevitably, there's a certain amount of your personality that comes into play when you're doing a manual threat model, right? Every threat modeler has a particular technical bias, right? Like, they have a technical history, so they may focus on one thing over another, right? They might find one thing to be a little bit more concerning over another. The benefit that we thought we could get out of a machine-assisted process is that we have a well-defined class of threats. And every time we put an application through that tool, we know that we are getting coverage across all of those threats. So nothing is getting missed. And I apologize, I've been trying to think of a more polite analogy, but the one I can only think of is it's sort of like, you know, having a— firing a shotgun into the dark. right? And it's a scattershot. And then you see where you hear a thud, and where you hear a thud, you can go digging deeper, right? So that was sort of the idea behind using that machine-assisted tool. So it's not just saving time, but it's also providing greater threat coverage and having greater assurance in your threat modeling process. **14:37 Chris Romeo:** So I want to ask a philosophical follow-up question. As someone, you know, Robert and I are both people that have taught threat modeling to a lot of engineers and developers across the industry. And when I think about this, the idea of machine learning and using that machine-assisted kind of approach to threat modeling, how do you ensure that you don't get too stagnated in that approach though? Because some of the best threats that I've seen come out from developers, or when they'll say something and I'm like, wow, that's a cool idea. I never thought of that. It happens to me every time I do a threat modeling workshop in a 2-hour period. Someone will say something during that time that I've never heard before. I'm like, wow, cool. How do you still— how do you balance that against machine-assisted where it could be not being— taking out some of that creativity? **15:37** Yeah. So that sort of like operational pipeline is something that, you know, we have continued to work on because we're still in a POC stage in terms of what we are building with the tooling. **15:50** Okay. **15:51** But so one, there's no machine learning involved here. So it is, and I very carefully chose the word machine-assisted and not use the word automated because What we found is that that human in the loop is quite important. So I'll give you an example of where that human in the loop can be important. And again, I'm going to use an analogy. So if you were threat modeling, for example, a home, right? How can you break into the home, right? And you could say, oh, I can break in through a window. I can break in through a door. I can break in through the wall, right? Let's say 3 particular threats. Now, a machine will give you that, right? There are these 3 different threats. But that's where the threat modeling architect is going to bring in the specific knowledge about the technology stacks at Comcast, right? So if you're doing this analysis in India, you're not— your wall is not going to be a big concern because homes in India are built out of brick. You're not breaking through that. Whereas in America, you have basically stick houses, right? All you have there is some vinyl and insulation. So, actually, your door is one of the strongest parts of the house. If you want to break in, what you really want to do is break in through the wall. But that's something that you only know if you know how the house is constructed. So, I think that's where the threat modeling architect's intuition and knowledge comes into play. Similarly, the threat modeling architect has a high-level knowledge of the technology stack, but then the developer or the product people have the deep knowledge about their technology stack. So they are pretty actively involved in this threat modeling process. We engage with those developers in multiple different ways to keep that knowledge up to date. like the thought about thinking about security and privacy as they're designing these systems very much upfront. So, I completely agree. The threat modeling session is still there. It's not going away. The only thing the machine-assisted tool is doing is that it's allowing you to go through a lot of kind of high, you know, a lot of kind of common threats relatively quickly. So then you can focus on the ones that are going to be unique for your particular application. Or it's, you know, in privacy, it's not necessarily always the application, right? It's sometimes the customer, it's sometimes the context of deployment. So it's allowing you to focus on the thing that you wouldn't necessarily always be able to focus on if you had to spend 16 hours doing the threat model. **18:46 Chris Romeo:** Mm-hmm. **18:48** Yeah, and I'll— If I could add— **18:49 Chris Romeo:** Please, yeah. **18:50** I was just gonna— if I could add one thing to the end of what Vijay was just saying about sort of focusing on all of this, all of these high-level threats that a human might not have time to focus on, I think this sort of lends itself to a comparison to the Lindon Go methodology where the cards are really intended to guide the conversation. So there's no way that you're going to have the same conversation every single time, but you have this set of guidelines that shape what you're talking about. So almost like a checklist, how checklists are intended to make sure that you go through all of the important things before you take flight, for instance. So I think that what we're doing here is just to help the threat modeling architects guide their threat models such that they have touchpoints of the most important issues that we saw, but they obviously have the human knowledge that we couldn't bring to the table. **19:36 Chris Romeo:** Yeah, that's one of the important parts. And I'll share just very quickly, You know, I just mentioned doing these virtual— I do these virtual threat modeling sessions all the time, and somebody always says— somebody always gives me an answer that I've never heard before. And I use that same example, Vijay, that you were just talking about, you know, the house is the first way to help people transition into threat modeling. And, you know, how could you get into the house? And one time somebody says— because I always make it kind of a joke, I'm like, imagine the world's best chocolate cake sitting on a table inside this house, and you have to get— **20:08** Yeah. **20:08 Chris Romeo:** through— you have to use some of the interfaces that are available to you, and then they go through windows, doors. And then somebody in one of these sessions, they're like, I'll just knock on the door and ask them for a piece of cake. And I was like, I'd been doing this for years at that point. No one had ever come up with that idea. Just knock on the door and ask for it. Don't worry about breaking in. Why would you go to that level? So yeah, so I mean, that creativity and that creative burst is so important in this threat modeling process. And it sounds like that's where you're going— I like that machine-assisted. It's not machine learning, machine-assisted. Great terminology to use for that. **20:45** Definitely. So, Kristen, maybe you can walk us through— very curious about the methodology. You know, as you were thinking through and researching, what was the methodology you followed? **20:59** Right. So as we briefly touched on before, we wanted to look at existing solutions to see if there was anything that would prevent us from having to start from all the way at the bottom. So we started to look at just these existing tools that were out there. And initially, we were looking at proprietary tooling. But after just almost a few days to maybe a week of research, we quickly realized that these tools didn't necessarily allow us to add in custom threats that we might be looking for that might not already be accounted for. So We kind of switched gears then and began to look at open-source tools. And as with any open-source project, what we were really looking for was a community of support around that project such that if we needed to reach out, that support would be there, and also the appropriate licensing so that we would be able to modify these tools and use them for the purposes that we were hoping to use them for. And then once we had this list of tools together, we had to put together a list of evaluation criteria to understand whether they were actually applicable for our use case And I'll highlight now, I guess, that we were really looking at our use case of having this idea of putting together a privacy threat modeling process that augmented our existing security threat modeling process. So in other places, that might not look exactly the same because I doubt any 2 companies' security threat modeling process looks 100% the same. But for us, a couple of those evaluation criteria were A couple of things I guess that I'll talk about now. So we started with complexity of logic. So how complex of a threat could this tool detect with just the code? So once again, we're looking at humans intervening later, but what could the code do on its own without any human intervention? Then amenability to custom threats was the next big one. So like we talked about, we were looking at the specific privacy threats that we had identified as important to our security, or sorry, privacy threat modeling at Comcast. So we put together a library of custom threats from a couple of different sources. So we were looking at the Linden methodology. We were looking at the OWASP Top 10 privacy threats. And we were using all of these for inspiration for the privacy threats that we wanted to consider. Then, the next big one was operational usability. So this goes back to what I mentioned before about users getting easily onboarded. We wanted to make sure that this was a tool that people would be able to learn and would actually be willing to use. If there were extreme technical learning requirements on the front end of all of this, it wouldn't really be something that people would be willing to adopt. So we wanted to make sure that we were taking the end user into consideration there. And then the final 2 characteristics or evaluation criteria that we were looking at were use in security and use in privacy. So how could this tool, or could this tool detect security threats and categorize them using any framework? such as Stride. So what was the categorization around that? And then, in terms of privacy, because a lot of existing tooling only really looks at security threats, our main question was, can this tool look for privacy threats, and can it successfully detect those privacy threats? And if so, is there any categorization around that, or can we introduce that? So the potential that we could do a little work on our own and then have that functionality exist there. And With all of those criteria in mind, that's how we then went through and evaluated all of the tools and then came to a decision about what seemed to work best for our particular use case. **24:24** And then, if I can, so one quick thing to add there is one of the reasons we also wanted to go with open-source tooling is that, from our perspective, if you're going to do all of this legwork to add custom threats, we figured, you know, why not just have it be available for any other company who wanted to do something similar as well. So there are certainly proprietary tools where you can add custom threads, but then it's sort of like, you know, us giving them, you know, us doing free work for them. And from our perspective, we are like, okay, well, these nice people in the open source community have built similar tools that maybe are as good, Why don't we just pay them back and add some of our own work to theirs? It is something that we do quite a lot. So it's not just unique to this project. We do that with a lot of our projects. We try to focus on either if we are writing a custom tool ourselves, we try to open source it, or often we try to build upon existing open source solutions and then try and give them back. **25:40** That's very cool. **25:42** As much as we can. **25:43 Chris Romeo:** Yeah, very cool. That's your kind of using something open source and then being able to give back some of that threat content to the project is something that's huge. Now, I've been waiting for— go ahead, go ahead. **25:57** Sorry, one thing. And, you know, there are some advantages to this, right? So when I build— I say I, I mean, you know, people smarter than me like Kristen, When they write the logic for detecting, let's say, a privacy threat, right? If it's in an open-source tool, that code and that logic is available for other researchers to look at. And they can say, well, you've written in this way, that's lovely. But what if you did it slightly differently and you get slightly better hits? So, there's, you know, by contributing there, you know, we also get something back because people can improve upon what we've come up with because sometimes we may not have the best answer, although it's rare. **26:45** Yeah. **26:47 Chris Romeo:** Yeah, makes sense. I want to get to the results of the analysis, and I don't know if you're going to answer this for me, but I really want to know what's the best threat modeling tool on Earth or in the galaxy, we could even say. I don't know if you're going to get me there, but curious to know kind of the results of going through this methodology. And as I'm scanning the list, I think we've talked to almost everybody on this list has been on the podcast at one point or in the past. So yeah, walk us through kind of the results of this methodology. **27:22** Yeah, so, I mean, obviously the best threat modeling tool is the one that I'm building now, which is available. only proprietary and it's available for like $1 million a license. But aside from that, it was a little bit challenging to say, you know, what is the best threat modeling tool? So what we were really looking at, you know, what fits best in our sort of SDLC lifecycle, right? Like our secure development lifecycle, what aligns most with our existing security practices. what would be the kind of tool that our customers, whether they're developers or threat modeling architects, would like the most. So from that perspective, we found, you know, THAADgile to be the most useful. So one of the things is that THAADgile, the kind of information it collects It's not sort of your traditional data flow diagram information. It collects a lot of additional details about the different individual assets in your system architecture diagram. That just makes it a little bit richer for privacy threat modeling because it's collecting additional data points about your data assets, which other tools don't necessarily collect. So that was sort of one of the criteria. But Kristen has sort of more, has more to say on that because, yeah. **28:59** Yeah, sure. First of all, I hope that VG is planning to share his profits with me because I think that that would be only fair. **29:06 Chris Romeo:** Well, you announced it, you announced it here. So we get a cut of like 2% or something. That's how it works. **29:12** But in terms of the tools that we had the opportunity to analyze, As Yuji mentioned, right? So we identified Fragile as the candidate that seemed to work best for our use cases, but there was definitely a lot that went into that. So just to talk about some of the other tools that we looked at, the one that we had highlighted in the paper, Keras, that tool has a lot of features. And in our use case, we sort of felt that this complexity would bog us down in terms of just executing the threat model because there are a number of features that have functionality that isn't necessarily directly related to identifying threats. And in addition, the UI felt a bit confusing to us. So there were a number of these features laid out, and then I think there were maybe 7 at the top, and then of those, 6 had dropdowns that had a number of other options. And we felt a little bit of difficulty just navigating through that to get to the content that we were looking for. So in terms of that, we were looking for something that lent itself to a more simplistic approach. And then, another issue that we encountered— so as Vijay mentioned, we did a lot of work on understanding what information we needed to provide in a data flow diagram to identify threats. So in Keras, constructing that data flow diagram actually took place on a different screen from viewing that data flow diagram. And this just led to a little bit of confusion in terms of consumption of the information that we were looking at. So that was why we had a bit of difficulty with that. For other companies where you are looking for that extra functionality that's maybe not explicitly related to identifying specific threats, Chris could be a good candidate. It's really just dependent on what the end goal is there. And then, to touch more on Thragile, what was great about Thragile for us was not only already what Vijay talked about, but Thragile was also super amenable to integrating custom threats. So it's a Golang application. So if you're familiar with Go, you're already on your way to just being able to add threats to Fragile. And it's just composed of this library of files where each file could represent a single threat. So you're able to go ahead and just write your own Golang threats that then can be integrated into the logic that Fragile uses when it parses any input that you give to it. In addition to that, Fragile seemed to us more useful in terms of the ability to— sorry, I'm stumbling on my words here. Fragile seemed to us more useful in terms of the ability to as Vijay said, parse more specific information. So in terms of collecting data flow diagram information, we're not only looking at the components and their flows, but also specific attributes of those components and flows. So for instance, what type of encryption is being used at a particular component, or what protocol is used when data is being passed from component A to component B? So because of that specificity, Fragile is able to make decisions about more complex threats that you might not understand just from looking at a traditional data flow diagram. **32:15 Chris Romeo:** flow diagram. **32:17** To sort of wrap up the discussion on Fragile and look at maybe the flip side of it in terms of cases where maybe it's not the most ideal, like we just talked about, it does require a lot of information. So if you're the person sitting there filling out that input to pass into Fragile, you might get frustrated. But what we're working to do is come up with a system that uses existing application databases. So information that developers and engineers have already provided that lives somewhere in a system, that we're able to pull and then just automate the creation of the input that Fragile is looking for. So if there's the potential to do that at any given company, then this might still be a good candidate. But if those databases don't exist, then it might not be the best candidate for that company if the developers aren't really willing to sit there and fill out all that information that's needed. So this is all to say once again that it's very case-by-case basis. There's no, unfortunately, one best threat modeling tool in the galaxy as much as we might want there to be, but ideally, you know, we can each figure out what works best for us. **33:21** And I think the sort of main thing to think about is look at the evaluation criteria. So the evaluation criteria— so the criteria is listed in the paper, right? But what that criteria means in your particular operational context is going to be slightly different across organizations. So another example that we sort of didn't talk about is PyTM. which we found is as good as anything else. It certainly did everything that Fragile did, but it basically did not generate a list of threats against Stryd. Now, if you don't care whether your threat falls in— which threat falls in the Stryd category, then you probably do not care about that at all. If you do, if your main development is in Python, that's probably a good threat modeling tool for you. So for us, you know, one of the benefits of Thragile was that the input is in YAML. And we felt that our developers would rather probably write YAML files than fill out forms. So we'll learn that in time. And then finally, As Kristen was talking about the UI for Keras, so we kind of don't need a UI. We just need a rules engine because we have our own UI that we have already built out that people are already familiar with. There's already sort of like an end-to-end process around that. So from that perspective, we kind of, you know, that sort of feature rich UI that CAIRIS had was not particularly important to us. But that, again, might be important if you don't have that kind of UI already existing in your organization. **35:14 Chris Romeo:** So just to summarize kind of what I heard about the best threat modeling tool is really the best threat modeling tool for your organization. So what you provided in the paper is your set of requirements that made sense. And so if somebody else is out there, another listener's thinking, hey, I'm gonna try to find the perfect open source threat modeling tool for my company, they can look at your requirements and maybe that, maybe some of 'em are the same. They might have slightly diff— they might have something that they're, they want to add into it. And so where I can see this research moving forward is so they can use this as a starting point, make any adjustments for what's really important to them. Maybe somebody out there is like, oh, GUI, I gotta have a GUI. Because I don't have one. That would be another criteria that could be added here. And then they can build on top of your research to help them. But at the end of the day, they're finding the best tool for their organization, which you've shown us how you got to the best tool for your organization. Yeah. **36:15** And yep, that's pretty much it. **36:17** I was gonna say that the other thing I noticed is in going through, like you said, YAML and so on. So your focus was on visual versus non-visual. So you really— visual was not as important because of the fact you already have that covered. And so that's where I see a lot of the tools, especially the open source, they're either non-visual where you can, as a programmer, you can build YAML files or Python with PyTM. Or very visual, which, you know, really comes from the threat modeling tool originally. And then some of the commercial tools, they sort of follow that sort of approach as well. So, but very interesting. It does, like I said, it just depends on the needs of the company and what works best for them and what they're looking for. **37:08 Chris Romeo:** Yeah, and I want to pinpoint one other thing, Robert, that you just reminded me of, and that's, Vijay, is as I was listening to you talk about you know, using YAML files because that works best for the developers. I just want to highlight, like, that's one of the most important things that we can do as security professionals is not continue to dictate the way we think things should be. And that's a perfect example, Vijay, of where you and Kristen have said, hey, our developers, they want— they're used to working in YAML files. That's something that's going to make their life easier. Because so often, and I've been in security for 25 years, it's always been about how do we make our lives easier as security people, not how do we make our constituents' life easier. And in that case, you're making their life easier. It might make your life a little bit harder. That's okay because we're not the ones that are going to run these tools all day long. We're the ones that are building them and then pushing them out. **38:00** Yeah. I mean, for us, you know, we work in a customer-centric model. So, you know, we really think of ourselves as like a, mini sort of startup inside a company, everything we build, we then have to go find customers for it. So in this case, so we will have the option for people to fill out a form if that's what they want to do, but they can also fill out a YAML file. And ideally what we are trying to get to is that they don't have to fill out anything because we are automatically generating those inputs. And then we— and this might be getting a little bit ahead of the curve, but what we want to do is sort of create templates for them because, you know, the tech stack for, you know, classes of applications is relatively the same, right? So why don't I just give you a template? We already know most of the answers that you were going to answer anyways. And then in fact, if you're answering something differently, Then we can go in and ask the question, hey, why did you use this protocol instead of this other protocol that is always used in this tech stack? And, you know, understand why that decision was made and see if that introduced— that introduces any security risk into the equation. So what we are really trying to do is that for all the inputs in Fragile, we are trying to generate those automatically. And where we can generate those automatically, work off of templates. And then really as the last recourse, should the developer or the product owner have to provide additional information. But we are a little bit away from that. We are still in that process. There are more papers coming down the line and hopefully we'll have more updates then. **39:46** And just one more thing on the idea of really making sure that this is a customer-centric, user-centric experience. **39:54 Chris Romeo:** Yeah. **39:54** I just wanna, I guess, share a little story. When I first sort of stumbled upon Fragile and came to understand it myself, I was like, wow, this is great. Like, look at all these threats that I can get out of this. And like, then I think Vijay and I went and showed it to one of the threat modeling architects with whom we worked closely. And he said, no one's gonna wanna sit there and fill out that whole input form and fill out all that information. And all of a sudden it was like, wait, like maybe this isn't as great as we thought it was at first pass. So it definitely is still helpful. And I think that we see a lot of potential, but there's, There's still more work to be done. **40:24** So thank you both for joining us today. It's been really great just to talk about your work and to understand more about your approach and the problems you were trying to solve. I know it's going to help some other folks who are maybe thinking along the same lines and trying to think through these things as well. For our audience, could you give us a call to action? And also, what would you like for the audience to do? What are some things that you would like for them to do next after we've talked about this, maybe to go read the paper and so forth? What are your thoughts? **41:04** So, we, you know, as we're doing this work, you know, fortunately, you know, I have a lot of friends in other companies, and I've been talking to them and asking them, you know, how do you do threat modeling? And not a lot of people are using automated tools or machine-assisted tools. I apologize. And I think a lot of people have been burned by them a little bit, right? Like people probably tried them 5 years ago, 6 years ago, and they were not there just yet. But I do think the tools have matured. I certainly do think, you know, Fragile, does a very, very good job in terms of the complexity of logic that it finds the threats. So what you're not getting is, you know, a list of 15,000 threats for an application. So then you can't really work through those. It gives you a fairly, you know, narrow list of threats. So my first call to action would be, you know, for people to go and check these tools out. Out, even the ones that maybe didn't get the best rating in our paper. That's, you know, they still do, you know, from the actual threat modeling perspective, they are pretty good. And then see, you know, how you can, instead of thinking about, oh, I have to do all of this work to operationalize this, think about what can you engineer in the backend that you don't have to do all of that additional work. So, like I said, right, you have so many security tools, logging tools that are running in your environment. They are already collecting a lot of this data. Think about how you can parse that data and find the inputs for your threat modeling tools. That's the thing that we are working on next. And we are getting— it is a little bit slow going and it is a bit of work upfront, but that means that it's so much less work down the line. And that's part of this mantra, I guess, for a lack of a better word, that our chief security and chief security and— it's a long title— our CISO and product security officer and privacy engineering officer, So, she has this mantra of like, you know, moving, of shifting left. So, what we really want to do is shift that burden of that legwork to us, try and automate, you know, that input generation. And that way people don't have to do as much work down the line. I do think the tools now, if you can get that input, they do give you pretty meaningful output. **43:52** Yeah. **43:54** And at the very least, if you think of threat modeling as a journey, it gives you pretty good signposts to follow. You don't have to stay exactly on the path. You can go and wander around a little bit, but it gives you good signposts. And then I, like I said, I'm particularly biased towards open source projects. So as you're looking at these tools, check out the open source ones. If you If they're not there where you want them to be, you know, contribute back to them. Fix what you think is, you know, missing. And then we can all benefit from it. But yeah, I'll get off my soapbox now. **44:36** I think Vijay summed it up pretty well. And I think that there's probably not a lot to add here. And we've said this time and time again, but I guess just if you're reading the paper, what I would recommend is really focusing on the methodology, and this ties into what we've said. So really just understanding what we did rather than exactly the outcome of what we did, I think, is important here. So this is not a paper to say this is the tool for you. This is a paper to say this is how you can find the tool for you. And just looking at, can you make your process more robust using this tooling? Can you lessen the work for your architects using this tooling? So answering those kinds of questions with a process rather than making an explicit recommendation. **45:15** Very cool. **45:16 Chris Romeo:** So, Kristen, VG, thank you for taking us through this paper. Like I told you, when I first found it, I was like, I'm not going to read anymore because I want to interview these people and I want to understand it from their perspective. So, but thank you for sharing this insight with our audience. And we look forward to hearing about the new open-source tool that's going to come out at the end of this. And so let us know when you're ready to talk about that. We'd love to introduce that to our audience as well. So thank you very much. **45:43** Thank you. Thank you for the opportunity. This has been pretty exciting. We've been longtime listeners. It's great to be finally here. **45:50** Yeah, thank you so much. This was a lot of fun. **45:52 Chris Romeo:** Thanks for listening to the Application Security Podcast. You'll find the show on Twitter @AppSecPodcast and on the web at www.securityjourney.com/podcast. **46:06** resources/podcast. **46:06 Chris Romeo:** You can also find Chris on Twitter @edgeroute and Robert @roberthurlbut. Remember, with application security, there are many paths, but only one destination. --- Source: https://appsecpodcast.com/kristen-tan-and-vaibhav-garg-machine-assisted-threat-modeling/