--- title: "Kim Wuyts — Privacy Threat Modeling" url: https://appsecpodcast.com/kim-wuyts-privacy-threat-modeling/ date: 2020-03-23 duration_seconds: 1656 guests: ["Kim Wuyts"] topics: ["Threat Modeling", "Privacy and Compliance"] audio: https://www.buzzsprout.com/1730684/episodes/8122610-kim-wuyts-privacy-threat-modeling.mp3 transcript: true --- # Kim Wuyts — Privacy Threat Modeling *March 23, 2020 · 28 min* with [Kim Wuyts](https://appsecpodcast.com/guests/kim-wuyts/) on [Threat Modeling](https://appsecpodcast.com/topics/threat-modeling/), [Privacy and Compliance](https://appsecpodcast.com/topics/privacy-and-compliance/) [Audio](https://www.buzzsprout.com/1730684/episodes/8122610-kim-wuyts-privacy-threat-modeling.mp3) ## Show notes A system can protect data from attackers and still violate the privacy of the people using it. Kim Wuyts, a privacy researcher and contributor to LINDDUN, explains why privacy deserves its own threat modeling questions. She distinguishes security goals from harms to individuals, then walks through the framework’s approach to finding privacy risks in software designs. The conversation explores how linking ordinary pieces of information can identify a person, why non-repudiation can be undesirable in a privacy context, and what disclosure, awareness, and compliance mean for a design. Kim also discusses the relationship with privacy by design and offers a practical starting point for teams that already use diagrams and security threat modeling. The Application Security Podcast is brought to you by [Security Journey](https://www.securityjourney.com/). About Security Journey Security Journey provides application security education for developers and everyone in the software development lifecycle. → [Learn more about Security Journey](https://www.securityjourney.com/) Connect with Kim Wuyts: → [Kim Wuyts on LinkedIn](https://www.linkedin.com/in/kwuyts/) Mentioned in this episode: → [LINDDUN privacy threat modeling](https://linddun.org/) → [LINDDUN research and publications](https://linddun.org/publications/) Chapters: 00:00 Privacy threat modeling with Kim Wuyts 01:44 Kim’s research background and LINDDUN 04:08 How privacy differs from security 05:55 Modeling harm to the data subject 08:50 Using diagrams and the LINDDUN framework 10:50 Linkability and identifying people from data 14:32 Lessons from the AOL search-data release 15:43 Non-repudiation as a privacy threat 16:44 Detectability and privacy 17:52 Information disclosure and the remaining categories 20:14 The connection to privacy by design 21:06 Getting started with LINDDUN 23:46 Future directions for privacy threat modeling 25:57 Kim’s closing advice ## Transcript *4,118 words · assemblyai* **0:00 Chris Romeo:** Kim Vats is a postdoctoral researcher at the Department of Computer Science at KU Leuven, Belgium. She has more than 10 years of experience in security and privacy and software engineering. Kim is one of the main forces behind the development and extension of Linden, a privacy threat modeling framework that provides systematic support to elicit and mitigate privacy threats in software systems. Kim joins us to explain the difference between security and privacy and introduce us to Linden and how to use it. I hope you enjoy this conversation with Kim Vats. At Security Journey, we believe security is every developer's job. We work with our customers to help them build long-term, sustainable security culture amongst all their developers. Our approach is to provide security education that's conversational, quick, hands-on, and fun. We don't do lectures. Instead, we let the experts talk about what's important. Modules are quick, 10 to 20 minutes in length. We believe in hands-on experiments, builder and breaker style. that allow your developers to put what they learned into action. And lastly, fun. Training doesn't have to be boring. We make it engaging and fun for the developers. Visit www.securityjourney.com to sign up for a free trial of the Security Dojo. Hey folks, welcome to this episode of the Application Security Podcast. This is Chris Romeo, CEO of Security Journey. I'm also joined by Robert Hurlbut. Hey Robert, how's it going today? **1:40** Hey Chris, yeah, it's Robert Hurlbut, good to be here, Threat Modeling Architect. **1:44 Chris Romeo:** And we are about to talk about something that fits right along with Robert's title there, right? We're gonna dive a little deeper into the world of threat modeling. And so a little bit of a backstory here as to how we got to this interview. So I was actually doing a talk at a local event here in the Raleigh-Durham metro area, and I was talking about threat modeling, And somebody raises their hand at the end of my talk and says, hey, what do you think about privacy threat modeling using this Linden method? And I said, hmm, that's a great question. I have no idea what you're talking about right now. And so the gentleman gave me a couple of sentences about what it was, and I thought, hmm, this is something I want to learn more about. And so our guest today is Kim Vats, and she is going to explain this idea of privacy threat modeling. But first, Kim, Our audience is always on the edge of their chair, wanting to know what is your security/privacy origin story? How did you get into this world of security and privacy? **2:47 Kim Wuyts:** Well, first of all, really great that I can join your podcast. So, I'm a postdoctoral researcher at the university, KU Leuven in Belgium. And I kind of just rolled from my thesis into some more security research. Well, I think one of my first projects was to apply STRIDE to some architecture, which I really liked. And how I shifted to privacy is actually also a similar project where I again was asked to do security assessment of an architecture. So, of course, we used STRIDE. And then my colleagues who were asked the same to do for privacy, they said, well, Wow, you guys have a method to systematically analyze security. That's really interesting. For privacy, we don't really know how to approach this. That brought us to the idea of, well, that's interesting and we should see how we can combine something like threat modeling and privacy knowledge. That's where my research took off and focused on privacy and threat Very, very cool. **4:08 Chris Romeo:** And so, one of the things I wanted to state and kind of work out very early is, what's the difference from your perspective between security and privacy? Because I see a lot of folks still, even in the security industry, that I don't think they really 100% understand the differences between these 2 things. And so, from your perspective, what's the difference? **4:33 Kim Wuyts:** Yeah, we, we get a lot that privacy is just confidentiality, but of course it's a lot more. I would say it's all like encompassed within the Linden acronym. So, um, it's a mnemonic for linkability, identifiability, non-repudiation, detectability, disclosure of information, unawareness, and non-compliance. So the information disclosure is just one of the 7 categories there. It's sometimes difficult to, to draw like a strict line be— between security and privacy. because, well, when you, when you talk about confidentiality and access control and limiting what attributes you want to share or you want to store, then there's a thin line between security and privacy. But I, I think one of the main differences is, yeah, kind of the way you approach it, as security is really about protecting the assets of yourself, of you as a company, or whoever is doing the assessment, while privacy is protecting the assets of the data subject. So you kind of need a different mindset to think about privacy. And it might conflict also the way your business would look at it because maybe the data is useful from a company perspective, but it would violate the data subject's privacy. **5:55 Chris Romeo:** Okay. So yeah, that's very helpful to kind of get your perspective on the differences there. And I certainly want to dive deeply into the different areas that make up the mnemonic behind Linden, because I'm fascinated. I'm fascinated by what those things are. But I think it would help our audience who have heard a lot about STRIDE and threat modeling from the security perspective. I mean, Adam Szostak's been on the show a number of times, Joff. There's been a number of different people that have been here to talk about threat modeling, but we've never talked about it from a privacy perspective. And so what are the kind of What are the differences between security and privacy threat modeling? **6:38 Kim Wuyts:** So content-wise, we have the different mnemonics, we have the different categories. Yeah, I already mentioned like the perspective is different. So when a typical security threat modeler would look at something, they will not see it as a threat, while from a privacy perspective, it would violate the data subjects' rights. attacker perspective is a bit different as well. From a security threat modeling perspective, you would talk about a misactor or an attacker, an external attacker. From a privacy perspective, a lot of the— well, I still call it threats, but I'm not sure whether that's the right term anymore. It's more about internal misbehavior where the system processes more data than it is supposed to do or collects more data or shares more data. It's thinking about what is your system or the organization behind it doing wrong with respect to the data subject? **7:39 Chris Romeo:** So it sounds like one of the differences you're describing here is that while security threat modeling tends to look from the outside in, we're always thinking about the trust boundary and what interfaces are crossing that and what particular threats exist around that trust boundary. It sounds like you're saying from a privacy threat modeling perspective, you're more interested in almost an inside-out view or look at how the issues are coming together. **8:08 Kim Wuyts:** Yeah, sort of. Although, I would still state that maybe even more so than for security, many threats are located on that crossing of trust boundaries. It is precisely when you collect data or you share data that crosses that trust boundary, you have like the main privacy violations. Of course, when you look at compliance and minimizing internal processing, then it's clearly within the trust boundary. But I would say a lot of issues still are at that trust boundary. **8:47** Hmm. Okay. **8:50 Chris Romeo:** Well, that's definitely good to know. So, before we dive into the, The kind of the mnemonic and walk through each of these categories. I'm curious about the process behind Linden. And so meaning, is it similar to what we do on the security side where we're using data flow diagrams and we're drawing pictures as our way to try to unlock where the threats are coming from? Or have you taken a different process approach to actually using the Linden mnemonic? **9:21 Kim Wuyts:** Yeah, no. So we were very much inspired by STRIDE. So we also believe that you will need to have in parallel a security assessment to be sure that the entire range of security and privacy is covered. So that also made it, according to us at least, useful to follow STRIDE quite closely. So you have one DFD that you can use both for the security and privacy assessment. So we follow the same steps. We have the data flow diagram for which you make a mapping table for each of the landing categories and then you just systematically go over each of the cells of the table and look at the Linden knowledge, Linden threat trees, to determine whether for that specific DFD element, a privacy threat occurs. **10:13 Chris Romeo:** So, I'd like to walk through the list of categories within Linden. I see one that looks familiar, but everything else looks like it's kind of some new stuff. And so, if we start with linkability, I'll just read. I pulled the list off from the Linden website, which we'll provide as a resource. to folks so they can go dive into this a little closer. I will go ahead and just read the description, and I'd love to have you kind of give us some more details behind what this actually means. But when I look at linkability, it says an adversary is able to link 2 items of interest without knowing the identity of the data subjects involved. **10:50 Kim Wuyts:** Yes, kind of already sums it up quite nicely. It can be seen as kind of the prerequisite for identifiability. So basically, the more data you collect on one person without necessarily knowing who that is. So, when you link all those attributes, you get a more scoped view on who that might be. So, you have something that is called the anonymity set, which is a collection of possible identities or persons that might correspond to the data. The more information you have, clearly, the smaller that anonymity set becomes. So by linking more information, you would more quickly go to identifiability. But also linkability could lead to attributability, to basically singling out one identity without knowing that identity. So you can all link something to a pseudonym, but you don't know who that is in real life. So it's not identifiability, but you still create a profile. And also, it can also be linking data. the same type of attributes, but about different people. So for example, when you link information about people who have the same disease, that's also linkability. And that can lead to, yeah, like some more societal harm where an insurance company might use that information to say, well, a lot of people in that area have a higher chance for this kind of disease. So we will ask them to pay more or something like that. So there are different consequences coming out of that linkability category. **12:34 Chris Romeo:** So you gave an example of kind of disease as a, I guess, an item of interest. What are some of the other items of interest there? Is it something as simple as like name, address, phone number, or is there something else that you're specifying? **12:49 Kim Wuyts:** It can be basically everything or any combination of attributes even. There's not really like saying when you just remove the name, then you're safe for linkability and especially also not for, for identifiability. But it can be basically all types of combinations. You have also something called pseudo-identifiers or quasi-identifiers, which are on their own nothing special. They will not reveal any information, but when combined, they are sufficient to identify with almost 100% certainty everybody in the anonymity set. But it's really hard to state like What type of attributes that could be precisely. It really depends on the data you have and also the additional information other people might have, which is in the current age a lot. So it's not really easy to say it's this particular attribute. It can be, I don't know, your hair color and your size and whatever. **13:55** I've seen an interesting example of this, and it may also fall into identifiability as well. But for example, I've seen some sites, social media sites, that on one hand you were supposed to be anonymous, but on the other hand, if you had an Instagram account or something like that, they would automatically link to it and show photos or something like that. And so people could figure out who you were just simply by connecting the dots, by making those connections without getting your permission to do so. So those sort of fall into that linkability, but also, I guess, potentially identifiability as well. **14:32 Kim Wuyts:** There are a lot of examples, like a couple of years ago, I think it was a couple of years, it's probably more, AOL released a whole list of search queries and some researchers really managed to identify one of the people doing the queries Based on, I don't know, it was some things about her dog and something about her son. And just the unique combination of search queries were sufficient to identify somebody. **15:04 Chris Romeo:** So that brings us then to identifiability. An adversary is able to identify a data subject from a set of data subjects through an item of interest. So it sounds like there's a lot of connectivity there between linkability and identifiability. **15:17 Kim Wuyts:** Yeah, yeah. So, you link pseudo-identifiers or you link some anonymous or de-identified item of interest to something that's already identified, or you further narrow the anonymity set until you eventually get some unique identity. Yeah, it's related indeed. **15:38 Chris Romeo:** Okay. And then the next one is the familiar one to those that know Stride. **15:42** Yeah. **15:43 Chris Romeo:** Non-repudiation or, you know, repudiation is how it appears in Stride, but non-repudiation, Identity subject is unable to deny a claim having performed an action or sent a request. And so, how does that play out on the privacy side? **15:56 Kim Wuyts:** Yeah, so, so from a privacy perspective, sometimes plausible deniability is something we need. If you think of online voting, you will want to be able to deny who you voted for. Some kind of whistleblowing systems would also benefit from these kind of things. It's not necessarily mainstream requirements, but for, for specific applications, plausible deniability will be an important requirement. So it kind of conflicts with repudiation from STRIDE, but I don't think that it's really a conflict for any application because I don't see an application where you would need to have some, some really strong proof of something and at the same time have for that same action plausible deniability. **16:44 Chris Romeo:** Yeah, that makes sense. And so the next one is detectability. An adversary is able to distinguish whether an item of interest about a data subject exists or not, regardless of being able to read the contents itself. How does this one work? **16:58 Kim Wuyts:** Yeah, so this is again where the line gets thin between privacy and security. It's kind of a bit like side-channel attacks where you— so just by observing that some information is being sent maybe at an irregular time, you know, that maybe an emergency situation occurs, or based on the size of information that is being transmitted, you can deduce more information that will— that can be linked to the profile and provide more information there. And there are so many things that you can deduce from side-channel attacks and, and those kind of things. Like a couple of years ago, I think it was even shown that that based on the electricity or the power consumption of a TV, you could deduce which show was being watched. So, I mean, the possibilities are endless. **17:52 Chris Romeo:** And then the next one is disclosure of information, which is one of the OWASP Top 10, right? Disclosure of information. **18:01 Kim Wuyts:** Yeah, so it's the category we kind of borrowed from STRIDE. Ideally, you don't share any data, but of course you will need to have some data in your systems and share it. But Then the least you should do is to properly protect it. So you will need confidentiality, clearly. **18:17** Yeah. **18:18 Chris Romeo:** And then unawareness. The data subject is unaware of the collection, processing, storage, or sharing activities and corresponding purposes of the data subject's personal data. So this is from the perspective of me as the user of being unaware that somebody's collecting data when I should know that they're collecting. **18:36 Kim Wuyts:** Yes. Yeah. **18:37** Yeah. **18:37 Kim Wuyts:** So, so this one and the next category are more the, called the soft privacy categories. So this is more about transparency, like as a data subject, as a user. But even if you're not a user, but data is being collected about you, you should be aware of that. You should be able to, to get access to that information, to know what is being, why it is being used. You should be able to intervene to some extent to the processes that are be— that the data are being used for. So it's, yeah, it's more about data subject rights, basically. **19:13 Chris Romeo:** Okay. And then non-compliance. I think I know where this one's going, but the processing, storage, or handling of personal data is not compliant with legislation, regulation, and/or policy. So I'm guessing there's a GDPR connection. There's potentially, there's this new California CCPA law that's affecting us here in the US. So I'm guessing that's where you're going here. **19:34 Kim Wuyts:** Yeah. So, yeah, so the category is already almost 10 years old, I think. So it was pre-GDPR. But indeed, when we created our framework, we saw that need to link privacy more closely to legislation. And we're already working on aligning it further with GDPR and looking at different regulations. But basically overall, the main concepts of all data protection regulations are kind of similar and relate to minimization, privacy by design as a whole, consent kind of concerns, and so on. **20:14 Chris Romeo:** So you mentioned privacy by design. Does Linden— is this a framework that gets me privacy by design, or is privacy by design something different? **20:24 Kim Wuyts:** That's almost a philosophical question, I would say. So yeah, it gives you as much privacy by design as Stripe would give you security by design. So I think it's something that will help you. So for me, privacy by design means that you really think of privacy early on in your development. So by applying something like Linden, you will systematically integrate privacy in your design. Ideally, you will also document your decisions, which will give you the accountability you would require from laws like GDPR. So in that sense, yes, I would say it's it would definitely facilitate privacy by design. **21:06 Chris Romeo:** How would you recommend somebody get started using this Lindon framework? A lot of times our listeners are relatively or brand new to a topic, and so some of them might be sitting there thinking, hey, I've never heard of privacy threat modeling, didn't even know this thing existed. How would you recommend that they get started using this and being successful with this methodology? **21:27 Kim Wuyts:** Great question. And that's actually something we have recently been working on as well. So Lindon has been around for almost 10 years now, and we're getting great response, and, and people seem to like it. But when, when you then go deeper and ask how they apply it, well, we get some feedback that maybe threat trees are— can be a bit complex to, to understand. And people tend to use the acronym as like inspiration for brainstorming, but really don't go further than that and kind of miss the intended systematic approach similar to STRIDE. That inspired us to look for something more lightweight. We recently launched LindenGo, which is a toolkit that aims to provide some more lightweight support. It's basically a set of threat types, so descriptions of things that can go wrong categorized per Linden category, and it provides a lot more guidance than the 3s did. We have some guidance questions. We have some examples, some more information, some hints towards likelihood and impact. And it's bundled in a toolkit. It's actually a set of cards that should get you started. It was driven by industry demand. So, we're really looking forward to get feedback from the bigger community. We've tested it out on a number of occasions, but still, it would be really great to get input from anybody who is getting started or is already familiar with privacy, with threat modeling, to see what the next steps are to further improve it and to make it easy and lightweight to use. **23:13 Chris Romeo:** Where would you recommend that our listeners connect in order to provide feedback? **23:20 Kim Wuyts:** When you go to our website where you download the, the Lindengo PDF, so it's all publicly available, there's already a link to survey, an online questionnaire. So, that would already be be very helpful. There is also a link to the Linden email address, which is info@linden.org, or you can just connect to me personally through mail or LinkedIn or Twitter or whatever. **23:46 Chris Romeo:** So, when you're thinking about the future then of privacy threat modeling, what are some of the areas that you think we need to explore in the coming year? **23:55 Kim Wuyts:** From a more meta level, what I think is interesting to see is that even though threat modeling has been around for, for over 20 years now. It's still kind of ad hoc, and there's not really like a formal definition. As I mentioned before, the feedback we got from Lyndon, we get the same things when we talk about STRIDE to people. There is such a big variety of how you would apply something like STRIDE, how you would apply threat modeling I think everybody kind of has a different definition. So kind of streamlining this would, especially from an academic perspective, be very interesting. And then further than that, I think focusing on making it more lightweight, as threat modeling is quite labor-intensive. I think you need a lot of time, you need expert knowledge. Mostly a manual process. There are some tools like the Microsoft tool which will help you get along, but it's more about generating documentation than really about automation. I'm not that familiar with the commercial tools, so maybe there are some more features there. So I think looking at how we can give more support there from automation tooling support, what do you need, how do you need to extend your models because your models would then need to contain more information to automatically extract risks or threats from that. Also, something interesting I think is evolution or change management because your model will change and it would be great if you would not need to redo your entire threat model analysis over and over, but you only need to do the the delta analysis. So, I think there are still quite some interesting topics to further explore. **25:57 Chris Romeo:** Yeah, we totally agree. Both Robert and I, as big fans and practitioners of threat modeling, I think there's a lot more work to go. And it's exciting just to think that there's a lot of folks that are out there thinking about the future and even specifically how privacy comes into this. And so, Kim, any last final thoughts or conclusions that you would offer, maybe a call to action for our audience? **26:22 Kim Wuyts:** Yeah, I just hope I've inspired some people to think about privacy. Feel free to have a look at the website. And I really hope that people will give some feedback on Lin and Co and tell us how we can improve it, or maybe even share what they see as challenges and give us some more interesting research areas to look to examine. **26:45 Chris Romeo:** Kim, thank you for taking the time to educate us and our listeners about privacy threat modeling and Linden. And it's definitely piqued my interest to go dive even deeper into this topic. And so, thank you very much for sharing your knowledge here. **27:00 Kim Wuyts:** Thank you. It was my pleasure. **27:03 Chris Romeo:** Thanks for listening to the Application Security Podcast. You'll find the show on Twitter @AppSecPodcast or on the web at www.securityjourney.com. application-security-podcast. You can also find Chris on Twitter @edgeroute and Robert @roberthurlbut. Remember, security is a journey, not a destination. --- Source: https://appsecpodcast.com/kim-wuyts-privacy-threat-modeling/