Skip to content
AppSec PodcastThe Application Security Podcast — home
42 minSeason 9, episode 16

Will Ratner -- Centralized container scanning

With Will Ratner

Security TestingSoftware Supply ChainCloud and InfrastructureDevSecOps and CI/CD

Will Ratner is a software security professional with extensive experience building and implementing security solutions across a myriad of industries including banking, media, construction, and information technology. In his current role at Atlassian, Will focuses on improving the vulnerability management process by building highly scalable and automated solutions for the enterprise.

Audio hosted by Buzzsprout. Nothing loads until you press play.

Episode chapters · 13 chapters
  1. 00:00Meet Will Ratner: Centralized container scanningAudioVideo ↗
  2. 01:31Yeah, and as we do that, let's turn over to WillAudioVideo ↗
  3. 04:38Very cool. And so again, we were talking about container scanningAudioVideo ↗
  4. 06:59Yeah, I love— I love that we're talking about something that'sAudioVideo ↗
  5. 10:12If only everyone would start from scratch with their base imageAudioVideo ↗

About this episode

Will Ratner is a software security professional with extensive experience building and implementing security solutions across a myriad of industries including banking, media, construction, and information technology. In his current role at Atlassian, Will focuses on improving the vulnerability management process by building highly scalable and automated solutions for the enterprise. Will joins us to discuss a centralized approach he built for container scanning. We explore the challenges and lessons learned, building a scalable, enterprise-grade solution, and how to build something that developers will see value in. We hope you enjoy this conversation with… We also introduce the idea of container scanning and building secure images that are then used to generate containers.

Connect with Will Ratner:
Atlassian
Docker

Resources
Atlassian
Docker
Kubernetes
Snyk
Go
Slack

Actionable

From this conversation

  1. Make scanning hands-free

    One of the big things that we try to do, and like I've tried to do at any company, is like make it as seamless and hands-free as possible.

    5:03
  2. Use multistage container builds

    The big thing you can do is use multistage builds, which is super valuable, right?

    11:02
  3. Avoid breaking builds at scale

    I would say starting out breaking the build is not scalable, right?

    19:34
Transcript · 42 min conversation

0:00Chris RomeoWill Ratner is a software security professional with extensive experience building and implementing security solutions across a myriad of industries, including banking, media, construction, and information technology. In his current role at Atlassian, Will focuses on improving the vulnerability management process by building highly scalable and automated solutions for the enterprise. Will joins us to discuss a centralized approach that he's built for container scanning. We also introduce the idea of container scanning and building secure images that are then used to generate containers. We explore the challenges and lessons learned in building this scalable enterprise-grade solution. We also explore how to build something that developers will truly see value in. We hope you enjoy this conversation with Will Ratner. You're about to listen to AppSec Podcast. When you're done with this, be sure to check out our other show, Hi5.

0:59Robert HurlbutHey folks, welcome to this episode of the Application Security Podcast. This is Robert Hurlbut. I'm a threat modeling architect, and I'm joined by my co-host, Chris Romeo. Hey, Chris.

1:10Chris RomeoHey, Robert. Great to be here once again talking about all magical application security things. I know we got a great conversation today with a topic that I think a lot of people are trying to figure out. Container scanning at scale is what I'm going to call it right now, but we'll let our guests kind of get into that as well.

1:30Robert HurlbutYeah, and as we do that, let's turn over to Will. Thank you, Will, for joining us. Will Ratner, appreciate it. I know you just put out a— or actually did a talk recently about this, but we're going to sort of dive into that a little bit today. But before we get started, could you let us know what is your security origin story? How did you get into that wonderful world of application security?

1:53Will RatnerYeah, definitely. Well, first, thanks Chris and Robert for having me on the podcast. So my origin story, it's probably similar to like Spider-Man. So, you know, he gets bitten by the spider and then he's trying to figure out his powers and like he'll be climbing up walls and falling off or jumping across buildings and taking a beating most of the time when he's starting off. And I didn't have a lot of physical beatings, but there was a lot of like trial and error throughout my origin story. So I had always been interested in computers. I went to school for computer engineering because I thought that would be a good opportunity to see the hardware and the software side of things. And I took my first analog systems course and found hardware was not for me. So I focused on the software side of the house. And as part of that, I had an internship as an application developer at a big enterprise, where I worked on a very large, like enterprise Java application and found very quickly that I did not enjoy that either. But while I was there, I took the opportunity to go and shadow different teams at this company and went and shadowed somebody on the security team. And I found it to be really interesting. This was a pen tester at the company. And I thought that was so cool, really interesting. And so like, that's what I wanted to focus on or try to learn more about. And so In colleges at the moment, there's a lot of schools, most schools don't have a whole lot of security education opportunities, but where I went, we had a couple, so I took every one I could. I eventually got an internship on the application security team at Capital One, where I learned a bunch, really got to dig deep into static analysis tooling. But from that, I learned that I missed the development aspect. Like, I still really like to code. So after graduating school, I went back to Capital One and was on their kind of security software development team, which was pretty new. So what we focused on was building security tooling for the enterprise. And I really found like, that was my sweet spot. From there, I bounced to a couple different companies, kind of doing the same thing. Eventually, like, I wanted to really get like, see what other opportunities in security there were. So I went and like did the startup route. And as like the sole security or like a very small security team, kind of wearing many different hats from like security operations, vulnerability management, risk and compliance, you're kind of covering the whole gamut and found that like, again, you learn that you don't love all of those things. And that like, I decided to go back to my roots. And that's how I ended up back at Atlassian, where I am on the product security team. focusing on vulnerability management and helping build out more security tooling at scale.

4:38Robert HurlbutVery cool. And so again, we were talking about container scanning today and looking into that and scaling it out. What problems are you trying to solve whenever you're looking at a centralized approach? Now, I know there are some tools that you can run just if you have a few containers and so forth. But if you're centralizing, what are some of those things you're trying to solve?

5:03Will RatnerYeah, so when you're implementing like any kind of security solution at a company, you have to like get the developers on board or like the engineering teams on board with using those solutions. And the harder you make them to use, the less participation you're actually going to get, right? So One of the big things that we try to do, and like I've tried to do at any company, is like make it as seamless and hands-free as possible. So when you're working in like trying to build out a centralized solution, you find out that there's like, you have hundreds of teams, thousands of engineers that all do things differently, right? And they might build their containers differently. They might build their code differently. It doesn't matter how they're doing it. But it's important, depending on the size of the company you're at, to like try and find like a general pattern between all these things and build off of that. So at Atlassian, where I'm currently at, we have a really great platform as a service to deploy any code. And in order to deploy your code and have an application, it has to be built in a container. So that made building a centralized container scanning solution pretty easy for us, just because we know, like, if they're deploying something, it's going through this platform as a service. And we have an enterprise event bus that we can use, and I can dig into that a bit later. But like, it makes it very easy when you're centralizing like the actual development process as well, or the infrastructure process. So I would say like, yeah, that's pretty valuable. And then again, like considering like working with the developers or the engineers, like team-by-team basis, like maybe the ones that might be affected the most. And you can kind of do that research initially, which is what we did, and finding out like who's going to be affected, how do they build it, how can we build them a more seamless solution?

6:58Chris RomeoYeah, I love— I love that we're talking about something that's scalable because a lot of times in application security conversations, people aren't thinking about, well, but how does this work for 1,000, 2,500, 5,000 developers? Like it's one thing to have— well, we have a development team of 8 people. It's easy to tell them, hey, scan your containers, and hey, I'm going to check in at the next standup. Did you scan the containers? You did? Great. All right. Awesome. But when you have the number of developers that you're working with at scale, you don't have that opportunity. So I want to make sure we back up for a second and just introduce this idea of container scanning. I know all 3 of us know what that is, and we're super excited about it as a category, but I want to make sure any of our listeners who might be sitting here going, okay, I'm You're talking about something I don't really understand. So, Will, can you give us your take on when you say container scanning? Like, what are you looking for? What are you scanning for? Kind of give us a little bit more on the details of that.

7:53Will RatnerYeah. So maybe like can start off with the definition of containers, which is like we're packaging up the code, the dependencies, the OS. It makes it much more portable to deploy your code and run your code. And when you're scanning these containers, there's different scanning aspects, right? In my case, like my focus has been on container image scanning. So when you build a container, you're picking some base image that you're building off of. And these can be like a bare bones, you know, Ubuntu or Debian image. It could be like a scratch image, which has nothing on it, but then you're building your application on top of that. And so what container image scanning focuses on is actually scanning like the libraries dependencies packages that you're building in that container that exist in that static container image. There's container runtime scanning as well. I haven't delved too far deep into that, and that's not what this will be about. But that's when you're scanning like what's going on in the container while it's running, looking for anomalous behavior and flagging those kind of things.

8:54Robert HurlbutOkay.

8:55Chris RomeoSo you're really— this is a software supply chain play. Like you're scanning to see, hey, the dependencies that are part of this image that's been defined by a developer somewhere. Are there any known third-party or open-source security vulnerabilities in these particular components so that the end goal is that you can then send some type of an alert to a developer saying, hey, your image has issues, you got to fix this in order to protect our production environments?

9:25Will RatnerExactly. And what you find with containers and what's so great about the open-source world is there's a base image or a library for everything you want, right? If you're building a Python app, there's the Python base image container that you can now use, and it'll have all the dependencies that you want. But what developers might not always know is that in those base images, because they're trying to cater to everyone and make them as generic as possible, they contain hundreds of packages that you might not need or know exist there that don't always get updated, or CVEs are discovered after the fact and you need to update them. So that's what this scanning solution is really focusing on, is like trying to make sure as developers are using these open source base images that we're keeping them, the packages that come with them, up to date and secure.

10:12Chris RomeoIf only everyone would start from scratch with their base image, right? But we all know that the challenge there is that takes a lot of effort, takes a lot of work to build an image. And, you know, one of the things that we always talk about is like, you know, you have a number of different places you can start. You can start with a minimal, you know, like Alpine, kind of a minimal image base. You can start with one of the big Linux distros. But to your point, that means you're bringing, even if you don't want them, you're bringing in hundreds of packages into play here. If we could just start from scratch, but the only problem is it takes so much time to start from scratch and get, and then you try to run it and like, ah, it won't run. It's missing something. Like, what's it missing now? You know? And so that's definitely the challenge. So, Any kind of insight on from your perspective? Like, what do you recommend to people? Do you recommend scratch images as a base, as a starting point? Or what do you recommend? Yeah.

11:02Will RatnerSo for me, like, I usually build— anything I build is typically in Go. So that makes it really easy for me to use a scratch image because I have a statically compiled binary that I can just pop onto that container that has nothing on it and it'll run just fine. Maybe I have to like copy over some, some other files from like another build, which brings me like, the big thing you can do is use multistage builds, which is super valuable, right? So you can use and build your application in like a Python 3.9 image. And then once you have that, you can actually like move over your built binary or like after you've collected all the dependencies and move them over to like a more bare bones like Alpine image as well. And that like really will help you slim down those images, which also leads to like faster deployment times and faster build times. But also like it's more secure because you have far less just like junk packages around.

12:01Chris RomeoThis is just bonus content for the audience here. This wasn't even on the agenda. We just went into this whole idea of like images and stuff. But I think it's important. It is a crucial component of having a solid application deployed to production. You mentioned multi-stage builds. If you're going to have something more complicated, you can throw away pieces of the container. You don't need to drag everything to production. You could use some pieces and then throw them away when you don't need them anymore and end up with a slimmer image at the end of the game. But let's talk about what you built as kind of our first step. So I understand that from listening to the talk you did at SnykCon, You built this centralized approach to doing container scanning. So kind of at a high level, walk us through what that looks like.

12:47Will RatnerYeah. So, you know, I kind of mentioned at Atlassian, we have this platform as a service, and we have this enterprise event bus. So one of the nice things about that is like, anytime someone makes a deployment on that platform, an event is generated, it goes on the bus, and it can be ingested somewhere else. What we built was this platform that basically would ingest those deployment events. And in those events, we have every container that's being deployed for a particular team service. So it'll have like the container name and then it'll have like the image digest as the tag. So we typically don't use the tags and this is safer this way, just because obviously if you update the image, but don't change the tag, maybe as part of like a staging build. And then in production, like you do a redeploy later on and it pulls the newest images, then you have things that break. So we always have the image digest as the tag. So we know the exact image that's being deployed at a given time. So we ingest these. Now at Atlassian, we have thousands of developers who are deploying, you know, thousands of containers a day. So we needed something that could scale and scan these on a regular basis. So we chose to go with a Kubernetes-based infrastructure, and we use this tool called Argo Workflows, which is a Kubernetes-native workflow orchestration tool. So it allows us to basically run, like build a workflow as a Kubernetes file and run many containers in parallel, which is what we needed, because typically with these deployment events that are happening, Not only are you deploying the application container, but there's also various like platform sidecars that go around with it. Like you're going to have different traceability tooling, you're going to have different like authorization tools as well that like they run in parallel with your application containers and help you interact with other services within Atlassian.

14:47Chris RomeoSo is Argo— is Argo Flow, is that an open source or is that a commercial?

14:52Robert HurlbutOkay.

14:52Chris RomeoSo is everything in your— is this— yeah, this whole thing is open source. So there's nothing— you didn't buy anything that you plugged into this?

14:58Will RatnerThe only thing that we use, which, you know, if you're scanning some open source tools, like we use Snyk, right, as our container scanning solution, which is a commercial tool. But if you're scanning open source things, then it is free. So there is that benefit as well. But what this allows us to do is, so we can, we can just basically have, say there's 6 containers being deployed for a service, we can scan all 6 of those at the same time. And so that's, that's what the platform was doing. And then behind the scenes, we, we have this thing at Atlassian called the vulnerability funnel, which is where all of our vulnerability tickets for teams go. So we have like, again, like I mentioned, it's important to centralize as much as you can to have a centralized scanning solution. So we have a centralized like ticketing system that goes out to teams, and there's a bunch of automation rules built on that and who gets assigned and how long they have to fix the issue. So, you know, basically the workflow looks like we get these deployment events, we build an Argo workflow, which then triggers off the scan. We have to download the container. We download the container archive because we're using a shared Kubernetes cluster. And because of that, you don't have like the privileged mode. So Docker in Docker doesn't really work very well, but we can download the image. And then we pass that to the scanner, which ones to scan. We get the results, we process those results, create tickets to the teams, and then they have, you know, X amount of days based on the severity to resolve those issues.

16:29Robert HurlbutSo one thing I noticed also in your talk was that you focused on, or shifted, I guess, a focus on event-driven scans versus in the CI/CD pipeline, as many people will advocate. So is that— tell us about that experience and why you chose that instead.

16:47Will RatnerYeah, so we kind of, we were talking about scale earlier. And so, like Chris mentioned, like, oh, if you have just like 8 engineers, you can just ask them at standup, like, did you, did you solve that issue? Yes or no? And then you move on. But that doesn't really work for like thousands of people. So there's kind of this challenge, right? So you can go and, and what you'll find is like a lot of tooling, scanning tools. They often build very pretty UI and they expect you to work out of that UI. And that never works very well, especially when you have a lot of people. And then there's the CI/CD option, but trying to ask when you have thousands upon thousands of repositories that would need to have that integrated into them, it's not really the easiest ask to go and say, okay, engineers, please go and do this and you have to do it. And And this will start like blocking your builds, especially in the case where you're like rolling out a scanner for the first time. And so like maybe you've never scanned these containers before, there's going to be tons of issues. And then when you put it in the CI/CD pipeline and you're blocking builds from, from going forward, like it's going to take a while for anyone to get anything out. And maybe that's good, maybe that's bad. What we've chosen to do is, again, make it as seamless as possible to The engineers don't have to do anything at Atlassian. Like, it's just as part of, you know, you deploy something to production or staging, it's going to get scanned automatically. And so that makes it like very seamless. The event-driven architecture is also very useful because with those events that we receive, we get a lot of metadata that we might not have. We get like where they're deploying it, who's deploying it, what service it belongs to, and we can use that to assign the proper tickets to the folks. and get it resolved in the timely manner that we, we see fit. So, and then also, like, the feedback is much faster, which you would get this in CI/CD as well. But because it's like event-driven, like, as soon as they, they deploy their fix, you know, we're scanning again, and they're getting their ticket automatically closed out if they, if they fixed it. So those are kind of the big benefits of doing that.

18:53Robert HurlbutDoes that—

18:56Chris RomeoI'm just trying to I'm trying to— I'm totally understanding kind of where you are and kind of the design decision that you made in doing this. I'm just trying to quantify, qualify, I guess, whatever the risk in, you know, because classic DevOps people are going to say, break the build.

19:16Will RatnerYeah.

19:16Chris RomeoBreak the build. But what, you know, you've kind of experienced a scalability challenge that says, And I guess I'm kind of summarizing what I think I heard, and correct me if I'm wrong here, but that breaking the build isn't scalable, I think is what you're saying at the end of the day, right?

19:34Will RatnerSo I would say starting out breaking the build is not scalable, right? And what we've seen in the process that we've done is when you start sending tickets to teams and now they have to fix them like out of cycle, right? So they're, they're getting like they're going to have to modify their sprints to now fix these issues. And teams don't like to do that. And so what they do is say, oh, I don't want to do this. I'm just going to integrate it into my pipeline and break my build so I don't have to do this anymore. And that's what we see at Atlassian is most teams are doing because they don't, they don't want to get those tickets in the first place. It causes interruptions. So what we've done is like, we're not forcing you to do it yet, but the way we're encouraging you to do it is actually just as effective because you're getting these tickets after the fact. We're not breaking your build, we're not slowing you down. But now you're taking it upon yourself to want to integrate this into your pipeline on your own volition because it's going to get you less tickets and it's going to make your life a lot easier down the line.

20:31Chris RomeoYeah, that's, yeah, that's interesting. It's just, it's a, it's a scalability to, you know, what we've mentioned a couple times, right? It's a scalability. It solves the, you know, greater than 1,000 developer populations trying to use a centralized shared service. You can't just be like the quickest way to get decommissioned with this tool would be to be the tool that shuts down our ability to push to production. Exactly.

21:01Robert HurlbutIn terms of another, actually, that leads into another question I had related to trust. You mentioned the idea of the breaking and not breaking and pushing it out later. developers don't want to see a lot of tickets. And I think that the other thing I got from the talk is that initially you had thousands and thousands of potential tickets that were generated. So are there some good recommendations on how to handle that in terms of a lot of noise and trying to reduce that as well?

21:32Will RatnerYeah, so that was probably the biggest challenge we faced with this implementation is this was like And this happened— this will happen anytime you're rolling out a scanner for the very first time, right? Like, if you've never scanned any of your thousands of containers before, you're going to see— find out a lot of stuff, a lot of interesting stuff. You're going to see a lot of tickets that would get created, a lot of packages that are out of date, a lot of OSes that are out of date. And those things need to be fixed. So the approach that we took was initially, like, we just started scanning without filing any tickets, right? Like, we haven't scanned before. So getting those tickets out immediately isn't going to make a huge difference. So, but so we can get, we can gather that data to understand like what teams are most affected, who's going to get the most tickets, and can we work with them? Because this isn't the first scanner that was rolled out at Atlassian. So we have a lot of experience with rolling out like other SCA scanners or SaaS scanners, network scanners, those kinds of things. So we know what can happen if you don't do this, and you just flood people with tickets. So that gives us like, we can look at the top 10 teams that are going to have the most container tickets and go work with them, try to resolve the issues before we and file tickets. So that, that kind of like, uh, relationship with the teams is, is very helpful. And having that open line of communication, reaching out to them, explaining like, we're going to be turning this on, you're going to get 1,000 tickets. Do you want that? No, we can work with you and make sure we get those resolved.

22:57Chris RomeoWhat was the, what was the amount of time, Will? I'm just curious, leading up to that point before you went live while you were working, because I think some of our listeners out there that might think about doing this are going to be curious about how much, how much time do I need to take in coaching my teams getting ready for the launch of the tool?

23:12Will RatnerYeah, so that, that can, again, like depend on the size of your organization. Like for us, we on the product security team, we also have folks that like are embedded with teams. So it wasn't just like myself and my team that had to go and do a lot of this work, we could educate some of the product security engineers as well as like what needs to be done. And they can go work with their teams directly. So we only needed about like a month and a half lead time for this, but we had the people to like make that happen quickly. That could obviously vary based on like how many people and the size of your security team.

23:46Chris RomeoAnd I would say the maturity as well, the maturity of your, your security culture in your organization is going to be— it's going to drive how fast you can make— you can get all this stuff ready for a production launch of the tool, avoiding, like we said, the thousands of tickets and stuff.

24:01Will RatnerYeah. But the mistake that we did make when we, when we did launch this tool into production was that there was a lot of tickets created. So I'll take a step back and say resolving container scanning image issues is typically like really, really easy. It's just a matter of because the packages that are being found that are out of date are OS packages. So it's a matter of like just updating the package via the package manager.

24:28Robert HurlbutOkay.

24:29Will RatnerAnd then you're pretty much done. That's pretty much it. It's not— and most of the time those packages are not very ingrained in the application. So it's not like a code library update where you can make a breaking change and it's going to cause a bunch of issues down the line. Like, oh, maybe you just need to update APK or libcurl or something like that. And people don't even know what those are, but it's like a—

24:57Robert HurlbutYeah.

24:57Will Ratnerimportant OS library that needs to be updated anyway. But what we found is that like, so we shipped a bunch of tickets. And a lot of these tickets didn't have fixes for them. And so the engineers get the tickets, and they're like, okay, what do you want me to do? And I'm like, that's a great question, because supposedly, this is a critical severity CVE with like a 9.8 CVSS score. But then you actually go and look at the security tracker for that specific Linux distro. So like, you can go look at Debian, and they've done the research and say like, well, actually, this isn't like an issue that affects us. So we're not even going to put a fix out for it. And that's why there's no fix. So looking at that, we were like, well, why, why are we filing tickets for these if the distro themselves don't think this is an issue? And this was kind of like a change of heart that we had had previously because with like software composition analysis, like when you're looking at, you know, repo packages or like you're using some third-party dependency and those have to be updated depending on like, you could have a library that's like supported by one person and a CVE might've been found and there hasn't been a fix because that person just might not have the time. It's open source. They don't have the resources to handle it yet. But thinking with like these, these Linux operating system packages, I mean, like Debian, Ubuntu, these have like full paid security teams. And I'm not an expert on Linux by any means. So I'm not going to tell them like, oh yeah, that's, that's the right way to do it there. I trust them. So if they're saying that like, it's not an issue, then I'm going to take their word for it. And then there's no point in filing tickets to the teams to fix something that the distro themselves isn't going to fix. Now, that's not to say for like these library scanning packages that if there's no fix, we're not filing tickets for those. We still are because half the time, like, it's a library that's just been deprecated or isn't supported and the teams just need to move off of it. But for these operating system ones, you know, what's the point of trying to get them to do it when the Linux distribution themselves is never going to have a fix? So— That ended up like clearing out a large portion of tickets that teams got. So now like every ticket from a container scanning perspective that teams had to fix, there was a very clear-cut way to fix the issue. Like you had the version you had to upgrade to, and it was a done deal. So that, that helped. Also how we grouped the tickets also helps. So instead of making a ticket for like every single individual CVE that was discovered, were grouping them by packages, right? So you can have— curl might have like 20 vulnerabilities depending on the version they're using. But really, all you have to do is upgrade curl to fix all of them, right? So there wasn't really any point. And it confuses the teams when you're giving them like 20 different tickets. It also makes them think like there's a lot more work that has to be done. So that was a challenge. We've like went back and forth, like, really, is it should we just create one ticket for a container that just, you just need to run like, you know, apt-get upgrade, apt-get update, or do we do it on like a per-package basis? And we decided to go a per-package basis ticket design because that gives the teams a bit more flexibility in terms of like, maybe there is a package they're using that upgrading it is going to cause breaking changes, but they can fix everything else that helps them like clear out those tickets and also focus on like which one they need to be working on fixing. So doing all of that helped cut down on tickets significantly. The other thing that also helped cut down on tickets for teams was that as teams got a lot of these tickets, they were like, oh, this is bad. We should like do something to fix this. So you saw many teams start moving to scratch images. You saw many teams start moving to like much more bare-bones images. And then you saw many teams move to like our Atlassian-managed golden images. So like those are managed by, again, a centralized team that worries about the security vulnerabilities and take that out of the actual engineers' hands who are using them.

29:20Chris RomeoWhich, I mean, that's a really big positive security move to have golden images that somebody is responsible for other than the developer who's building on top of them. And then they can inherit, they can build their image on top of it. and get, you know, that goodness without having to, you know, just start from, you know, their own Ubuntu-style container or something along the way.

29:43Will RatnerExactly.

29:44Chris RomeoSo, so yeah, we talked a lot about tickets in this perspective. One follow-up question on the tickets. I think I heard you say this and I want to understand this a little bit more. You automatically close the tickets for— through this, this centralized service. So That, like, and I'm guessing that was a, that was a design decision you made going into this to improve the lives of your developers. But tell me a little bit more about that because nobody ever does that. Like, nobody uses that as a, as a way to— they don't think about, hey, how are we going to make our developers not think this tool is terrible and is a waste of time for them? And when I hear you say automatic closing of tickets, like, if I'm sitting there as a developer, I'm like, woohoo.

30:27Robert HurlbutYeah.

30:28Chris RomeoLook at 7 of my tickets went away and all I did was fix the problem. So tell me a little bit more about that.

30:32Will RatnerYeah. So when I, when I said seamless, like, I mean, like the developers don't even have to go into the platform or see this, right? Like all they get is a ticket in the same system that they're working in every day for all their other tickets. And it comes across their desk. They're in that ticket. We're very descriptive in like what needs to be done to fix the issue. So they don't need to like enter the platform. They don't need to use it locally. Like, it's, you know, very, like, thorough in what needs to be done to fix it and close it out. So we have— I mentioned, like, this vulnerability funnel where all the tickets go. We also have a platform that was built that— so, like, whenever I'm creating tickets, whenever I run a scan, and I get my, my output of all the issues, I send it to this service that we've built internally, which is what's creating the tickets. And with that, like, you basically get a hash saying like, this hash is assigned to like this container and the, the like specific vulnerability that might be there. Then the next time it's scanned, if they fix the issue, and that hash doesn't exist anymore, because the issue is no longer, no longer persists in that output from the scan, right? So now what happens is, we send the results to that tool, in the absence of that vulnerability, tells the tools like, okay, this doesn't exist. So I'm going to close out that, that ticket now. And so that makes it like immediate. And with the event-driven architecture, like as soon as they deploy and the scan completes, ticket's closed, they get the immediate feedback that you fixed the issue and you're good to carry on with the rest of your life. But yeah, that, that makes it much, much easier. And like the fact that the engineering teams don't have to do anything, they don't have to set anything up themselves. it makes scaling that process much more enjoyable for me and also for the development teams that they just get it out of the box now as part of their build process.

32:28Chris RomeoYeah.

32:30Will RatnerOkay.

32:30Chris RomeoYou've mentioned vulnerability funnel a couple of times, and I have to know what that is now. I've heard it and I'm like, I got to know more about this. So give us some more context on how are you using this term vulnerability funnel?

32:43Will RatnerYeah, so it's, it does what, like, it, you know, it speaks for itself, right? Like any vulnerability from all our scanners that are discovered and tickets are created go into this one project in Jira, right? So like, if I want to search for any team that has like, might be affected by like a specific CVE, I can just go into this project and search. And I have like a full list of all the tickets associated, who they're assigned to, what teams they belong to. And this also helps teams like, We have different dashboards if they want to see how many tickets are associated to like a specific product or a team related to that product. All that information is there in one centralized place. So from like your SCA scanners, your, your host-based network scanners, like any tickets that are created from those scanners all get funneled into this one Jira project. And having that centralized project makes it easy because one, everyone knows where to go. see if they have any vulnerability tickets. And we can set up a ton of automation around like using Automation for Jira in this project to make sure like they're assigned to the right person, the team and like product information on the ticket is populated. You can build notification alerting off of that, which is also really valuable, right? Because, you know, we mentioned like, we were talking about earlier about that, we're not breaking builds, right? But we have very strict SLOs at Atlassian, like If it's a critical issue, 14 days to fix it. High, 4 weeks, blah, blah, it's publicly posted on the website. And so like, they get due dates on these tickets. And this is when they need to be due. And as those dates approach, like you get notifications on Slack telling you, you have a ticket that's approaching your due date, you need to fix this. And it'll escalate as needed. And so like all of that automation is built into that centralized vulnerability funnel, which makes like our rate of breaches of that service level objective to be very, very, very low because people do really well in getting these issues fixed on time.

34:41Chris RomeoAnd you've got great metrics, I'm guessing, coming out of that because you have everything in one place. You don't have to go build another metrics infrastructure. Your vulnerability funnel is where you're getting those metrics from, and you can then dashboard and manage individual teams against you know, where they are and kind of see how products stack up. So yeah, that's, that's pretty, a pretty fascinating thing. So we talked a lot about tickets. So from a kind of lessons learned perspective, I'm curious if there's any other, like what's one or two other lessons learned you would take away from this that could help some of our audience members who might be, maybe they're not doing container scanning, but they're thinking about a centralized approach to something inside of their big company. What, what are some other lessons learned that you would share with them?

35:26Will RatnerYeah, I mean, I think just getting— like, you hear this all the time, but like, having the data is so, so important to figure out like how you're going to actually— like, what needs to be built, what needs to be focused on. So, um, I think like regardless of what scanner you're using, like, so I would say if you're rolling out any kind of scanner, make sure their APIs are good. Um, if you're building a centralized, uh, tool or like similar funnel, you need a good source of attribution data. So whether that's your Active Directory or like some Workday database that you can access, like you need to be able— you need to have good attribution data to know like what teams have what people and how you can get that information, because that's going to make your metrics that you're able to get out of this much more useful. If you just have like a list of people that aren't assigned to like a specific team, or you don't know like what division or department they're in, it's not super valuable. Like you'll just have this person has this many tickets and maybe they need some help, or you can find some single points of failure. But having that, that attribution data is, is massive. And then just kind of like trying to build things as generic as possible that'll fit as many like options of like various ways people might be doing things in your company. So like you mentioned, I mentioned like the platform as a service that we're reading off of, but we also have people that aren't using that platform as a service. So what we built is we have an API that they can put in their CI/CD pipeline because it's a much smaller amount of people, but then it goes through the same process that the people that are using the platform get as well. So an extra step because you're not using the platform, but again, like building things out and thinking of like every possible avenue where people might like be building something or deploying code, you want to try to consider all of those. And then obviously like communicating with the teams is like a huge piece in building that trust to make sure like you can build a centralized solution, you can have all of that information, but If you're, like, filing a bunch of garbage tickets that aren't very useful, or just, like, you're misfiling a lot of tickets, you'll lose their trust, and then you're not going to get the remediation efforts that you're hoping for.

37:54Robert HurlbutSo, Will, what kind of resources would you recommend? Somebody's trying to do the same thing here. They're trying to build out a centralized capability, and they don't know where to get started. what would you recommend?

38:07Will RatnerYeah, so this can depend again, like, on your company size and maturity. What I would say is, like, really important, especially for, like, building out a centralized container scanning solution at scale, is to first, like, understand, like, how containers work and how to build secure containers in the first place, right? Because you're going to want to be giving that direction to the teams of who you're going to be creating tickets for, or however you might get that in front of them. So there's, there's a lot of different resources out there. I mean, you can, you can Google it. I think like Docker has a really good page itself, just around like how to use multi-stage builds and how to build like smaller, more secure images. And that's like what I use to just like get myself familiar and see the best practices. Also just like doing research. I don't have any like specific websites off the top of my head because I was just reading a bunch of stuff. But yeah, that's really like where you want to start from. Like if you're building, trying to build like a scalable solution, I think Argo Workflows are really cool. They have their own site. It was really useful for like building a very scalable scanning solution. And then just like, yeah, I think those are like the big resources we used when we built this. Also relying heavily on like communicating with the engineering teams was like, which again will vary per company and everything, but like talking with them and understanding their build processes, communicating with like the platform teams that build those golden base images or the sidecars that we mentioned are also really important resources in just understanding like how things work within that platform, which again, specific to Atlassian, but it would be useful in any company where you might have something similar. But again, like you're going to get a lot of your information from the teams that are doing that actual work and building those containers before you go and build out that scanning platform yourself.

40:18Chris RomeoYeah, and I love the fact you keep coming back around to communicate with your engineering teams, get feedback from your developers. That's such an important piece of any solution that you're offering to them because so often, in general, people will build something and then the developers are just supposed to live with it. It's like, we don't really care what you think. You're just going to have to do your work out of it. And you keep reiterating, communicate with them. Get their input, get their feedback, you know, be, you know, a source of value for them, not a source of like, ugh, this platform again sending me a ticket. So yeah.

40:55Will RatnerAnd if you lose that trust, it's going to make your life as a security engineer like so much harder.

41:01Robert HurlbutYeah, that was my takeaway. It just so much relies on trust between the teams, the security, the developer teams, and that communication, keeping those lines of communication open.

41:11Chris RomeoYeah, definitely. So, Will, thank you so much for sharing this insight with us about this centralized scanning solution. I know I certainly learned quite a bit in this process, so thank you for educating me a little bit about how you do this at scale, because something that we all got to be— we all got to be thinking about, how can we— how can we do these things better and at scale so that developers think about that tool and they're like, yeah, it's okay? If we get developers to say that platform you built is okay, We all know that's like a win. We won. Like, we love it. So, Will, thanks for taking the time and look forward to following what you're working on into the future as well.

41:46Will RatnerAbsolutely. Thanks, Robert, Chris, for having me on. I appreciate it.

41:50Chris RomeoThanks for listening to the Application Security Podcast. You'll find the show on Twitter @AppSecPodcast and on the web at www.securityjourney.com/resources/podcast. You can also find Chris on Twitter @edgeroute and Robert @roberthurlbut. Remember, with application security, there are many paths, but only one destination.

7,929 words · transcript by assemblyai

More on Cloud and Infrastructure

View all episodes →

Get Reasonable AppSec: new episodes and useful picks from the archive.