Aaron Rinehart -- Chaos Engineering and #AppSec
With Aaron Rinehart
How do you know a security control will work when the system around it fails? Aaron Rinehart introduces chaos engineering as a way to test assumptions about complex systems before an unexpected incident exposes them.
Audio hosted by Buzzsprout. Nothing loads until you press play.
Episode chapters · 14 chapters
- 00:00Chaos engineering and AppSec with Aaron RinehartAudio
- 03:31Security as a reflection of good engineeringAudio
- 06:52What chaos engineering meansAudio
- 09:57How Chaos Monkey worksAudio
- 12:30Principles and the wider movementAudio
- 14:13The relationship with site reliability engineeringAudio
- 15:56Examples of complex-system failuresAudio
- 20:28Applying chaos engineering to securityAudio
- 21:13Introducing ChaoSlingrAudio
- 24:50Testing failure conditions rather than adding flawsAudio
- 25:58Resources for learning moreAudio
- 29:17Release It and resilience lessonsAudio
- 31:33The SOC-less model at NetflixAudio
- 35:04Aaron’s writing and further readingAudio
About this episode
How do you know a security control will work when the system around it fails? Aaron Rinehart introduces chaos engineering as a way to test assumptions about complex systems before an unexpected incident exposes them. He explains its relationship to good engineering, Netflix’s Chaos Monkey, and site reliability engineering, then brings the discussion into application security. Chris and Robert explore ChaoSlingr, the difference between testing a failure condition and introducing a vulnerability, and how experiments can reveal gaps in detection or response. Aaron also discusses a SOC-less model that connects security signals with the people best placed to act. Books and foundational resources round out an introduction to learning from controlled experiments and building confidence in how systems actually behave.
The Application Security Podcast is brought to you by Security Journey.
About Security Journey
Security Journey provides application security education for developers and everyone in the software development lifecycle.
→ Learn more about Security Journey
Connect with Aaron Rinehart:
→ Aaron Rinehart on LinkedIn
Resources
→ Principles of Chaos Engineering
→ Chaos Monkey
→ ChaoSlingr
→ Release It! Second Edition
→ Google SRE books
Actionable
From this conversation
- 7:05
Study chaos-engineering principles
Uh, so I'll lead it with, uh, if you go to principlesofchaos.org, it's the marquee seminal, uh, doctrine when it comes to chaos engineering that everyone sticks to.
- 10:00
Run experiments during staffed hours
What Chaos Monkey does is during business hours, that's a key part, during business hours, it will, it will randomly bring down a VM on one of Netflix's production systems.
- 24:08
Partner on failure hypotheses
That's what the focus is of chaos engineering, is to sit down with your product team, your app team, and say, okay, hey, we think we're seeing these failures happening.
Transcript · 37 min conversation
0:00Chris RomeoSeason 4, episode 11 of the AppSec Podcast. On this episode, we talk about this thing called chaos engineering, which I really didn't know a whole lot about before we started. Robert and I get a chance to talk with Aaron Bryson, who's somebody who's done a lot of research and a lot of thinking about how do we embrace this whole chaos engineering thing in the world of security. So we hope you enjoy.
0:23Robert HurlbutThe Application Security Podcast.
0:32Chris RomeoHere we go. Hey folks, welcome to this episode of the Application Security Podcast. On this episode, we are joined by Aaron Reinhardt. who is going to talk about something called chaos engineering. And, uh, cards on the table, I'm not even really sure what that means at this point, but that's okay. I'm here to learn too, just like, uh, the rest of you in our audience. So Aaron, we always start with this single question that our audience knows and loves, and that is, what is your security origin story, or how did you get involved in security? How'd you get into this discipline?
1:28Aaron RinehartWell, first of all, I want to say thanks for having me on the show. It's really an honor to be here. So my background, my story. Well, I wasn't bitten by a spider, right? Nor do I have superhuman powers, but my story is I got into security— let me back up. So I started off in more of the system engineering, network engineering space. And early on with my work within the Marine Corps, Department of Defense, had some great, great experiences, sort of unique to that field that sort of propelled my career early on. Soon after I— soon after that, I sort of found myself getting into— I had trouble getting actually jobs because of my age. at the time in that space. So, but what's funny is I couldn't be hired on as an experienced network or system engineer, but I could be hired on as somebody, a software programmer that didn't know how to write software. So I found my— I began in the sort of the backend database space with, you know, PeopleSoft and Oracle and soon found myself into the frontend sort of design and then application design and ended up becoming a software engineer at NASA. I spent quite a bit of time there designing applications for NASA, and while I was there, I had the opportunity to start experimenting with security, with security. And it was— that's sort of where my career in security began. But I like to think that I'm— my background in the— my multidisciplinary background is what makes me a good security professional, because from my perspective, security is really just a reflection of good engineering.
3:31Chris RomeoYeah, what do you mean by that though? That's an interesting statement. Tell me more about what you mean about security as a reflection of just good engineering.
3:45Aaron RinehartWell, so for example, a lot of A lot of, like, certain— like, let's say, for example, we used to do sort of— when it comes to traditional, like, network design— I don't know why this first thing popped in my head— we would do— we would separate out components by trust systems, by trust boundary. Well, really, that has always been the sort of good network engineering, was to, you know, is to sort of create those sorts of separated topologies. I guess that's not the best example, but writing good code, right? Should that be— flaws in code, should that necessarily be a security thing, or is that just a good practice? I don't know.
4:39Chris RomeoYeah, and I think that— yeah, that makes sense to me, and that's It's kind of an interesting way of thinking about this. Like, we've always thought of security in the industry as, and even in how it's approached in organizations, both small and large, as a separate entity that's always been, well, that's the security department. And you still see that a lot of different places. And it sounds like what you're kind of thinking about is security really is embedded into everything that's happening. And good security just means good engineering practices, and they're really reciprocating to each other.
5:17Aaron RinehartYeah, it's really that reflection of the quality and depth and introspection in the engineering, and it's really about quality. And I think most of the industry, given the whole DevOps transformation, what that's brought to security, is we're really coming back to our core roots in engineering. I think we began out of the engineering problem, and we went on this weird compliance trip for 10 years, and we're coming back to our roots. It's really exciting to see security not be this weird compliance focus. I mean, compliance is important, right? But the real problem is still the engineering.
5:53Chris RomeoYeah, and you can go ahead and make fun of compliance. It's okay. It's an accepted thing here. on the podcast of— but I mean, you know, compliance serves its purpose. And it's, as long as it's not the basis of your entire program, there's a certain amount of things that we have to do from a compliance perspective. But I always am a little leery when people are leading with compliance as the driver and what makes us tick as a security organization.
6:23Aaron RinehartI mean, you're chasing your tail. I mean, like, you know, I mean, look, I don't think there— I think the first Docker container security guidance came out from NIST 2 years ago. Containers have been around for 8 years. That means there was a good 6-year period where there was no compliance guidance or standards around it. If you weren't focused on the good engineering, you were just open to vulnerabilities. Yeah.
6:52Chris RomeoLet's transition into this whole idea of chaos engineering. If you can give me just a basic definition that we can kind of build off from. When you say chaos engineering, what does that actually mean?
7:05Aaron RinehartSure. Well, just, just, uh, so I'll lead it with, uh, if you go to principlesofchaos.org, it's sort of the marquee seminal, uh, sort of doctrine when it comes to chaos engineering that everyone sticks to. It's the principles laid out by Netflix, the Netflix team led by Casey Rosenthal. But the definition of chaos engineering is the discipline of experimenting on a distributed system in order to build confidence in the system's capability to withstand turbulent conditions in production. All those words in that definition are very important. We can actually get into it at some point in the discussion about what some of the differences between security practices like red teaming, purple teaming, and chaos engineering, and there are major differences. But the idea of chaos engineering sort of was born out of a particular need from Netflix. So there are actually roots of chaos engineering that began at Google in the SRE program at Google, as well as Amazon.com. Netflix, as many folks may know, consumes about almost 30%— I don't know the exact numbers today, but they consume about 30% of all North American internet. On top of that, they consume about the same percentage of Amazon Web Services. When you're that heavy from an engineering perspective, stuff starts breaking. What Netflix decided to do is that they decided that they weren't that their business heavily depended upon their services being available for people. You can't watch a movie if it's not there. They decided they're going to build their system to be resilient to failures from Amazon and network traffic type of failure modes. They went forth and they built to the best of their capability circuit breaker patterns and resilient types of engineering approaches to combat the problem, but they ran into a scenario where, okay, now that we— okay, now that we have done this, how do we prove it? So they went to Amazon and they said, hey, will you break your system for us? Well, no. Amazon was like, well, you know, that kind of violates our agreement with you. You know, you want us to what? And but Amazon said, you know what, you can— you know what, you can go ahead and Break it yourself, right? Have at it, right? And so out of that was born Chaos Monkey, which is still considered the sort of the first tool, the seminal tool in the space. Most people, if they've ever heard of chaos engineering, sometimes they've heard of Chaos Monkey just 'cause it's kind of a crazy name.
9:57Chris RomeoOkay, what is that? What is Chaos Monkey and what does it do?
10:00Aaron RinehartSo what Chaos Monkey does is during business hours, that's a key part, during business hours, it will, it will randomly bring down a VM on one of Netflix's production systems. It is during business hours for the fact that during business hours, you have all the right people there if something goes wrong. But the idea is that when they bring down that VM or the AMI, it's to measure things like, okay, did the service failover? Did we lose traffic? Was there data loss? It's testing all the things we built, but it's also— chaos engineering isn't about validation and verification of things you built per se. I mean, that is one— you are validating certain things, but the idea is actually to learn new things you did not know before. Let me explain how I break down instrumentation. I break down instrumentation in 2 major categories. One, testing. So testing is the verification or validation of something you already know. In security, it could be a CVE, an attack pattern, a signature, what have you, right? Experimentation is rooted in the scientific method and rooted in the idea of seeking new information that was previously unknown, because with chaos engineering, we're trying to We're trying to learn about what we didn't know. So things like security incidents, often I find that people are trying to guess— usually first and foremost, the focus is on the remediation, getting the thing back up and running more than it is the root cause. But I find when we are doing that sort of postmortem, post-incident, root cause, whatever you're calling it, analysis, one, it's usually poorly done, but two, we often associate the problem to what we knew without taking into consideration the fact that there may be lots of things in the equation that we did not know. Because most often, especially in the world of complex adaptive systems, it's not usually one thing, it's a combination of multiple things. And it's about— so with chaos engineering, it's about the learning about how your system really works versus how you thought it did. That's a lot to take in.
12:30Chris RomeoYeah, most definitely. It's, uh, but it's, it's making sense to me. So just to kind of recap the— so Principles of Chaos is where it all is, kind of where it all came from. And this is Netflix, they've documented what they were— they're, they're doing internally to share that with the rest of the industry, right? That's, that's the point of that.
12:54Aaron RinehartExactly, exactly.
12:55Chris RomeoAnd then they use this, you know, and I had heard of this Chaos Monkey tool before, um, but didn't really associate it with a bigger movement. So, um, so how much of this is a movement? Like, is this—
13:08Aaron Rinehartis this—
13:09Chris Romeodoes chaos engineering have the same amount of velocity that like DevOps does, where everybody's doing DevOps now? Is chaos engineering gone? Has it gone mainstream now, or is it still something that most people don't know about?
13:23Aaron RinehartI would say it's going mainstream now. It's like the next evolution in that whole category. I have a— so my All Day DevOps— I'm one of the All Day DevOps speakers this year, um, on the SRE track. My, um, I guess I'm like the only security SRE kind of combination, um, but, uh, uh, I'm gonna talk a bit about how the relationship between SRE, DevOps, and chaos engineering, it all kind of centers around failure, to be honest. Okay, go ahead. Sorry.
13:53Chris RomeoYeah, no, and just so for, for the listeners and, and, uh, and even my understanding, SRE is Site Reliability Engineering.
13:59Aaron RinehartYep, Site Reliability Engineering. Uh, the definition comes from Google. Um, uh, but like many people move beyond the word site and use more System Reliability Engineering.
14:11Chris RomeoSystem.
14:12Aaron RinehartOkay.
14:13Chris RomeoAnd so, so that's kind of the discipline for which chaos engineering Chaos engineering is kind of the movement, and SRE, or system reliability engineering, is kind of the process and the moving parts that you assemble to do chaos engineering?
14:30Aaron RinehartWell, you made a great distinction. Chaos engineering was kind of born out of— I mean, it kind of was in the same vein as SRE.
14:42Chris RomeoOkay.
14:43Aaron RinehartAnd there is a lot of similarity, and I have an article and a talk coming up about this just to sort of drive some clarity around it, where they overlap. But it is about, you know, it's about the core tenet is failure, right? It's using failure as a tool and treating it the right way and building things the right way. And in the end, it's about building a learning culture about how your systems really operate.
15:15Chris RomeoAfter the break, Aaron explains the business value of an investment in chaos engineering and SRE. The Application Security Podcast operates with support from Security Journey. A security belt program provides the 3 pillars of successful AppSec training. Learning, application, and experience. Visit us on the web at www.securityjourney.com to learn how you can teach and empower your developers using a new kind of security training. Aaron dives back in with the tangible results of chaos engineering and SRE.
15:56Aaron RinehartSo that's a good question. So, um, I can give a couple case studies examples.
16:03Robert HurlbutOkay.
16:03Aaron RinehartSo the nature of the problem, and the nature of the business case, is that we— our systems are no longer these linear 3-tier apps running out of data centers, right? We are in the era of the dawn of the complex adaptive system, and there is science and clear definition on what that means, meaning that complex adaptive systems are are made up of many different parts. Those parts are networked, a network of relationships that as a whole, the relationship and the nature of these things and how they're interconnected create nonlinear outcomes. What I mean by that is, so, is that linear outcome would be cause plus effect equals outcome. Meaning 1 1 2. In the world of a complex adaptive system, we reach nonlinearity, which— what that means is that 1 1 may equal -90 or 100, in that it's the relationship of feedback loops between things. Think about a system like Amazon.com with thousands of microservices running to facilitate that app.
17:24Chris RomeoRight.
17:26Aaron RinehartRight? You're having things— things are spinning up, they're spinning down, they're interconnected, and, you know, everything relies on each other. And so it's the nature of the cascading effects between those things that cause exponential waves. It's not— so complexity science and complex adaptive systems is where the chaos theory idea is born.
17:52Robert HurlbutRight.
17:52Aaron RinehartThe butterfly effect, right, is that one change can ripple throughout an ecosystem and create nonlinear outcomes.
18:03Chris RomeoThe example I've seen, and I'm by no means an expert in anything related to chaos, but isn't there an example with the butterfly, like the butterfly flaps its wings and 1,000 miles away a hurricane or something?
18:15Aaron RinehartYeah, that's it. That's it. Really, actually, Jurassic Park, Malcolm, what's his name, does also a great example of it in pop culture. When you ask me a question, what is the business outcome? The business outcome is that large-scale distributed systems have unpredictable outcomes. Look at Amazon's Prime Day. I think we all remember that was about a month and a half ago.
18:40Robert HurlbutYes.
18:44Aaron RinehartA 1-hour outage cost Amazon.com— I mean, this has been— this is sort of what publications have determined what the cost was. I mean, Amazon was not— just to be clear, they were not openly saying we lost $33 million an hour, but it was roughly calculated that's about what they lost. It's $33 million an hour and they're out for 3 hours.
19:04Chris RomeoOuch. That'll make for a tough Christmas this year, you know.
19:08Aaron RinehartYeah, but here's the thing is it's like is it used to be easy to know what a system was doing. But now, think about how our release cycles, delivery cycles, and our build cycles are working. Let's say you've got an app and it's got 10 APIs, for example. There's not actually 10 APIs, there's probably at least 10 for each API for throughput. There's probably a blue-green deployment or a type of scenario where they've got multiple versions they're testing.
19:35Robert HurlbutYeah.
19:36Aaron RinehartAll those things are interconnected. They're spinning up, spinning down, and teams are releasing, hopefully at the same schedule, but some at different schedules. So things are rapidly changing, spinning up, spinning down, and they're all interconnected. We're at a point no longer where the human can ascertain what exactly is happening at any given point in time.
19:54Chris RomeoYeah, I'm kind of coming to the conclusion here that the business case or the reason is in this complex adaptive system, there is really no way to do functional testing. Because there's so many different components, your test plan would be tens of thousands of pages long to try and cover all of the possible scenarios that could happen.
20:14Aaron RinehartExactly. I mean, we can do some, right? I mean, like chaos engineering doesn't replace existing testing methodologies, but you're, you're on the right path for sure. Um, is, I mean, we, um, yeah, I mean, that's, that's, yeah, definitely.
20:28Chris RomeoSo I think, I feel like we've got a good introduction now to, or probably, you know, at least an intermediate level introduction to, uh, chaos engineering. But since this is the AppSec Podcast, let's now transition into what are the ramifications for application security in combination with chaos engineering? Where do these 2 things intersect with each other, and what do those connections look like?
20:54Aaron RinehartSure. So a good example of— so I could explain a bit about Chaos Slinger. which is the first sort of open source tool that we released that does some of these chaos engineering type of activities.
21:13Chris RomeoWas that Chaos Slinger?
21:14Aaron RinehartYeah, Chaos Slinger.
21:16Chris RomeoOkay, got it.
21:17Aaron RinehartYeah, so originally, okay, full disclosure, this is a cool podcast, so I'm going to give the cool name out. Definitely, it was not okay with most people to release under this name, but Originally, the name of the project was Poo Slinger because we figured if we were going to create a chaos engineering tool for security, I mean, what better? There's all these monkey tools like Chaos Monkey, Gorilla, Kong that we're thinking about. Okay, well, what do monkeys throw, right?
21:48Robert HurlbutYeah.
21:49Aaron RinehartThey throw poo. We called it Poo Slinger. It made the project fun. 5-year-old jokes, you know, every meeting. But the idea of— so from an AppSec perspective, or from a security perspective in general, is that the nature by which an incident occurs is somewhat subjective. You don't know where it's coming from, who it's coming from, when it's gonna happen, and all your preparation, it could be the logging, it could be, you know, the monitoring, or the attempts to make your system more observable, which is a, a very new concept that also overlaps with chaos engineering, by the way. But no matter all your efforts, preventative or detective engineering focus, you're still at a level of unpredictability and subjectivity to an incident. Well, and so when an incident occurs, the focus is, like I said earlier, it's mostly on the remediation, not on actually what happened. But even if you were, seeking what happened, or information about what happened, you most likely do not have the right data or enough data to make sense of it. But what if, what if, okay, what if, you know, you were able to cause the objective event itself, right? What if you were the root cause and you could inject the signal and measure it throughout the cycle? And what I mean is you're not With chaos engineering, in terms of its response to this example, you would inject the failure mode. It could be a misconfiguration or misconfigured IAM policy in AWS. It could be a misconfigured port, or it could be a dynamic— another example could be a dynamic WAF rule that was created through the machine learning on Amazon's WAF. Anybody who's dealt with that knows that not all ML and AI makes very— they don't all make good decisions.
23:47Chris RomeoRight.
23:48Aaron RinehartMost often, they don't make very good decisions at all. But failure happens. But what if we were able to control when it did and learn about how well we were prepared for the situation when the world is not on fire?
24:07Robert HurlbutYeah.
24:08Aaron RinehartThat's what the focus is of chaos engineering, is to sit down with your product team, your app team, and say, okay, hey, we think we're seeing these kind of failures happening. Are we seeing these types of events happening? We're not sure why it seems to overlap in this, and let's say it seems to be report-related or seems to be related to this type of activity where we're seeing a block. We want to learn more about how— what other pieces didn't tell us that should have told us. what's happening so we can make changes or increase our visibility. Does that make sense?
24:50Chris RomeoYeah, that does make sense. I guess from the incident perspective, you're not actually causing a security incident then per se. You're not injecting a SQL injection or something into a running application so that then you can track and see what's happening. You're more generating failures that would be the logical result of an incident and then using that to measure how well the system responds, how adaptive it is, does it fall over and completely die? Do I have that right? Am I understanding that correctly?
25:27Aaron RinehartYeah, I mean, I guess in a gist, yeah. I mean, it's failure injection predominantly, right? You're injecting failure modes. I mean, they're key. This is where it comes back to some of the chaos principles. It's important to stick to them. Anytime you inject failure, you need to ensure that you control the failure domain and the blast radius, per se. Let's see. I think I answered your question.
25:58Chris RomeoYeah. What are some resources then that somebody who might be listening to this episode and might be thinking, I want to learn more about this whole chaos engineering and kind of figure it out. And I know we talked about the principlesofchaos.org. We'll put that in the show notes. We'll definitely put a link to your article and your DevOps Day talk as a place for people to get more information. But what other resources are available out there that you recommend to people?
26:33Aaron RinehartOh, I've got a bunch. I think it's important before you dive into chaos engineering to understand nonlinear systems and complex adaptive systems. On YouTube, there's a series. It's super— it's freaking awesome. Somebody broke it down easy cartoon style.
26:54Robert HurlbutOkay.
26:55Aaron RinehartIt's a 55-minute course. You can actually break it down, maybe watch half of it, but it's from Complexity Labs. They just make it easy to consume. It explains what these systems are. Maybe if you're out there and you're trying to figure out why unpredictable things keep happening with the system, maybe this is your first best way to dive into the understanding behind why chaos engineering and observability are important. On top of that, there's also another— there's a seminal paper that should be a must-read for anyone in today's modern engineering. It was written in 1983, 35 years ago, by Lisanne Bainbridge. It's called The Ironies of Automation. There's all this focus in security, especially on automation, on ML and AI, but it's important to understand the problems that come with automation. Automation— I don't say automation is bad, but we assume it's the answer or an answer that is going to just solve it for us. Automation comes with a tax, right? With a tax being that— so this paper is called The Ironies of Automation, written in 1983. It should be free out there. And the third resource is Why Complex Systems Fail. It was a paper written by MIT.
28:21Chris RomeoOkay.
28:22Aaron Rinehartthat also describes— it's important to grasp more of the domain of resilience engineering before diving into chaos engineering. That way, you understand the nature of network theory, emergence theory, chaos theory, and why chaos engineering is a thing. But also, I think we're in an era where if security doesn't start understanding that problem, we're going to continue to design stateful security in a stateless scenario. Cool.
29:05Chris RomeoRobert, you have any— you've been kind of taking this all in.
29:09Robert HurlbutYeah, this is great.
29:12Chris RomeoA lot of information. You got any specific questions for Aaron?
29:17Robert HurlbutThere are a couple of resources there that I was not familiar with, so I definitely want to follow up on those. One other resource I was going to mention, and we kind of talked about it before we started recording today, was the Release It book by Michael Negard, which there's a second edition that just came out, and he added a chapter on chaos engineering as well. That's really where I started first learning about 10+ years ago when the first edition came out about failure. in architecture and how you need to think about failure. How much of that has influenced where we are today as well in chaos engineering and so on?
29:52Aaron RinehartOh my gosh. I mean, it's right in there.
29:59Robert HurlbutI remember that's where we first learned about circuit breaker and bulkheads and all those kinds of things, those patterns that it seems like that's where That's where we need to be thinking about, and now it's come to fruition with this chaos engineering and so on in terms of building resilient systems.
30:18Aaron RinehartOh, precisely, and security is a core part of resilience, in my opinion. Simply, resilience, all it means is the ability to respond, you know, so failure is a tool. NyGuard really, you know, hits hard at. And, and I was just actually had a great conversation with someone writing an article on the topic of detection engineering and with chaos engineering within it. And chaos engineering is very, from a security perspective, is very much a part of that, where detection engineering is really focused on, you know, trying to empower the human in sort of the response. Depends on if you follow Alex Mistretti from Netflix's methodology with the SOCless model, which is more DevSecOps, but it's the idea of creating more context around alerting and reducing alert fatigue and things that come along with that. Chaos engineering adds to detection engineering in the fact that it's about what you did not know. It's about the unknown unknown.
31:33Chris RomeoNow, what is this? Can you tell us more about this SOC-less model at Netflix? And you mentioned in regards to DevSecOps, I'm just curious what that is and kind of what goes into that.
31:43Aaron RinehartSure. I was originally— there was an article that was posted on LinkedIn from Alex Maestretti. He's Netflix, I believe, engineering manager. It's about not having the need for a SOC, a core sort of core central component responsible for responding to alerts from systems and things like that, launching into response forensics, lacks very little context about systems that the alerts are coming from. And what I mean by that is, so what Alex is articulating, which I have sort of articulated more as incorporating security into the SRE function. I think those two fields, the fields of application security and SRE, are a perfect-made match in heaven, right? It's like SRE solved everything but security. Well, it's frustrating, but like— But the point of it is that distribute the SOX functions to the a DevSecOps model where, uh, what if you had a security person, right, that had innate depth, depth and knowledge about how the product itself worked, where they could ascertain quickly whether or not the alert was just junk, right? It made no sense. Well, we get all these alerts that are guessing somewhat about what might be happening. Remember, our products and services are intellectual property, right? They're what makes the business money.
33:17Chris RomeoYeah.
33:18Aaron RinehartThere's uniqueness to it. Often security, historically, through alerting, through AppSec scans, through whatever, we get this telemetry that says, okay, hey, App team, product team, you did something wrong. You go back to them and say, well, no, not entirely. That's not how that works. Then they explain it to us and then we say, no, that's not what we're seeing in the tool. We're explaining a tool, they're explaining the engineering they built. It's this back and forth. What it really lacks is there's empathy lacking from the security side of what was done to build the depth in engineering and the product. But what if there's a security person, by the nature of their role, they had the depth in the engineering, they knew how the thing was built, they knew how the security interacted, and they could easily ascertain, well, no, that's not how that works. Actually, they did this. which is kind of cool, kind of unique, but it makes it kind of this scenario not really feasible. The tool just doesn't understand it. And we've, in the AppSec space, we've all been there. We've all seen false positives due to uniqueness in product.
34:23Chris RomeoYep. Yeah, definitely still a challenge even today, almost 13 years after I stopped working in the world of incident response and event management. It doesn't seem like— I mean, the tools have gotten faster and they can collect more data, but it still seems like the model is still traditionally by— of having a bunch of people in a room 24/7 that are watching this stuff. And so yeah, that makes sense in the new DevOps world, put the alerts closer to the people that can actually triage them and really know what's happening behind the scenes. So yeah, I think that's a good approach.
34:57Aaron RinehartThey're gonna know exactly the person that needs to get involved, right? You know what I mean? Like, It just makes sense.
35:04Robert HurlbutHey, Arun, you also have a series of articles on opensource.com, is that right?
35:10Aaron RinehartYeah, that's right. That's right. It's sort of been sort of a healthy exercise for me to help other folks understand. It's a culmination of a lot of questions I get from folks about the differences between, let's say, purple testing and chaos, or SRE and security. A lot of stuff we talked about today on the podcast, you can dive deeper if you go to opensource.com and just search for my name. You'll see quite a few resources.
35:40Chris RomeoDefinitely. Aaron, thank you for taking the time today and sharing what is obviously a large amount of expertise in this area, and thank you for summarizing it in such a way that I took away a real good set of notes as far as what what this chaos engineering thing is and some great resources to take a look at. So we thank you for taking the time, and we'll have to have you come on again in the future to maybe unpack one of these more complicated topics that we just kind of got a chance to just introduce. Have you come back and we'll dive deep into one of these pieces in a future season. So we thank you very much.
36:17Aaron RinehartI love that. Thank you. Thanks for listening to the Application Security Podcast. If you enjoy the podcast, please do us a favor and visit the iTunes Store and give us a 5-star rating. Our intro music is 8-Bit Kung Fu by Born and TJ, and the outro is Southern Delight by Stefan Cartenberg. You can find us on Twitter @AppSecPodcast or on the web at www.appsecpodcast.org.
5,417 words · transcript by assemblyai
More on Privacy and Compliance
View all episodes →- August 20, 2021 · 36 minEran Kinsbruner -- DevSecOps Continuous Testing
- January 30, 2020 · 38 minDJ Schleen — DevOps: The Sec is Silent
- May 14, 2020 · 25 minMarc French, Steve Lipner, Maya Kaczorowski, DJ Schleen, Kim Wuyts — Season Six Wrap up