Aaron Rinehart -- Security Chaos Engineering
With Aaron Rinehart
Secure DevelopmentSecurity TestingCloud and InfrastructureDevSecOps and CI/CD
Aaron Rinehart is expanding the possibilities of chaos engineering to cybersecurity. He began pioneering security in chaos engineering when he released ChaoSlingr during his tenure as Chief Security Architect at UnitedHealth Group (UHG).
Audio hosted by Buzzsprout. Nothing loads until you press play.
Episode chapters · 12 chapters
- 00:00Meet Aaron Rinehart: Security Chaos EngineeringAudioVideo ↗
- 03:46RightAudioVideo ↗
- 16:56Yeah. Let me ask a couple of clarifying questions too, becauseAudioVideo ↗
- 21:31Yeah, but I mean, you think about how much time, moneyAudioVideo ↗
- 26:05You say it succeeds, you mean that the system detected anAudioVideo ↗
- 28:59Most people doing these manual right nowAudioVideo ↗
- 33:05Let the record show I'm about to ask a product-related questionAudioVideo ↗
- 34:14WhatAudioVideo ↗
- 36:59Yeah, that's helpful. So, I want to stop for a secondAudioVideo ↗
- 39:40There's entropy too, rightAudioVideo ↗
- 42:30I want to do is we're— as we're kind of comingAudioVideo ↗
- 46:07Great. So, um, one, one kind of last question. So What'sAudioVideo ↗
About this episode
Aaron Rinehart is expanding the possibilities of chaos engineering to cybersecurity. He began pioneering security in chaos engineering when he released ChaoSlingr during his tenure as Chief Security Architect at UnitedHealth Group (UHG). Rinehart is the O’Reilly Author on Security Chaos Engineering and has recently founded a chaos engineering startup called Verica with Casey Rosenthal from Netflix. Aaron joins us to explain what the heck security chaos engineering is. We explore the origin story of chaos engineering and security chaos engineering, and how does a listener get started with this new techniques. We hope you enjoy this conversation with… Aaron Reinhardt is expanding the possibilities of chaos engineering to cybersecurity.
The Application Security Podcast is brought to you by Security Journey.
About Security Journey
Aaron Reinhardt is expanding the possibilities of chaos engineering to cybersecurity.
→ Learn more about Security Journey
Connect with Aaron Rinehart:
→ ChaoSlingr
→ Verica
Resources
→ ChaoSlingr
→ Verica
→ Security Chaos Engineering
→ Chaos Monkey
→ Kubernetes
→ Apache Kafka
Actionable
From this conversation
- 7:41
Continuously verify security controls
What we're doing fundamentally with security chaos engineering is we're asking the security mechanisms, are you doing what you're supposed to do?
- 15:37
Expect controls to need recalibration
It constantly has to be recalibrated to meet the conditions of the system, and we're not doing that.
- 22:20
Run your first experiment as a game day
It's best done in a game day exercise.
- 27:16
Automate outcome validation with control telemetry
You want to hook into some KPI with the control.
- 30:09
Run experiments independently after deployment
It's better to continuously run it on its own schedule and inform you than run from CI/CD.
Transcript · 49 min conversation
0:00Chris RomeoAaron Reinhardt is expanding the possibilities of chaos engineering to cybersecurity. He began pioneering security in chaos engineering when he released Chaos Slinger during his tenure as Chief Security Architect at UnitedHealth Group. Reinhardt is the O'Reilly author on Security Chaos Engineering and has recently founded a chaos engineering startup called Verica with Casey Rosenthal from Netflix. Aaron joins us to explain what the heck security chaos engineering is. We explore the origin story of chaos engineering and security chaos engineering as a discipline. And we also get into how can you as a listener get started using these new techniques. We hope you enjoyed this conversation with Aaron Rinehart. Are you trying to build a security champions program?
0:49Robert HurlbutEveryone is these days.
0:51Chris RomeoOne challenge of rolling out security champions is How do we educate all these new folks? Security Journey has your answer. We provide a Security Dojo environment with level-based security education that gives your newfound champions a path to follow.
1:07Robert HurlbutAnd the best part? It requires almost zero administration by you. Visit www.securityjourney.com to set up a demo and learn how you can use the Security Dojo to connect with your security champions. Hey folks, welcome to this episode of the Application Security Podcast. This is Chris Romeo, CEO of Security Journey, and I'm flying solo today. I guess I've got the reins, which should be slightly frightening for Robert and anyone else involved with the podcast, but that's okay. We're going to have some fun today. Excited to have Aaron Reinhardt join for the second time on the podcast. So, he joined way back in season 4, a number of years ago. He and I both had various entrepreneurial journeys happening throughout that— since that time. But he talked in those days about chaos engineering and application security. That's what he brought to us. You can go back and listen to that initial episode, but we're here today to catch me up on security chaos engineering and You as our audience, you're going to get a chance to listen in on this because I've known Aaron for a number of years. I've known what's been happening with chaos engineering and security chaos engineering, and I still don't feel like I could explain it even to a 5th grader. And so, I reached out to Aaron. I said, all right, I got to get you back on the podcast because I got to understand this better. I got to be able to talk about this intelligently. And so, Aaron, welcome back. And I want to jump right in with my first question. What the heck is security chaos engineering?
2:50Aaron RinehartHey, Chris, thanks for having me again. You know, I'm a big fan of what you do, and so thanks for having me on the show. Yeah, so, you know, I think, you know, I even saw a lot of people struggle with what is chaos engineering sometimes, you know, and I will— so let me explain what chaos engineering is, and I'll explain kind of where security, how it evolved to, or how I kind of evolved it into cybersecurity. And just to give you a background on myself, I've been a builder most of my career. I'm a builder, I build things. And so I was a software engineer for most of that, most of my career before I got into security. So you can't really lie to me how things are built. I know the process. It's hard, it's experimentation, it's exploration. It's a process of trying, that worked, that didn't work. That didn't work. That didn't work. Oh, that kind of works. Yay!
3:45Chris RomeoRight?
3:45Aaron RinehartLike, over and over and over again. Right? Like, and anyway, so chaos engineering is, it's the technique of proactively, I guess, discipline, the discipline of proactively introducing turbulent conditions into a distributed system to try to determine the conditions by which a system or service will fail. before it actually fails. So early on in my journey with chaos engineering, I ran into— I had a conversation with one of the large payment processing companies. And they were talking to me about how they wanted to move this legacy application that was like the flagship. It did all the payment transactions for the company. And it was legacy. It rarely had an outage or an incident. The engineers felt competent. How it worked was well-known, and people felt comfortable with the system. I've been through this thought evolution ever since I started with the chaos engineering stuff. To experience this conversation, I started thinking to myself, was that system always stable? Was it always that way? I mean, when they first built it, did it never have errors or outages or problems? Did the engineers always feel like they really understood it and competent? Well, most likely not. Because we learn about what we didn't know about how the system really worked through outages, incidents. Those are surprises. If they weren't a surprise, you would've fixed it already. You're surprised by that behavior. You can summarize all incidents, outages, and breaches into surprises. It's a surprise. But those surprises informed us, wow, crap, I didn't know that worked that way. Or, hell, I'm surprised that ever worked. It tells us, informs us of what we didn't know. And so we learned through a series of outages, incidents, hopefully not breaches. Well, that brings us to the next question. That's how we learn. That's how systems become stable. Well, that is a very expensive process. What I mean by that is that it comes at a cost of customer pain or business productivity. And chaos engineering, you can think of as a proactive way of accelerating that process, where we're not responding to an outage or an incident. We're not— 'Cause when people are in an incident, They're freaking out. If it's a security incident, it's not the security response people that are freaking out. It's everyone else like, crap, I knew I shouldn't push that code, or, I knew I shouldn't have made that change at 3 AM, whatever. People are worried about the blame-name-shame game. They're worried about being named, blamed, and shamed for being the person that caused the demise. So with chaos engineering, we don't do chaos engineering under those pretenses. Chaos engineering is about learning. It's about continuous learning and continuous adaptation. So what we're trying to do is proactively learn about the system through asking it a series of questions. So the series of questions are in the form of, I believe under X conditions, the system is supposed to— I've designed the system to Do why? And you never do a chaos experiment you know is gonna fail. If you already know it's gonna fail, you're not gonna learn anything. It's a waste of time. But when you do a chaos experiment, it's on what you think you know to be true. And so there's this fundamental first principles problem in all security that chaos engineering solves. I'm getting to the security pieces I'm helping building here.
7:40Robert HurlbutMm-hmm.
7:41Aaron RinehartIs that, well, the first principle, it took me a long time to really understand where the value was coming from in what I was doing. And one of the issues is, is that in modern engineering, what we have is we have this sort of first principles dilemma where, as an engineer, I need the flexibility and convenience to constantly change something. And I'm doing that because the business is asking me to do it, 'cause they wanna deliver value to the customer in the market. So I'm constantly being asked to do that. That's my job. But security, security is a context-dependent discipline. What I mean by that is you have to understand the context of the thing that you need to secure in order to know what needs to be secured about it. And so security is kind of forced into sort of a staple context-focused discipline. in order to properly design it? Well, what's happening is that we're in a situation where when we originally designed the security based on the context we had, it was working. And then what's happening is that you're constantly getting this curve towards business value, constantly changing and updating and trying to deliver product value to the customer. And then the security's kind of staying with it, and you don't really know you've kind of diverged from a drift perspective, from a misconfiguration kind of perspective. Until there's a problem, right? Or some sort of instrument or error, observable error, tells you that there's a problem. Well, that's kind of— it's a little bit late to do that. So what we're doing fundamentally with security chaos engineering is we're asking the security mechanisms, are you doing what you're supposed to do? And we're not doing it, we're not hacking, we're not attacking it, we're not mimicking an adversary, where our goal is not to try to penetrate or get in. Our goal is to proactively verify the security conditions are working. And you can get some of that. Now, the argument can be made with purple teaming or red teaming or pentesting, you kind of are getting that, but it's not necessarily the same goals, right? And the way we do chaos engineering for security is very— we do one experiment, one failure injection at a time, because we have to be conscious of how we're building systems today is that we're building these large-scale microservice types of environments. I feel like the monolith might have been a good idea at this point. It was funny, I had this thought the other day, it's like, will the monolith live in the Hall of Fame of computing like the mainframe is right now? I don't know.
10:14Robert HurlbutWe're gonna get back to it though, because just think about computing in general. We went from mainframes to personal computing to cloud, which if you think about it, cloud's just a big mainframe that's happening. So maybe the monolith's going to come back. I mean, we should make some t-shirts and try to bring it back. I don't know.
10:31Aaron RinehartRight? Bring back the big ball of mud, on with it. No. Yeah. So there's this problem, right? There's this fundamental first principles problem, right? And what we're doing is we're injecting, we're asking the question, hey, does Is that thing we designed, is that IAM rule set that you put in AWS still effective against the conditions you designed it for? Most likely not. IAM is a target-rich environment. Even if you have the greatest security tools in the world, it's still gonna have this problem. And that's what I try to tell, even for customers, like, hey, it's not that Twistlock, Aqua, Capsulate, whatever are bad. They're not. They're great companies, great products. Everybody's got an opinion, okay? But what I'm saying is it's a first principles problem, is that things are changing and there's no feedback loop between the change and the control. The control doesn't know what it's missing. We're making many, many changes today, arguably more than we've ever made before, just because of the way we're distributed. With microservices, you got a team per microservice. You've got 10 microservices, you've got 100 of each of them running, and it's very complex and difficult. This is what we're doing. We're trying to make the system more stable, trying to correct our understanding of how we believe the system is actually functioning. Chaos engineering, people also, I want to address this, people also think it's some super elitist kind of engineer, like, oh, we can barely do the DevOps or CI/CD. Why can't we? I mean, Netflix couldn't do that either. I mean, Netflix was going from a data center to AWS when they developed chaos engineering. They were doing it because there was a lot of unpredictability in what they were building because the cloud provider's AMIs were just disappearing.
12:29Robert HurlbutYeah.
12:29Aaron RinehartLike, crap. Well, our code, We don't know if our code can survive that failure mode. They tried to design for the failure mode, and then they invented Chaos Monkey to actually exercise the failure mode. Then they decided, okay, we're going to do this constantly and just ensure that we're ready for this when it happens, because we start streaming to customers and they start experiencing a bad time, they're not going to hang around for $12.99 of subscription. Anyway, so I just wanna address that. This is something that accelerates. The most common driver for chaos engineering is cloud transformation still. There's a lot of SRE, but really, it's the cloud. Chaos engineering for security. It's really not a whole lot different than chaos engineering. I mean, the use cases are a little bit different, I can explain some of the use cases on where I've gotten value. I'm hearing about other companies out there that are using it in different ways. That's the coolest part of being at the beginning of something. One, it started as an idea. I was like, I wonder what this looks like, trying chaos engineering at security, because it made sense to me as an engineer. I'm like, well, there's no security system in the system. The system is secure. It's not. So it's like, you're just exercising the other conditions in the system. I can exercise the security conditions as well. And the difference between chaos engineering for security and, let's say, so in IAM, so IAM in AWS, for example, just give an example. So let's say, even if you got evidence.io or you got, oh, evidence.io's gone. But if you got Palo Alto's tool now or RedLock or whatever you got, that you need to have those tools.
14:17Chris RomeoYeah.
14:18Aaron RinehartDo not hear me incorrectly. You have to have those tools, but they are reading a config and validating the config. We exercise it. It's interesting enough, we exercise it by introducing the condition that the system is supposed to account for. You would be surprised how often it doesn't account for it. Kubernetes, I'll give you an example of Kubernetes experiment, and then I'll get into the origin story of how I started with this. And there's a good Kubernetes example. Believe it or not, this is a great example. We found a lot of great results with this at a large telecom. And it's just launching a vulnerable image in a container. And the idea is to try to see how long it takes Twistlock, Aqua, whatever solution you have, to identify it, to identify the attack, prevent it, block it. whitelisted, blacklisted, whatever it's supposed to do. We have it tuned, configured. And when we originally launched this experiment, I'm like, you know, it's gonna catch it all the time. It's not a very valuable thing, right? Like, there's no way, right? I was expecting about 80% of the time to be caught, 20% of the time would be the difference. And, you know, I'm learning actually other people doing the similar experiment at where they work, 'cause they're having the same results. About 80% of the time it's not even caught, right?
15:37Robert HurlbutMm-hmm.
15:37Aaron RinehartI mean, 'cause what happens is, Well, you install the tool, right? It gets configured for the environment. I mean, it's hard. Security people, we gotta do a lot of stuff, right? I mean, a lot of stuff, but, like, recalibration is required, right? It constantly has to be recalibrated to meet the conditions of the system, and we're just not doing that. And we don't know we need to recalibrate because we don't know there's a problem yet, right?
15:59Robert HurlbutYeah.
15:59Aaron RinehartAnd so there's drift, right? There is, you know, I'm sure you've been in your career, Chris, in the situation where You know, you bought in a great— you brought in the intrusion prevention system or the awesome blacklisting solution, and you get it going and you're excited. All of a sudden, it blocks 60% of all traffic to the flagship application. They're like, turn that shit off.
16:22Chris RomeoRight.
16:23Aaron RinehartLike, you know, and then, you know, that happens a lot with security controls too. It's like, you know, like, no, too much overhead. No, because we have to— a lot of times as security people, You and I are engineers by trade, but not all security people have an engineering mindset to understand that we impact latency majorly. In a large-scale distributed system, the network is already not reliable. If you're introducing more problems into that, it just makes it harder and less likely that they're going to keep the security route.
16:55Robert HurlbutYeah. Let me ask a couple of clarifying questions too, because like I said off the top, My goal is to understand security chaos engineering. And so, you've helped me, you've kind of laid the foundation for us about the things around it. And so, I want to hit on a couple of things that you said. So, it's not pen testing. So, what is the difference then between security chaos engineering and pen testing from your perspective as someone who is very entrenched in security chaos engineering?
17:30Aaron RinehartSo, I mean, first, just caveat, not caveat, but I put a disclaimer out there. I'm by no means saying pen testing, red teaming, purple teaming, breach attack simulation. I'm not saying any of those things are bad at all.
17:42Chris RomeoYep.
17:43Aaron RinehartRight? They're still good. They give up good, they give us good objective. They're instrumentation, they're forms of instrumentation. We need more of that. We need less subjective kind of human, poor human assessment of a machine. We need an instrument telling us what is happening and give us data to make better decisions because engineers don't believe in 2 things, hope or luck.
18:04Robert HurlbutOkay.
18:05Aaron RinehartSo, like, it works or it doesn't, right? And then we fix it. But, like, so you can think of security chaos engineering as, like, security fault injection. You can think of it as fault injection, failure injection, maybe if that helps you wrap your mind around it, but it's a little more involved than that. But that's kind of like maybe distinguishing factor for you. And it's not every show based, right? It's not, we're not trying to, you know, game the blue team where, you know, it's so interesting. A lot of conversations anymore in security, you have to play a word bingo game a little bit, you know? And I feel like, I don't know if you've ever seen the, I'm writing the new O'Reilly book with Kelly Shortridge right now. And like I'm finding I have to just read it. I have to redefine everything for everybody on everything.
18:50Robert HurlbutYeah.
18:51Aaron RinehartThere's a security color wheel. I don't know if you've seen that thing before, but—
18:54Robert HurlbutYeah.
18:54Aaron RinehartIt's got like 60 different colors. And I'm like, somebody once asked me, where does security chaos engineering fit on the color wheel? I don't know, pink? You know, pink testing, green testing, blue. I was like, oh geez, you know? But anyway, so yeah, you can think about it as fault injection, failure injection. So we're pen testing. So pen testing is, so pen testing, red teaming, purple teaming, And breach attack, well, I'll address breach attack simulation a little differently. But those techniques, so one, red didn't really evolve to purple. Purple became something else, right? And 'cause remember we had this problem with red teaming, like the DevOps problem, dev throwing the results over to ops, vice versa, you know? And we had that problem with red teaming, right? Is that red teams would attack it and they would always have the result and throw it to the blue team and say, hey, you suck, you suck, you suck. Go fix it, right? And they're like, no, I tell you what, you try doing this job, right? Like, it's, you know, it's just not— it was a very toxic kind of approach. And typically red teaming, purple teaming, pen testing is only done on maybe top 5% of all applications. And typically it's either the most highly regulated or it's the flagship kind of core revenue generating kind of, you know, applications or products.
20:12Robert HurlbutYeah.
20:13Aaron RinehartSo, the coverage is somewhat poor, and it's somewhat— the tools for automation are good, but are good mostly in corporate IT, right? So, like, when it comes to— because software is a little more dynamic and difficult to address when it comes to automation because you're dealing with automation. And, you know, so there's— I'm just thinking of some of these pieces. But what I'm saying is all these, all those techniques are still viable. Don't, don't hear me wrong. Please don't hear me wrong. I know I'm gonna get on Twitter and somebody's gonna say Aaron was talking shit on red teaming or purple teaming. And I'm not. But like also like we're not trying to, like this is more purple teaming and breach and attack simulation. We're not trying to simulate an attack. We're not trying to see the effectiveness of attack, this kind of attack. That kind of attack. What we're trying to do is we're trying to exercise the security directly. Yeah.
21:11Robert HurlbutAnd that makes sense. I mean, I'm somebody who's taken more of a stand against the— and so people can send me hate mail or whatever, but I think as an industry, we focus too much on breaking things. I'm sorry, I'm just going to say it. I don't care what anybody— you can get mad at me.
21:25Aaron RinehartBreaking is about the making. Yeah, but it's about the making. In my opinion, that's where I stand on it.
21:31Robert HurlbutYeah, but I mean, you think about how much time, money goes into the breaker kind of side of the world. And so all that to say, I see the value prop in what you're describing here though, is because first of all, to your point, a lot of times pen testing only happens to the top 5% of the apps. What about the other 95%? Oh, no one's ever going to try to attack them. Like, who's making that risk management decision? So I like where you're going in that I feel like the narrative is leading me towards Security chaos engineering is something I can do for the entire fleet because I'm not paying somebody 40 hours of $500 an hour time to sit for a human being to use tools in their own brain to try to break it. I'm using automation, which I can apply to 100% of my fleet. Am I going down the right path in the narrative here?
22:20Aaron RinehartYeah. So let's get more to some of the how, right? So the first time you ever run an experiment, we call it like the exploration mode, right? The word is experiment. It should always be experiment because experimentation is key to all engineering. I really dislike that— I don't know who did it in chaos engineering, but they started calling chaos experiments attacks. I'm like, guys, I was the first security guy on the scene. It was like, guys, security people are going to hate you for this. It was non-security people doing it. It's not an attack, guys. It's not an attack, right? Like you're not attacking the system. I could show you an attack. It's a little bit different. And anyway, so the how, so the exploration. So the first experiment we write, so, you know, there are a bunch of different kinds of experiments, but you wanna start with something that you just know, hey, without a doubt, we bought Twistlock and we, without a doubt, We know that this is covered, or we think we're covered against a misconfigured port or a vulnerable image or a container talking egress out to the internet or something. That simple. In the best security chaos experiments, you guys are going to look at me like I'm crazy. are the low-hanging, easy ones to write from a Python or Bash perspective. I mean, like, that's not hard to write this kind of stuff, you know. It's like, but that's— but it's like, if you can't handle the basics, you're not going to handle anything more advanced than that, right? Like, like, people come to Casey and I'd be like, hey, you know, you guys, you guys should do some AI machine learning, it's blockchain. I'm like, no, no, actually, just the basic— what we're doing is next to no— whether you're in Silicon Valley and you're you know, uh, FAANG to, uh, you know, to, to anyone else. You can't, no matter how, whatever engineering prowess you have, you have these problems, right? Uh, and so, um, so anyway, so once you do it, so first time you run an experiment, so you write an experiment, let's say a vulnerable image, um, launching a vulnerable image or a misconfigured port. Well, you, you write it, you do it, you— but it's really best done in a game day exercise. It was— game days were invented by Jesse Robbins at Amazon when he was there. It's not like a sick red team day or red team game day. It's meant to be an engineering-wide kind of perspective, meaning you got somebody from the product team, you got somebody from the call center, right? You got somebody from the response team, somebody from the monitoring team. It may be somebody from the security monitoring team or somebody from the IT monitoring or technology monitoring, whatever that team is, or corporate monitoring, because sometimes they're different.
25:08Robert HurlbutRight.
25:09Aaron RinehartAnd it's what a great, what a great opportunity to, to introduce, you know, something all of you believe to be true and see whether or not it's true. You get such a breakdown of empathy on those game days, you know, but you don't have to do it this heavy human exercise. But there's a lot of value in one, getting people together that really work together, but they don't know, don't know they do, uh, and seeing how the system doesn't, uh, how everyone is wrong. how the system works. And so, you know what, you don't have to do a game day. You don't have to, right? It's just a good idea to make sure everybody's involved, understands the reasons why you're doing it, and that, you know, you're trying to enhance value on all fronts, right? Like, you know, and, but anyway, so, you know, if the experiment succeeds, which it probably will not, you're probably gonna be wrong the first time when you run it. I really have never seen a chaos experiment succeed the first time.
26:02Robert HurlbutYeah.
26:02Aaron RinehartIt's just we're almost always wrong about how—
26:05Robert HurlbutWhen you say it succeeds, you mean that the system detected an open port, an unauthorized port, that's success or is that failure?
26:13Aaron RinehartThat's success. So if it did what it was supposed to do, success, right?
26:19Robert HurlbutOkay.
26:19Aaron RinehartThat's success. But almost, I've never really seen one succeed the first time. I've seen them succeed, but like, it's that there's always some kind of recalibration. Oh yeah, crap. You know what? We did move that over a different VLAN, or we, I don't know, we changed the security group template. I don't know, whatever it was, you're reminded, oh crap, we did that. Because humans, we have a lot of things we have to remember, right? And we have a lot of things in our systems and it's just not easy.
26:47Robert HurlbutSo is there an active agent in this? Let's use the open port example, okay? So you create a Security Chaos experiment and you say, I'm gonna open a port on this system that should not have that port open. Is there some infrastructure that's doing scanning? Like, how do you know whether you passed or failed? Like, what systems are giving you the data that lets you determine whether you passed or failed the experiment?
27:16Aaron RinehartDepends on how you do it. I mean, it really depends on how you set up your experiments. So, the majority of chaos experiments, Really, you know, like the— so Verica is different. I can explain how Verica is different. So we've kind of evolved this craft a little bit. But let me start where most everyone else is. A lot of chaos experiments are heavily dependent upon the observability tools you have, meaning like your Spunks, your Hiptio, or not Hiptio, Humio, your whatever, whatever you got. And, you know, you have to kind of figure out, okay, did it succeed? Did it not? One of the best things you can do too is you can hook into— what you could do is you could try to hook into TwistLock's API and validate that it caught it. I mean, there's lots of ways you could potentially do that, but a lot of people put it on like— So you're on the right track. If you want to automate it, that's the path you have to go, is you have to have something to trigger success or fail, or you're alerting on You know, from your, you know, from some kind of mechanism you put in the, in your, in what you wrote to do the experiment. Like Chaos Slinger put all the, tracked all the output from, that's the tool, that's the tool I was gonna explain in a minute, is it tracked all the output and sent it to Slack, right? So it would tell us whether they succeeded or not. And we tracked it there. 'Cause, but, But if you want to do it and automate it, you want to hook into some kind of KPI with the control.
28:59Robert HurlbutSo are most people doing these manual right now? I'm trying to wrap my mind around this. And I know this is kind of a fledgling industry, so there's probably not a lot of people that are like, wow, they are super mature when it comes to security chaos engineering. And I'm trying to figure out, is this a You talked about game days. Is this a manual process where I'm doing a lot of this manually and then I'm validating, or is it— is there an automated kind of way that I can say, hey, I want to just drop, you know, a lot of different types of experiments randomly throughout the 30 days of the month, and I'm really going to test where we are because nobody's even going to know when they're going to happen?
29:44Aaron RinehartSo unfortunately, you know, from an open source perspective, there's not a lot, um, and I've been trying to get more people to open source What they've written. You know, I believe Oracle's got a tool for the Oracle Cloud that they wrote. We've got a great story about them in the next O'Reilly book, by the way. 10,000— they're doing this stuff on 10,000 endpoints, and it's pretty, it's pretty, pretty, pretty cool.
30:08Robert HurlbutWow, that's scale.
30:09Aaron RinehartYeah, it's some pure scale. You have— I don't know, I'm not sure how Cardinal Health is doing it. I think they have an automated— I don't know what they use for their KPIs. for the experiments. But Chaos Slinger, how we tracked it was via the Slack notifying us that it failed. And that would trigger us to go in and look at why and triage it. But I'm not here to pitch my company or the tool, but that's kind of what the evolution of what Barrica was. It's like Casey and I collectively created chaos engineering as you know it today. He started Netflix, I expanded it with the security bits, and what we wanted to do was make it valuable for people. I mean, it's not that chaos engineering is not valuable, it's just like people get enamored with the idea of chaos engineering, and they wanna do it, and everyone's saying, you gotta do it because it's about resilience engineering, it's about stability. Yes, it is, but it can take you all with the open-source tool that really is just breaking something and getting to the point of maturity that Netflix had. And so what we try to do is accelerate time to value in building that maturity into a tool. So what we— so Verica, for example, we don't just break something and you have to go look at Honeycomb or whatever to figure out what happened. We're outcome— our experiments are outcome-oriented. We have an observability component that helps inform you, like, what actually happened? Like, what did you learn? It gives you those insights. So it's built with observations, from the beginning. And you could take that kind of approach. And that's, but it's just, that's kind of how it, but even if you did it, even if you did it cron, right? There's a way to do it with write some Bash, write some Python, run a cron, even running the experiments once a week or every other day. I mean, it's not gonna be a whole lot of like, it's not gonna be a whole lot of work to really go through that at some point. Like, it's just, it's informing you, right, that something, that something's wrong. And, you know, I see a lot of people hooking in experiments into CI/CD pipelines as well, running these, because really big job scheduler, right? And like Jenkins, I mean, like, but, but yeah, I actually, it's important also to recognize. So it's one of the things that's important is to understand that a lot of people depend too heavily on the CI/CD pieces because a lot of things happen outside of the build process, meaning network configuration changes and environment changes. It's better to continuously run it on its own schedule and inform you than run from CI/CD. But if you want to hack running it at a time, you can put it in that environment.
33:05Robert HurlbutLet the record show I'm about to ask a product-related question that I asked and then Aaron's going to answer. Just because I know, I realize, Verica, you guys are at the forefront of this world, this new fledgling kind of world of tech. And so, from your product's perspective then, when you talked about observability, do you have the ability then to tap into AWS through some API connection where you're able to then observe any of the things, any of the actions that are happening, say I'm using AWS? Is that how you're doing observability, or is there something else to it?
33:42Aaron RinehartWell, it depends on— so, the way we approach products is that— well, the way we approach this entirely is— you really hear us even say chaos engineering when you engage with Verica. You really hear, like, you guys are the chaos people. Yeah, but I'm not trying to give you another thing to do. What I'm trying to do is help you with the problem you already have. Which is Kubernetes is complicated. AWS IAM is complicated. Kafka is hard, because Kafka's number one dependency is the network being reliable.
34:14Chris RomeoWhat?
34:15Aaron RinehartThe number one fallacy of distributed systems is the network's never gonna be reliable. So it's just, you know, so what we do is, so we're very narrow and very deep on how we approach experimentation. So that also allows us to look directly at those targeted observations. So, experiments are kind of narrow and deep in how they operate, but they're focused on— our experiments are not made-up failure modes, right? Meaning, they're not like kill pod, kill VM. Netflix had that failure mode in real life, right? So, that's why they did that. Now, you'll see, I'll see all these people, versions of Chaos Monkey for some other system doing the same thing. I mean, yes, but did you actually see that problem, right? So, one of the things Verk is also doing is we're We're trying to create the Internet Incident Library, where it's like we're trying to study how complex technologies and systems fail. That's why Verica was really slow to market. To get to the market is where you spend a lot of time trying to understand what are the hard things people are doing now. Kubernetes and Kafka were 2 of the biggest. AWS is another area, but Kubernetes and Kafka are both on AWS. Since we take a narrow approach, not only to product, Because we're trying to help you keep Kubernetes stable and secure. We're trying to help you keep Kafka the same. You get the idea. It's more of a deeper approach. But when you're opinionated from an engineering perspective like that, it allows you— it makes it easier to observe what happened, what didn't. That makes sense?
35:44Robert HurlbutThat makes sense.
35:45Aaron RinehartBecause we know what to look for exactly. Yeah, that makes total sense. Experiments are opinionated too. If we're assessing resource allocation in Kubernetes, which is a huge problem in Kubernetes, that's from a normal chaos perspective or stability, we focus on where to go get that information and present it in a visual way. We really try to present all experiments if we can in a visual observed way. Because one of the biggest problems— this is why observability is such a big word right now, is that, you know, when you have 3,000 microservices running, and, like, all the things and stuff inside of it, it gets really hard to grok information, right? So humans, we need simplification. We need interpretation. We need story. Tell me a story, right? Tell me a story about what's happening. And that's what we try to do. We try to tell you a story, right, about what happened, what we did, right? And how— what really happened, what we really learned, how the system's really working. And so, I guess if you could take anything from what I'm saying about product is that these are some tidbits on what you can consider when you design your own experiments.
36:59Robert HurlbutYeah, that's helpful. So, I want to stop for a second and just say from a lifecycle perspective, so I'm a big secure development lifecycle person, you know, how does it apply to Agile? How does it apply to DevOps? I'm always thinking about that. How is secure development lifecycle evolving? Where do you see security chaos engineering in a software development lifecycle and even in the secure development lifecycle? What's the right place that it fits in?
37:23Aaron RinehartSo, all chaos engineering in general, whether security or not, is post-deployment, right? I mean, if the problem is happening post-deployment, it's not that like— so we still do our smoke tests, our abuse tests, our fuzzing, we're scanning, all these things are still valuable, right? And Where CI forced us to really write smaller bits of code and ship them faster, CD helps us deploy more often. You can think about it as building a better engine.
37:58Robert HurlbutYeah.
37:58Aaron RinehartIs that we built a better engine, and then we're able to now go faster. But what chaos engineering does is helps— we call it continuous verification, It helps you keep and know you're still on the road. It helps you identify the safety margin. Well, the one thing in all engineering that you also cannot identify is where is the safety margin? Where is the edge of the system? Because you don't know until you fall off of it. It's too late when you fall off of it. That means instant outage. So it's post-deployment. Post-deployment is where the complexity is. You got 10 different teams. 10 different microservices, but all microservices are not independent. They're dependent on each other. So, there's functionality dependent on other functionality. Teams are updating things, and it's very hard to keep track of everything. Post-deployment, it's a mess. Yeah, you built it, all the tests passed, but the properties we're now expecting from the systems have to be emergent. So what I say emergent purposefully, because we're now really dealing with complex adaptive systems, which is a term in applied science. It's where chaos theory comes from. It means that the thing we're working with is no longer linear in cause and effect perspective. It's dynamically nonlinear. That means is like, since everything is acting to reacting and changing all the time, you get kind of cascading magnifying impacts. Like you get, um, so, um, a simple change, you know, can ripple throughout the system and cause harm.
39:39Robert HurlbutSo there's entropy too, right? There's entropy in a running system like that. Like things are falling apart, wheels are falling off things. The longer the system continues to run, um, things are going to degrade over a period of time.
39:56Aaron RinehartWell, so you bet you mentioned it. So I wrote an article, probably the best writing I've ever done, I've ever done, hands down. I put it on the Verica blog, but it's Security Differently. And one of the things I talk about there is the concept of brittleness. Brittleness in all complex systems can never be zero. You will never prevent all of the incidents. It's not gonna happen. It just will not happen, flat out.
40:18Robert HurlbutYeah.
40:20Aaron RinehartSo, all you can do is try to increase your ability to adapt and to adapt to change. And that's what— so, if you listen to Don Allspaw, David Woods, Richard Cook, that's sort of the— David Woods, creator of resilience engineering, that's what it's all about, continuous adaptability and empowering the human to be able to adapt and respond. We've all been in those situations where we're where we're like, something weird is happening with the system that we didn't expect, and we're having to go to different systems. We're having to respond in ways we haven't responded to before, but in the moment, we're doing things that we never thought we'd ever have to do. That's what I mean by capacity. It drives me nuts when people talk about guardrails in security. Guardrails is a no-no word in resilience. It's a no-no word because you design guardrails typically with what you believe to be the system in your head. I'm putting this, no SSH. I'm putting whatever guardrail protection because you try to do the right thing. You try to do the right thing. But what happens when you do that is that you think you understand how the system really works, but you really don't. So when a system no longer works that way, your hands are tied. There's an outage, there's an incident. We can't log into it. That's the way we built it. We gotta rebuild. You understand what I'm saying? These concepts are really valuable, Chris. I mean, I've been on really a big journey in diving into cognitive systems engineering, resilience, because I was trying to understand why is this something so simple producing so much value? I really started getting into resilience engineering, and like, you know, understanding how they got there. Because it was the Netflix people that kind of pushed, kind of like pointed me towards the resilience engineering stuff. It's not resilience. Cybersecurity people keep saying resilience. It's resilience engineering. That is a completely different— that's a completely different field, right? It's a field of its own. And anyway, soapbox off.
42:30Robert HurlbutSo what I want to do is we're— as we're kind of coming towards the conclusion of our conversation here, Let's say, imagine, I know this is going to happen with a number of people that listen to this. They're new to this whole idea of security chaos engineering. You've been doing this for a few years now. How would you recommend they get started? If I constrain you to say you can only give them 3 steps for how they can get started, because I know there's probably a million things they could do, but if you could only choose, like, okay, here's the 3 things for somebody that's brand new. They just heard this. They're like, ah, this sounds really cool. I want to learn more. I want to do some of this. What's the advice you give them?
43:08Aaron RinehartI would say number one, read the O'Reilly Report. The O'Reilly Report will give you greater context on why, how, and how to get started in like how Capital One built their program, how Cardinal Health built their program. These are not small companies. These are big complicated businesses that had the need for it and did it. And it was amazing because there is no book for them, you know, and Number 2, I would say we didn't talk about Chaos Slinger. Chaos Slinger was the open source tool we wrote at UnitedHealth Group. Chaos Slinger, it's not on the managed project at UnitedHealth Group anymore. It was the first open source tool for the company, but I left the company and then they use it internally, a different version of it internally. But the point is still out there because people still use it because it's more of a framework. There are 3 different There are 3 different functions you need to write. There's the failure injection piece, that is Slinger. There is the target acquisition piece, which is Generator. What it does is it goes out and identifies which security groups have, for Port Slinger, the main example, it identifies the security groups that have an AWS reference tag of X. Could be, I think we called it, Chaos Slinger was originally called Poo Slinger.
44:27Chris RomeoOh, okay.
44:27Aaron RinehartSo I think it was poo. That's what we used for the tag. Yeah, we're children. And then, yeah, but you can, you can, oh, and then there was a tracker which tracks the output from the experiment and sends it to a Slack channel. But we, you don't need this. So you could, it gives you exactly what you need. It's there to code. It's not very much code. It gives you exactly the framework by which you have to do it. And you wouldn't be the first people that that have done it. People are, people are, there've actually been a couple of copies of Chaoslinger that are out there now that people have been using. So that's 2. Let's see, what would 3 be? 3 would be, I'm trying to think of a good 3. Good 3. I don't know. Reach out to me. Reach out to me. Now I have my personal, I have my phone on the internet. I've had it there for over a decade. I have Twitter, DMs open, I have my LinkedIn messages open. I am a horrible security privacy person. I do that because, you know, I want to make a difference, right? I really want to make a difference. And I, no matter how many times I tell people, uh, you can DM me on Twitter, you can email me, you can call me, whatever you want to do. I probably won't answer if you call me because I don't know you, right? But like, if you message me that you're gonna call me beforehand, I'll answer it. Uh, but, uh, I have great filtering, so, uh, but I will, I will, I will respond. I will reach out. I will, you know, um, you know, be responsive. So, um, I want to help and I also want to know stories, Chris. That's, that's one thing Kelly and I are really, uh, wanting to get right now is that we're gonna collect as many stories as we can for the, the big O'Reilly book. That's what's being written right now, uh, over the next year, year and a half.
46:07Robert HurlbutOkay, great. So, um, one, one kind of last question. So What's your key takeaway for the audience here? Like, now I'm going to limit you to just one. Like, it could be a call to action, you want them to go do something. Maybe it's one of the things off your top 3 list there. Maybe it's just a key takeaway, but what do you want to leave this audience with?
46:28Aaron RinehartUm, I want to leave with— I would honestly say, um, I could say a lot by— I hate to say, go read a book, right? But, you know, we spent a lot of time writing that book. We rewrote it 7 times. So, you know, it's meant to be read beginning to end, but read the book. There's— it'll— a lot of you, I think you'll find a lot of it evergreen is from a content perspective and a different way of thinking. And I think that will help you frame up a discussion around, we probably need to look at this. And I will tell you, if I'm leaving with anything, you have no choice. You really, in the end, you will have no choice because of the amount of complexity, speed, and scale we're dealing with. Chaos engineering is going to be the new normal. It's already becoming that, but it's like people are trying to figure it out. It's still, like you said, fledgling. People matured, they're way up here. People in the middle, there's people getting started. We don't have enough tools to help us with that post-deployment complexity problem, and chaos engineering is just one of the only ones we have.
47:39Robert HurlbutAwesome. Well, Aaron, thank you for joining the podcast again, and for— I feel like now I've got a good understanding of security chaos engineering. I think I've made some connections here in the way you described it. So, thank you for sharing that with me, but also with our audience as well, and look forward to catching you at a conference sometime soon as the world continues to open back up. So thank you very much, sir.
48:02Aaron RinehartThank you, Chris.
48:04Robert HurlbutThanks for listening to the Application Security Podcast. You'll find the show on Twitter @AppSecPodcast and on the web at www.securityjourney.com/resources/podcast. You can also find Chris on Twitter @edgeroute and Robert @RobertHurlbut. Remember, with application security, there are many paths, but only one destination.
8,392 words · transcript by assemblyai
More like this
View all episodes →- August 6, 2021 · 32 minJeroen Willemsen -- Security automation with ci/cd
- July 9, 2024 · 1 hr 5 minTanya Janca -- Secure Guardrails
- February 16, 2022 · 42 minWill Ratner -- Centralized container scanning