Skip to content
AppSec PodcastThe Application Security Podcast — home
30 min

Drew Dennison – Security should make the computer sweat more

With Drew Dennison

Security TestingSoftware Supply Chain

Drew Dennison is the CTO & co-founder of r2c, a startup working to profoundly improve software security and reliability to safeguard human progress. Drew joins us to introduce a tool called semgrep.

Audio hosted by Buzzsprout. Nothing loads until you press play.

Episode chapters · 10 chapters
  1. 00:00Meet Drew Dennison: Drew Dennison – Security should make the computer sweat moreAudioVideo ↗
  2. 01:42All right, so the topic for today is something called SemgrepAudioVideo ↗
  3. 07:24I'm trying to think. You spoke about this tool, I thinkAudioVideo ↗
  4. 12:22This a replacement or a— I mean, is this designed toAudioVideo ↗
  5. 15:21Yeah. I mean, that's the whole DevOps approach to the worldAudioVideo ↗

About this episode

Drew Dennison is the CTO & co-founder of r2c, a startup working to profoundly improve software security and reliability to safeguard human progress. Drew joins us to introduce a tool called semgrep. Semgrep is a fast source code analysis tool, potentially faster than anything you’ve seen before. Drew Dennison is the CTO and co-founder of R2C, a startup working to profoundly improve software security and reliability to safeguard human progress. If you want to see the live demo Drew does of Semgrep, head over to the Application Security Podcast YouTube channel to see the video. We hope you enjoy this conversation with Drew Dennison. At Security Journey, we believe security is every developer’s job.

The Application Security Podcast is brought to you by Security Journey.

About Security Journey
Drew Dennison is the CTO and co-founder of R2C, a startup working to profoundly improve software security and reliability to safeguard human progress.
Learn more about Security Journey

Connect with Drew Dennison:
Semgrep
r2c / Semgrep Inc.

Resources
Semgrep
r2c / Semgrep Inc.
Applied Cryptography
Ron Rivest
John Nash’s Declassified NSA Letters
Hack (programming language)
Flow
nodejsscan
Yoann Padioleau
Applied Cryptography
Hacking Exposed Web Applications
Tree-sitter

Actionable

From this conversation

  1. Write targeted checks for risky framework usage

    I want to audit all my codebase of where people are doing authentication, but it's happening outside the trusted path.

    12:36
  2. Add fast scans to pre-commit and CI

    I think the use case that we've been using internally is either as a pre-commit hook, which has worked quite well, or a GitHub Action, or a CircleCI, where you're like, hey, as long as it finishes in less than 3 or 4 minutes, totally fine.

    14:38
  3. Show developers only findings introduced by their change

    It is going to show you the findings that are in that 15 lines of code or whatever you submitted as part of that pull request, which I think is nice because then you're like, oh, yeah, or I disagree with the tool and I want to archive it and be like, add that to the whitelist file.

    16:25
  4. Inventory authorization coverage across repositories

    Show me across all my codebases, what are the routes that have any authorization, or more interestingly, which ones don't have any authorization at all?

    18:23
  5. Automate analysis and reserve people for judgment

    Let the computer do that work, find the hotspots, and then, you tell the developer directly, you send a Slack notification or open a Jira issue to the security team, and then they're going to go through and apply their intelligence.

    28:10
Transcript · 30 min conversation

0:00Chris RomeoDrew Dennison is the CTO and co-founder of R2C, a startup working to profoundly improve software security and reliability to safeguard human progress. Drew joins us to introduce a tool called Semgrep. Semgrep is a fast source code analysis tool, potentially faster than anything you've ever seen before. If you want to see the live demo Drew does of Semgrep, head over to the Application Security Podcast YouTube channel to see the video. We hope you enjoy this conversation with Drew Dennison. At Security Journey, we believe security is every developer's job. We work with our customers to help them build long-term, sustainable security culture amongst all their developers. Our approach is to provide security education that's conversational, quick, hands-on, and fun. We don't do lectures. Instead, we let the experts talk about what's important. Modules are quick, 10 to 20 minutes in length. We believe in hands-on experiments, builder and breaker style. That allow your developers to put what they learned into action. And lastly, fun. Training doesn't have to be boring. We make it engaging and fun for the developers. Visit www.securityjourney.com to sign up for a free trial of the Security Dojo. Hey folks, welcome to this episode of the Application Security Podcast.

1:33Robert HurlbutThis is Chris Romeo, CEO of Security Journey. I'm also joined by Robert. Hey Robert.

1:38Chris RomeoHey Chris, yeah, it's Robert Hurlbut, threat modeling architect. Good to be here.

1:42Robert HurlbutAll right, so the topic for today is something called Semgrep, and we're going to figure out what that is because I think I know what it is, but I'm not 100% sure. But first we have to understand our guests, where he comes from. And so our guest is Drew Dennison. And Drew, we're gonna jump right into what is your security origin story? How did you get started in this crazy, crazy world of security?

2:09Drew DennisonHey Chris, thanks so much for having me on the show. Let's see, like probably like most of you, I was kind of the designated computer troubleshooter growing up in my family's home. We had dial-up internet and I don't know, I ended up going to our local library and getting a lot of Windows XP books and trying to just configure and learn a little bit about the Active Directory and all the IT side of things. And then at some point, I stumbled on Hacking Exposed, which is a classic book. So I read through that in probably freshman year of high school, I think. And then what really, I think, got my interest in security though was reading some of the cryptography books. I read Bruce Schneier's Applied Cryptography and There was a great book about the history of cryptography. And so kind of just going into college, I was really interested in, you know, basically math and crypto, with a little bit of security just from some of those earlier kind of IT, like configure your network, turn off all your remote desktop, you know, tunnels and everything like that. So then, you know, went to MIT for computer science, and Ron Rivest was one of the like choices for freshman advisor. And I was like, oh man, childhood dream come true. I have to pick Rod. And Rod was just like the coolest, nicest, smartest guy and really didn't want to talk about encryption. He was much more interested if I was getting enough food at the cafeteria. And I was just like, no, come on, we got to talk about security. So ended up majoring in computer science with a concentration in encryption and applied sort of more practical security. Stack smashing attacks and stuff. And one of my favorite classes was 6.857, which was the like introduction to cryptography class that Ron Rivest taught. And I remember, I want to say it was a spring semester class, and I think like literally 2 weeks before the class started, the NSA had declassified some old papers, like these kind of almost like serial killer looking notes handwritten from Don Nash. And he sent them in back in the '50s and was like, hey, I designed this unbreakable encryption machine. I think it was actually like a physical machine, almost like an Enigma. Would you guys be interested in this? And I think someone must have FOIA requested it and they dumped all that out. But anyways, our first assignment was to go through and break this unbreakable encryption. And it was obviously like there were flaws in it. So I think we did the statistical analysis. Ron had, or Vast had, made a Python implementation of it as best as he could follow John's hand-drawn notes. And we got to go through and There's actually some real, like, I was actually just looking up on the system right before this podcast, and there were some really interesting things that, you know, John Nash did have, which was like, you know, you should basically have the security of your encryption based on like a parameter about how hard it is to make things, you know, computationally difficult. And so you could go through and say like, I want a longer key, right? And of course, modern crypto, we think of that like you want a 256-bit key or a 512-bit, you know, or 1024, and you can kind of tune that computational difficulty. rather than just like scrambling, you know, letters, if you will. So that was just a fun, fun thing that really got me hooked. And then after that, I graduated and went to one of the Silicon Valley companies, Palantir, which was a great first job. I was doing consulting work there. Probably wouldn't join Palantir now, maybe, but anyways, I like was assigned to doing a lot of incident response work. So I got to go see large banks and retailers, you know, think like the largest retailers, like Home Depot, Target, like those kind of—

5:41Robert HurlbutYeah.

5:41Drew Dennisoncompanies or big banks in the US and Europe, and we'd kind of go either post-breach or towards the end of my time there, we were working on some cybersecurity insurance products where I was like, I had this hunch that basically the cleaner your network is or kind of the IT hygiene or health would tell you your probability of getting hacked. And so rather than having the standard compliance questionnaires, do you have a firewall, do you have antivirus, do you encrypt data at rest? We were like, well, we've got to be able to technically assess some of these things. Could we just get metrics? And then we were partnering with an insurance company. They were going to do statistical analysis if those hypotheses were correct. And so I worked there for about 4 years, mix of basically being a software engineer and the security incident response work, worked as part of a team doing that. And then my old college roommate who had gone off to work at the Department of Defense doing more offensive hacking, and I had been chatting and we just were like, there's got to be something. can do in security together, you know, do you want to start a company? And so he was considering getting his PhD, and ultimately we started RTC together. And so now I'm the founder and CTO of RTC, but I've always just wanted to kind of bring more like software engineering practices to automate more and more security.

6:58Robert HurlbutThat's very cool. I almost wish we could play a game of like, since you're someone who studied crypto extensively, where you say something and then we have to guess whether it's actually crypto-related or whether it's something you completely made up. off the top of your head. But we didn't have time to prep for that. Our producers were busy doing something else. But maybe in the future as well.

7:19Chris RomeoThe topic we want to cover today is this idea of this tool.

7:23Robert HurlbutI'm trying to think. You spoke about this tool, I think, recently at a conference. That's how I saw even that this thing existed. I looked up on the GitHub page, and I went, Semgrep? Because I certainly know what grep is. Wait, this is like a static application, kind of static scanning kind of solution that has some of the fundamentals of grep in it. And so, there's my terrible, probably, description of what this tool is and what it actually does. But from your perspective, what is Semgrep?

7:54Drew DennisonSure. So, Semgrep is an open-source tool that our company's been working on. The history of the project's actually about 10 years old. So, somebody we hired last year, Yoan Patilou, who's this French PhD, static analysis genius, and was Facebook's first static analysis hire, was working on a tool called Escript at the time at Facebook. And I think, reading between the lines, this must have been the early era. He was like employee number 100 or 200 or something along those lines. This was definitely the move fast, break things. There was no types. There was nothing like that. So Facebook was having lots of bugs and lots of errors. And so the first tool he built was this basically a way to let you just search for code patterns, because he probably started with grep, but was like, grep's not very precise. It's designed for strings. If I want to look for this specific function call— and so he told me at Facebook, they wrote about 1,000 different rules over time, and they would say, hey, this method's deprecated. Please use this other method over here. And that way, if a brand new hire or an intern started, they could—

8:57Robert HurlbutYeah.

8:58Drew Dennisonkind of get some basic program analysis running on every single commit at Facebook. And then I know Johan hired Julian, who's the creator of Hack, which obviously adds types to PHP. And so Facebook went definitely down the typed language support, and they built other products like Flow and stuff. But Johan worked at Facebook for a while and open sourced Escrep as part of a larger program analysis package. And then he kind of moved, retired, did a career change, something, went to Italy, and was working on a couple different books, and was just kind of thinking about what was next for him when we got in touch with him. And he's like, oh, that's pretty cool what you guys are doing. And so now he's working full-time with us out of Italy. And, you know, we've kind of expanded the original fscrypt core from just one language, or two languages, like for PHP and some C and C++. to what are the modern programming languages people want to use: Python, JavaScript, TypeScript, Go, Java, and we're adding more. It's pretty straightforward to add things, but fundamentally, what it's doing is it's letting you just express, hey, I want to look for this pattern in my code, and I'm happy to show a demo, but if you're going to reach for grep to search, is the word password anywhere hardcoded? Oftentimes, you can be a lot more precise with tool like Semgrep, where you can say, is password— is something passed to the password keyword argument, or something like that? Yeah, it's pretty easy to use. It's all open source. We've been using it ourselves, and what we've found anecdotally is you can write a new pattern about 10 to 50 times faster than you could go write a Flake8 plugin or an ESLint plugin or some other more traditional, actual static analysis framework. And so what it lets us do, like one pattern I was just working on last week was, I didn't actually realize this until I found a bug in my own code, but in Python code, if you just use exit as a keyword, it works most of the time because it's an interpreter built in. But really what you want to do if you're going to, like, ever, you know, do some performance optimizations or something is use sys.exit, right? It's just a really simple bug. It's not even a security bug. But I was like, I can just go in 5 minutes, wrote the little pattern, and now it's checked into our open source repository of these rules, and it's just like anyone now can be like, oh, that's a small bug. And so, I think a really interesting application of that is application security and checking for misconfiguration or insecure frameworks, things like that.

11:30Robert HurlbutSo there's a list— so there's a rule base then, or pattern base. That's part of the project. So if I'm somebody new that's coming to this, I can take advantage of all the existing rules that you've written for Python or Java or one of the other supported languages. I don't have to figure out all those patterns myself on day one, right?

11:50Drew DennisonCorrect. Yeah, there's about 200 rules we have out there right now that we've sort of vetted and written, and we're getting other rules from a couple different people, one like from the Node.js scan tool. He contributed about 50 rules. And so, we're just, over time, building up this repository that's just open source for ways to be able to say, I'm an expert in, let's say, the way you do database calls in Django. Here's 5 rules that just catch common bugs that I've seen. And so, that's been really nice to just share these patterns.

12:22Robert HurlbutSo, is this a replacement or a— I mean, is this designed to be a tool that causes me to say I don't need traditional commercial static application security testing tools?

12:36Drew DennisonIt certainly could trend that way. I like to think of it somewhere in the middle between grep and, like, a full-blown commercial tool, or even an open-source SAST tool. I think the areas that we probably don't have right now is anything that involves state tracking, or if you want to do really deep taint analysis, a tool like Coverity, for example, is going to have 10 years of really great PhD research that's been done in that. But one of our hypotheses is that it's sort of a trade-off. If you want to go do whole program analysis, it's often very slow, and you end up looking for, I would argue, language-type bugs of, oh, this Java variable is null here, and you're going to dereference it, and that's going to be a bug. And that's cool, but it's often prone to false positives. And I think where Semgrep really shines is when you have Like, a framework, like, okay, we're going to use this library to do all of our encryption. I want to write something that says nobody should be importing SHA-1 or MD5 or any, like, Bouncy Castle or something like that, because I know that you should just use the NaCl Strongbox encrypt function. And so, I want to audit all my codebase of where people are doing authentication, but it's happening outside the trusted path.

13:48Robert HurlbutYeah.

13:50Drew DennisonSo we'll see. I mean, I think, you know, as, as we get more features, like we're working on some basic paint stuff, and, you know, I could see it kind of moving in that, that space. But right now, I think it's probably more of like the 80/20 rule. Get 80% of the value of a full SaaS product with hopefully less than 20% of the effort.

14:05Robert HurlbutYeah, and I got to imagine it's pretty quick from a performance perspective.

14:10Drew DennisonWe are targeting 100,000 slok per second, so that's about half the speed of grep. It's certainly never going to be as fast as, say, like ripgrep or any of those tools, because you're not able to just do ASCII or Unix character matching and regex engines. You are actually parsing things into a full abstract tree, but all the core work is done in OCaml, so it gets compiled to native code.

14:33Robert HurlbutSo, you said it was 100,000 lines of code per second? That was the metric?

14:38Drew DennisonThat's our target for just a single pattern. I want to look for someone imported this function or something like that. It should be fast, and I think the use case that we've been using internally is either as a pre-commit hook, which has worked quite well, or a GitHub Action, or a CircleCI, where you're like, hey, as long as it finishes in less than 3 or 4 minutes, totally fine. Happy to block the build, because that's going to be faster than my test suite.

15:02Robert HurlbutYeah.

15:03Drew DennisonAnd I think that is an order of magnitude better than— I've never personally used Fortify, or Checkmarx, or some of these more heavyweight things, but anecdotally, I've heard those will typically run for hours, and then you'll get a report, and you have to go through and kind of triage that, and it's not in line with what the developer is thinking about that bit of code.

15:20Robert HurlbutYeah. I mean, that's the whole DevOps approach to the world, though, right? What you're doing with Semgrep is the way things have to be to live in a pipeline. You can't tell me that, you know, I've got to run my SAST tool, and I need 2 days. Like, oh, we just pushed 300 versions of code in that 2 days. And we created a backlog of scan time of 12 years, all in that 2-day period, right?

15:49Chris RomeoRight.

15:49Robert HurlbutAnd so, you've really got— you know, you guys are— your team is— and I'm sure there's a lot more people behind this as well, but the concept of doing this at a high rate of speed is certainly the way to go. So, you mentioned pipelines as one use case. Like, you can include this in your CircleCI build pipeline, GitHub Actions. I'd have to configure my own GitHub Action. It's not something that's part of GitHub.

16:14Drew DennisonWe have an action on the GitHub Marketplace that is actually pretty nice. You can just go to install it, and you paste in 3 lines of YAML into the GitHub Action workflow, if you've ever done something like that.

16:25Robert HurlbutGot it.

16:25Drew DennisonIt's really nice, and we've improved it a lot in the last couple of weeks, so a new version is coming out soon. But what it'll actually do is it will run the tool twice. First one, before it applies the patch, and then applies the patch, and then runs it again, and then diffs the results. The nice thing is, rather than showing, thousands of findings, which you're like, I'm not even a developer. I don't even want to think about some other bug in a different part of the codebase. It just is going to show you the findings that are in that 15 lines of code or whatever you submitted as part of that pull request, which I think is really nice because then you're like, oh, yeah, or maybe I disagree with the tool and I want to archive it and just be like, add that to the whitelist file. Great. But if it's something I've done, I'm like, oh, actually, you're right. I forgot to close that file descriptor, or I can have a nice message that's like, hey, please don't use a float field for money. You should use a decimal field because it doesn't have floating point errors, stuff like that.

17:21Robert HurlbutYeah, no, I think that's a neat way to approach it. I think that's a— one of the things I've already said this like 3 times in the last couple of podcasts, but my 2020 2 words I'm putting together are developer empathy. As security people, That's what we struggle with, is putting ourselves into the mind of the developer. What you're doing there is you're creating a GitHub Action that's smart because it's a developer-first function in that, of course, the developer doesn't care about the 10,000 other lines of code. They care about what they just committed, the changes to what happened there. That's brilliant to give them just what happened because they're not going to be able to say, that doesn't apply to me. Well, yeah, it does because it was your code that changed. It was your change that made this thing pop and go on.

18:07Drew DennisonYeah.

18:09Robert Hurlbutreally neat way. So, GitHub Action pipelines— what other use cases are people running this? Are developers running this in their local build environment, their local development machine, to check things pre-check-in?

18:23Drew DennisonI run it as a pre-commit hook, and so I just use the pre-commit framework, and I have a bunch of other things, including Black and Prettier and MyPy and all your style and formatters and things like that, and then I just added this as a pre-commit hook as well. So, that's worked pretty well. I think the other major use cases we've gotten, one is security teams looking to do asset inventory of their code. So, they want to go through and say, show me, I've got 15 or 20 different repositories. I'm going to write 3 lines of a Semgrep check that looks for someone who's importing an AuthZ authorization decorator that says admin only. Show me across all my codebases, what are the routes that have any sort of authorization, or maybe more interestingly, which ones don't have any authorization at all? And then, they can go through that and do that auditing and at a code inventory level. Another use case we've gotten some interest from has been the security research community. These have been either professors, grad students, or often bug bounty hunter types. One thing that we've done— this part's not open source, but we've built out— and this probably goes more towards our commercial offering— we've built out a giant MapReduce framework platform using AWS, and we can scale up to 5,000x parallelization. What that lets us do is say, I want to run this pattern across a million JavaScript repos, or the top 10,000 Python packages, and just, it takes about, I don't know, 2 minutes or something for us to go through and scan that volume of code, because we go in the background and spin up thousands of these cheap Spot instance executors, and it costs us, like, I don't know, $10 or $20 to go do a scan like that. Okay. And what that lets us do is 2 things. One is we think about writing more and more rules into this open source rules repository. We can tune them on real code and be like, oh, yeah, actually, here's this edge case. we had to go fix. But it also lets people say, I'm gonna write this vulnerability thing, and almost do a variant analysis across the entire ecosystem of open source projects, and then go submit bugs. And so, we're just starting to talk with a couple of security researchers. One of the guys is, like, I think, number 5 on HackerOne for JavaScript vulnerabilities, and he's been pretty enthusiastic. And I think we're at the stage right now with Semgrep where we've gotten it to, I'd say, like, an alpha product, and we're just kind of looking for the teams and the people who are like, this is really cool, I wanna shape it exactly to my use case. use case. And so I'm really focused on reducing friction and just making sure it stays performant and is easy to use. So are you—

20:50Chris Romeoit sounds like you're mostly code agnostic then.

20:53Drew DennisonYou're not— because you're running outside of any IDE, out of any environment that has a dependency on Java, it's .NET, it's Python, it's whatever.

21:02Robert HurlbutYeah, you have patterns that still match certain code constructs as well, right?

21:07Drew DennisonRight. So we do care about the language because we need to understand the syntax of the language.

21:11Chris RomeoSure.

21:11Drew DennisonSo there is a universal abstract syntax tree for those who like to geek out about programming languages. And we're starting to move in the direction of using Tree-sitter, which is what GitHub uses for their Atom code editor, and that lets us support every language that Atom supports. But you're right, the fact that you don't need to have compiled artifacts, I don't need to understand how the Java bytecode is built, if you ever use a tool like SpotBugs or something like that, because you're just looking at the Java source. And it doesn't even technically have to be valid or buildable project, we'd just need one file or something, because you're just going to look for, from java.spring.webhandler, import servlet, or whatever the— I forget the exact Java pattern that you'd want to look for. And then, from there, you'd say, okay, cool. We're taking a request. We're getting the user data, and then, we're going to send that directly to exec, or something, or a file opener, or a database query. And so, is that safe?

22:08Robert HurlbutYeah.

22:08Drew DennisonWell, who knows? So, you can kind of look for these little simple code audits, right, when you think about just at a file level.

22:15Robert HurlbutThat's very cool, that example you had earlier there about the scale with which you're going to be able to scan things, like being able to scan 100,000 JavaScript repos at the same time. I just feel like there should be diabolical music playing in the background while you describe that. It sounds like a Dr. Evil kind of move, you know? You should be having your hands up like this as you're telling me. No, that's very cool. That's got to be like the most, you know, it's got to be a problem or a solution that will really resonate with people who are scanning lots and lots of source code, because I know that is the biggest challenge people have with static scanning is the timeframes that go into it.

22:58Drew DennisonYep. And that's why we were like, most of this stuff should be, especially if you kind of don't have to depend on building the entire project, it should all be very embarrassingly parallelizable. I can get a spot instance for an hour. It has 1 CPU, 2 gigs of RAM for a third of a penny. So I can do a lot of code checking in an hour for very, very cheap amounts of money. So yeah, having that scale is important. And I think oftentimes, the trade-off that I was unhappy seeing a lot of my AppSec friends and myself is you either say, I care about third-party dependencies and I'm going to go buy a really cool tool, let's say Snyk, or there's Dependabot or SourceClear. There's a bunch of third-party vulnerability databases, but they're typically just looking at, does this version match the CVE database? Or, I'm going to say, well, what's the code I can control, and I'm going to write? It's only the companies like Google, and maybe Dropbox, and a few others I've talked to, who actually go and vendor all the third-party code and bring it into their monorepo, and actually care about the security of their third-party libraries that get built into the final application. And I think everybody else just kind of throws their hand up and says, well, if there's some crazy JavaScript package that has hardcoded initialization vector, someone else is going to have to fix it and find it. Yeah.

24:10Robert HurlbutYeah, I think that's a common approach to third party that 99.99% of the industry follows, which doesn't mean it's right, but I think it's reality right now.

24:21Drew DennisonIt's rational because, you know, you're probably not going to get hacked unless you're like a cryptocurrency company. Someone's probably not going to take the time to like hack a third party module just to get inside your environment. I mean, I guess we saw that happen last year with JavaScript, but yeah, you know, if you're just me building a startup, you know, probably I'm safe just depending on Flask and React and whatever other third-party libraries.

24:43Robert HurlbutBut you're running Semgrep against anything you're building, so good luck. They're not gonna be able to sneak any of that stuff by you. So what does the future hold then for Semgrep? Like, is there a roadmap out in years in the future? Big, big plans for we want to add this big cool thing, or what are you thinking about the future?

25:00Drew DennisonWe have 4 key features we're adding. I'd say, I don't know if it's 3 years out, but in, say, the next 2 or 3 quarters, we're working on more code equivalences. So this is letting you define, almost like in a plugin architecture, there might be 4 or 5 different ways to do semantically the same thing. So what I mean by that is, let's say I'm opening a file in Python. There's the file open, there's the pathlib.open, there's with, file open, you know, I can rename things. So, like, as an application security engineer, all I want to say is, write one pattern that finds all file opens, and then I'd like to go kind of into this code equivalence world and say, oh, but this also matches, so I'm going to map it back into this, you know, kind of abstraction layer. That's something I'm personally really excited about. The next thing we're working on is, I was mentioning earlier, like, support for, you know, more languages. Let's go from, like, 5 languages we support to, like, 30 languages. That'll probably land in the next quarter or 2. And then the second or third thing we're working on is integrating external type information, because this is another way you can just be so much more precise. If you take— I don't know how familiar you are with JavaScript vulnerabilities, but for example, the setTimeout method, which is used everywhere, everyone is trying to use this polling architecture, you're going to hit the server every 400 milliseconds or something like that. Normally, it's given a function, but if you give it a string, because JavaScript's JavaScript, it will just eval it. for you. And so, what I want to be able to do is write a pattern that says, find all calls to setTimeout where the type is a string. And, obviously, type inference is a huge, hard problem. I don't want to solve that, but I do want to let the TypeScript engine, or the Flow engine, or some sort of external typer run over and say, here's what I think the type is, dump it out as a JSON file, and then I can now refine my patterns. And so, for languages like JavaScript, that's easy. Python, of course, has a couple of external typers, you know, mypy now. Sorbet exists for Ruby. And so I think that would just be a nice step up to let us write more precise things. And then fundamentally, once you have types, then it's relatively straightforward to use more of a tainting engine, because you're just defining a special type, which is user data. And so the tricky thing here though, and I think it's important to follow what, say, the TypeScript folks did, where you end up wanting this community registry to have, here are all the places that users can manipulate data. So if you take, let's say, a web framework, Here's a way you can change the headers. Here's a way you can change the user agent. Any of these things are dirty data that you can't fundamentally trust. And then you're going to have some sanitizers that clean that up. And then you're going to have some dangerous functions, which might be exec. It might be reading right to a database without parameterized queries. And so if you can now have a list of, these are the ways people modify things, here are the ways that these dangerous functions, then you can glue those together and do what's classically known as taint detection. taint analysis. So those are the 4 features that, you know, we're coming out with: code equivalences, more language support, having external type information, and taint analysis.

27:58Robert HurlbutSo if you had to leave our audience with one kind of final thought or a key takeaway from our conversation here, well, what would that be, or what is it?

28:10Drew DennisonWell, my thought is, just kind of zooming out a level from Semgrep, is We should make the computer sweat more in doing our security. And what I mean by that is I think typically people will go through and do like a pen test, or they'll look through and try to do manual inspection of code. I've certainly had it done on me, and that's great because you get that human intelligence. But what does the world look like where, as the computer's idle, as the developer's sitting there thinking, how do we run more and more analysis behind the background saying, oh, this is going to be a bug, this is a correctness issue? a security problem, or when you go submit it, let's just go run that scan really fast on every single commit. Let the computer do that work, find the hotspots, and then, you know, maybe you tell the developer directly, maybe you just send a Slack notification or open a Jira issue to the security team, and then they're going to go through and apply their intelligence. And so, how do we augment and just kind of move AppSec from a human process to, you know, fully automated? I think that's just the next leap forward.

29:07Robert HurlbutYeah, that's a great line. Make the computer sweat more. I love that as a line. So Drew, thank you for taking the time to be with us today to explain Semgrep, and we look forward to following along as you do that development and talking with you some more in the future to hear how Semgrep progresses and all the good stuff that's happening there. So thank you very much.

29:32Drew DennisonYeah, you can go to semgrep.dev. It's just a short redirect to our GitHub repo, start the project. There's a community Slack room, you know, follow us on Twitter. Very excited to be building this and doing it in the open and getting community feedback.

29:45Chris RomeoThanks for listening to the Application Security Podcast. You'll find the show on Twitter @AppSecPodcast or on the web at www.securityjourney.com/application-security-podcast. You can also find Chris on Twitter @edgeroute and Robert @roberthurlbut. Remember, security is a journey, not a destination.

5,794 words · transcript by assemblyai

More on Software Supply Chain

View all episodes →

Get Reasonable AppSec: new episodes and useful picks from the archive.