Transcript
Real stories, real defenses, and real recoveries, straight from the practitioners building and defending modern data environments. Hi and hello to the latest episode of Zero Downtime. My name is Shelley Calhoun-Jones, and I'm a Technical Marketing Director here at Cohesity. Today we're joined by Don Easton. Don, would you like to give an introduction? Yes. Hi, my name is Don Easton. I am currently a Cybersecurity Professor here in Eugene, Oregon. I have 25 years of experience in the industry, a majority of that in cybersecurity, cloud security, network security, and so on and so forth. And I've been working in the field, and data storage and compliance is one of my specialties. Don, it's great to have you. And that's really going to be the topic of our conversation today, because we know that data can stick around longer than we realize. Even after it's created, it just keeps on getting copied, moved, backed up, archived and recovered, sometimes ending up in places that we normally don't keep an eye on. And this is where security risk can hide. And so in this episode, we're going to take a look at what it really takes to secure data through its entire lifecycle, and how backup and recovery systems have become part of the attack surface, and why protecting data isn't just about prevention. It's making sure that you can recover clean, trusted data in case something goes wrong. Don, when organizations believe data has been deleted, where does it usually exist? Well, it's rarely gone in one click, as you just mentioned, it persists in secondary storage like automated backups, site archives, disaster recovery mirrors, and in something that we often don't think of as shadow IT. Employees when they don't have the resources or claim they don't have the resources, you know, they'll store data in personal Google Drives and OneDrives, you know, so they have downloaded copies for ease of access, and they've done it years ago, they may be residing on their local workstation or in their Google Drive account for 10 years. Beyond that, think about the incremental backup process many organizations follow, that's, you know, the most common and the best, in my opinion, way to perform data backups. Even when you delete a record from the production database, that record exists in yesterday's full backup, the day before is incremental, it may be in snapshots used by developers for testing. To truly delete something, you often have to age it out, all right, so you have to have some type of process to make sure that you meet your data requirements for retention. But that cloud bucket or that backup tape, you know, sometimes we have to look in places we didn't really think, and I kind of think of it as a set of footprints. So you're not just deleting a single piece of data, you're trying to erase a trail, it's been used by a dozen different processes, automated or otherwise. And in the age of AI, really, I mean, that's definitely something that we have to pay attention to. Now, those are all really good points. And I keep on thinking back to customers I've worked with, you know, previous roles, where they may need to keep their data for up to 30 years. And they may have, they may have like dozens of cloud accounts that they're managing, all with different data lifecycle policies. So it really kind of makes me think of, you know, herding cats in a way, having to keep all of those, all those cats herded. Yes. So, you know, from a security perspective, you know, how does long live data quietly change an organization's risk profile over time? So what's interesting about this question is there's a term that I have used in the past and I know it's been around a little bit and it's called data gravity. And so the more data that an organization accumulates, the more it pulls other risks towards it. So other applications and services and users lean towards it. In cybersecurity, we know that that's a lot of the, you know, the metrics that we use and we gather and all that information we use for threat intelligence, you know, all of that's long live data. You often end up having way more data than you have realized. So, you know, long live data will often lose some of its metadata of some of its purpose, the context of who owns it, why is it sensitive? Why do we even have it? So this long live data is a prime target when, when somebody breaches a system, they're not just looking for today's or yesterday's spreadsheets, they look for the 2018 folder that hasn't been touched in years, you know, because they know no one is monitoring the access logs for that. So think of it in the terms of like an addict, I like to do this when I use an analogy for this when I'm working with students is you got boxes of stuff upstairs you've forgotten about, but a burglar is going to find plenty of value in those old records while you're busy guarding the front door. So basically, you know, you have to watch everything and you'll have all this data. Even the old data is important, right? And so that really, the security risk that, that profile that we're looking at, it's, it's definitely increased that threat landscape is just much wider when we look at the old data. You bring up a really good point there, Don, about threats that could be dormant in the environment for months, you know, if not a year or more. So yeah, making sure that you have all of your data inventoried, making sure that you're cycling out, you know, old credentials, because for a lot of developers, they may have hard coded API keys may be built into their applications versus using roles. So yeah, those are all really good points. And I love, I love the idea of the kind of like the garage or attic analogy where you just have all this stuff, you know, maybe projects that you've worked on and that you've never finished in the case of my, my garage here at home. So that's great. So how often does data outlive security assumptions that was originally created under? So when I think of old data, I often think about cryptographic age, that's the term I always use, or, and some people refer to it as decay. So it's the gradual like weakening of security protocols because an encryption algorithms over time. So say for instance, we encrypted something with SHA-1 years ago, SHA-1 is considered insecure. So we have data from 10, 15 years ago with SHA-1 on it, it's encrypted with outdated standards. Cryptacker doesn't need a key, they just need modern computing power to brute force it. We know that cryptograph, cryptography is only as strong as the amount of time it takes to break that security algorithm or that, that to get that key. And if it doesn't take but a couple hours with a modern computing system, then it's not secure. So it's secure by 2015 standards, 2018 standards, but it's definitely instead of what I would consider vulnerable by 2026. That's interesting. Yeah, it really kind of drives home the point that if you have vulnerabilities, even if it's patched, there's the potential that it, that, that patch, if you don't have, have your environment, you know, updated, there's another way that the attacker could get in. So I love the fact that you bring up cryptographic standards. That's a constant race here in the security industry for as long as I've been in it, and it's just continues to be a battle, you know, making sure that your environment is secure and that you're using the proper encryption keys. So why does restoring older data sometimes reintroduce risk that a team may have thought was already resolved? And I feel like that kind of ties to what we were previously talking about with vulnerabilities or using older cryptographic keys. So I usually refer, when I, when I think of a question like this, I look back at my endpoint days. You remember those days with endpoint protection and managing large scale endpoint deployment. We get this thing where we do backups. And I call it, it's kind of like a time machine thing where when you restore a backup from six months ago or a year ago, and you're recovering from some kind of an incident or an outage, you might inadvertently restore dormant malware, strain, unpatched vulnerabilities, compromised credentials that your team worked hard, you know, they're trying to get them out of there. You know, all this work that you've had to do to make the system secure and you're restoring data back in, which could potentially be insecure. So I like to think back to the days of Windows XP when we had the restore system, restore feature. And oftentimes, you know, the user would restore that system back six months ago, but not knowing that that malware sat on that system for six months. And so they're just back to where they started from, right? More recently, I have dealt with a few of these in the past myself is ransomware, ransomware recovery. If that person, that attacker hasn't, you know, that dwell time that we all know is really, really a key factor in an effective attack, they can be in your system for three months before they trigger it, right? You know, your backups from the last 90 days are likely, they're already compromised. So you know, and we don't know, unless we keep really good records, we don't know how, you know, what five years ago, what happened, what did the security team have to deal with five years ago, and is there a chance that these backups or this old data couldn't hold some kind of a vulnerability that we're reintroducing into our network, right? Yeah, that's a really good point there, Don, related to the infected snapshots. How did you typically handle that when you were working in like a security operations role being able to pin down, you know, a clean snapshot versus an infected snapshot? We would usually, I mean, so in a lot of, we would have test systems because we did a lot of patch management and did a lot of security operations. And one of the things that we would end up doing is that we would test our backups on a regular basis, which of course is best practice. And we kept good records, we would make sure that we understood five years ago, maybe we did have a crypto wall infection, right? But it's really, it can be tricky. And so I always tried to play it safe and would restore into a test system and, you know, do a full scan. I would hope that current endpoint protection could detect an older compromised, you know, system or compromised files. When we look at encryption and all that other stuff, I mean, really it's, we have to assume that if something was encrypted in a certain format in 2015, then we know that it's not current now. Yeah. Right. And so that's something I'm going to take some extra work. Those are all really good points. And something for the audience, just so you're aware, Cohesity does offer the ability to scan your backup snapshots. So we will scan your backup snapshots. And if we do detect a threat or some type of anomaly on the snapshot, we'll actually notify you. So you can find that feature within the Cohesity Security Center. We also have the ability to do data classification scanning. So if you have a threat that's moving laterally across your network, potentially using a server as a staging site, we can scan that object for any sensitive data and you can actually plug it into your other security products. I love the idea of also doing those dry runs and rehearsals. That's something that I always tell our customers, you know, it's more than just being able to back up your data, but also making sure that you have the ability to recover, not just from a backup or recovery standpoint, but also from a DR standpoint, just to make sure that you're engaging the right stakeholders, that you have the right processes documented, all that good stuff. Because you never know when a threat's going to hit your environment. So just some last, you know, key takeaways here, you know, we've talked a lot about, you know, the life cycle of data and how data can live within your environment for years, if not decades. We've also talked a lot about some of the different ways that you can make sure that you have proper data lifecycle policies set up within your environment. Don, any last parting thoughts that you'd like to share with the audience? I like to think of a good way to begin to approach this is, I think, being very minimal about your digital footprint. In this day and age, when storage is super cheap, and we all have supercomputers in our pockets, and we have thousands and thousands of pictures on our phones that we never get rid of, because it just takes too much time, that it's good to consider creating policies to decide why are we keeping this data? What is it for? Do we really need it? I like to use a lot of analogies, and I think about this where I've had in the past boxes of cables in my garage, and I'm like, all these old cables, right? Do I really need a parallel printer cable? I mean, RS-232 serial cable? What am I going to do with this? A null modem cable? Do I really need these, right? And so you have to really avoid trying to keep that data past its prime. If it's not current, it's not relevant. It has what we call a carrying cost, you know, it costs money to store it. It costs money in legal discovery, it costs money in breaches. If we get a breach and we have old data that gets out into the wild, you know, it's on us, it's on the organization. So yeah, if you don't need it, let it go, right? It's just, it's not worth keeping it. Yeah, no, it really kind of makes me think of all the cables that I have at my back office. It's like, do I really need all of these cables? Well, I want to thank you for joining us today, Dunn, and just for the audience, if you want to learn more about how to strengthen your data protection strategy and see how Cohesity can help organizations recover with confidence, visit Cohesity.com for our latest insights, demos, and best practices. And if you found this discussion useful, please follow us for more Zero Downtime episodes where we'll explore the trends shaping data security in 2026 and beyond. This wraps up another episode of Zero Downtime. Thanks for watching, everyone.