Transcript
everything IT, all in a virtual environment. I'm Darren Thompson, your host and field CTO at Commvault. And I'm joined today by Stephen Foskett, founder of Tech Field Day, part of the Futurum Group. Now today, we're going to dive into a concept that's reshaping enterprises as we think about continuity and recovery, the subject of res ops. We're also going to spend some time discussing meantime to clean recovery and how new metrics like this are changing what good looks like in the modern resilience framework. Stephen, welcome to Strive. Thanks a lot. It's great to be here in New York for Commvault Shift. Great to have you. You've been tracking IT operations in its various evolutions for a lot of years now. Tell us a little bit about that experience, first of all, and I'd love to get into your initial observations of the differences really between traditional IT ops and now what we're calling resilience ops. Well, I think fundamentally, it's right to call it IT ops because traditionally, it has always been about IT operations. It's about operating IT infrastructure, operating applications and platforms for IT, and not so much about operating business functions and supporting the business. Early in my career, I was part of the IT operations staff, and it was true that in many cases, I didn't even know what applications were running on the infrastructure that I was supposed to be supporting, let alone actually building infrastructure that was capable of supporting their needs. Right. Interesting. We're seeing more and more a focus on technology as well as people and process and everything that surrounds that. In some ways, taking the IT away, I think, is probably a good thing. It gives people a slightly different focus. Talk to us about what you're seeing now in terms of this initial change and give us some comparisons from what you've seen, again, traditionally and the definition of resilience ops. Well, traditionally, data protection was handled by IT, and again, the challenge for IT was that, frankly, they didn't understand anything about what that data was, what its needs were. What we would do is we would essentially put together standard metrics like RTO and RPO. We would attempt to meet those requirements, and we would hope for the best. What I think is happening now, especially with the rise of DevOps, we're seeing companies looking at ways of more tightly integrating IT infrastructure and applications with software development and, ultimately, with the needs of the business. Right. Right. That makes sense. Talking of the business, how do you see this organized? I often get the question, well, if we're moving in this direction, we're bringing sec ops, infra ops, DevOps together, what does that do to the organization and the organizational structures? Are we going to see new roles develop, do you think, and emerge? Yeah. I think one of the challenges on that point is that, in many companies, security and cyber resilience are handled by a cyber person at the C-level, whereas IT infrastructure is handled typically by a different person, and in many cases, software development is handled by a third person. All of them, I think, though, are looking at this DevOps integrated approach toward developing applications, and I think that they're hoping that they can more tightly organize, more tightly coordinate, but that's always going to be a challenge because, ultimately, you have different reporting structures, different goals, and, frankly, different incentives for each of these different parts of what many would think of as IT. Yeah. It's interesting. Something we've been doing recently at Commvault is deliberately forcing those disparate teams into a room to workshop something out, to work on some problem, and the problem is often less important, actually, than bringing the people together. Are there other sort of tips that you have for people in terms of how we can start breaking down these silos and have these various disciplines start to work together? Well, ultimately, I think it would be extremely helpful if we could get those people together to work together, and that's one of the things that I've noticed that the industry overall is really trying to do, especially a data protection company with that sort of heritage like Commvault has, bringing in people who really understand cyber, who really understand resilience, and trying to build sort of a new practice of res ops, resilience operations, within the cybersecurity practice as a way to help them better work with the traditional IT folks over on the more traditional IT ops side. Right. Yeah. A really simple use case for that that we're involved in every day here at Commvault is building a cyber recovery plan. The cyber recovery plan needs those traditional infrastructure folk. That's really backup and recovery and data protection, but it also needs forensics so we know where the clean data is. That's traditionally a security job. It needs DevOps in terms of some of the sort of functions that you were mentioning there. So again, a good use case for building a workshop is what does a cyber recovery plan look like? How are we all going to contribute to that? What is that team going to look like? Well, it would be great if more companies had a cyber recovery plan because frankly, I think many of them do, but it's one of those things where I guess a secret of the industry is that most people don't really believe it's going to work. I think that they hope it would work. I think that they hope that it would meet the needs of the business, but in many cases, they really don't have that kind of confidence in it. But having that kind of workshop where you're sitting people down across those artificial divides between the different parts of IT could lead us to a position where we might actually have a cyber resilience plan that might actually have some hope of actually bringing the data back in the event of a breach. Yeah, interesting. One of the other things we've been working on is how do you measure that? So traditionally, that infraops team, they were talking RTO, RPO. Security teams are probably more interested in root cause analysis. Almost all of that and none of it matters in a cyber recovery plan. What actually matters is how quickly can the business bounce back? When do we get clean data that we can rely on back for our minimal viable company as quickly as possible? I was recently personally involved in writing a paper around this concept of mean time to clean recovery. I don't know if that's going to end up being the metric, but it's at least moving in the right direction and thinking about this. What are your thoughts around metrics? Have you come across any others that could be useful in this area? Well, not particularly, but I do like the idea of clean recovery because essentially, again, we're sort of addressing the elephant in the room, which is that in many cases, we don't have a clean data set. Sure, we can recover back to X period in time, but is that actually clean data? And also importantly, another thing that I've been looking at is it's often not all of the data that needs to be rolled back to that point in time. How do we handle a use case where we have a large data set, a bit of it has been corrupted or encrypted or deleted or whatever, and the rest of it is still clean? How do we adjust to that world? And I think that if we have more of an approach toward treating the data a little bit more granularly, then it'll allow us to have a better recovery than if you're just sort of rolling the entire company back to a point in time, because if you roll it back to this point in time and it's still corrupted, then you're going to roll it back more and roll it back more. Well, you're going to naturally be losing data at that point, and we definitely don't want to be doing that. Yeah. So that does affect that RPO piece, the recovery point objective. How much data do we lose by keeping it clean? And obviously, we're sitting here now in New York City at the SHIFT conference. And one of the things we're announcing is exactly that, is granular synthetic recovery. So the leveraging of threat intelligence, anomaly detection, and AI so that Commvault itself can construct you a synthetic backup based on what's clean, which hopefully is going to really help those recovery use cases. So looking forward then, how do you see this ResOp thing evolving over the next, let's say, 12 to 18 months or so? What are going to be the big moves there, do you think? Well, I hope that what's going to happen is that getting people together across these disciplines is going to help businesses be able to more effectively and more realistically recover. I hope that it will allow them to better respond to the really increasing threats that are happening right now with cyber. It's a terrible time because people are applying AI to develop ever more difficult to defeat cyber threats. We're seeing ever better phishing attacks. We're seeing all sorts of problems that are essentially going to pose a challenge to every business. And every business needs to be thinking about how can they respond when that happens. Allowing companies to approach that sort of resiliency as a practice instead of approaching the copying of data or approaching the, as you said, incident response or the root cause analysis, having an entire overall approach toward resilience operations, I think, would hopefully help make businesses more resilient. But at the very least, I think that it's important to do what you said, which is to get these folks together, have them sit down at the table and have them discuss really what their needs are and what their capabilities are. Because again, many people on the security side of the divide may not understand what the technology can really do at this point. And many on the technology side may not really understand what the cyber threats look like. And I think that that actually could be not just a positive outcome for businesses, but actually kind of, frankly, an enjoyable outcome for the folks that are involved in those situations because all of them have something to learn. And if I've learned one thing, it's that technical people tend to be very inquisitive. They tend to enjoy these new challenges. They tend to enjoy rising to the occasion. And who doesn't want to be the hero after a cyber attack? Right. Yeah, absolutely. And I'm a technologist. I work for a technology company. I'm incredibly excited about things like synthetic recovery and clean room and some of these innovative technologies that allow us to do some of this stuff. But without the changes to people and culture and process and governance, none of this is going to come to anything at all. So I think, you know, the ResOps thing is exciting for me because it could really be the basis of where this technology fits, how people collaborate, how people learn to measure things in different ways. So I think that's very exciting. So, Stephen, we're out of time. Thank you so much for being with us, first of all, here in New York at our Shift event, but particularly for being on the Strive podcast with us. Thanks for joining us. Well, it's great to be here. Thanks for having me. So that's where we'll wrap up today's conversation and this episode of Strive. Thanks again to Stephen for joining me and to unpack what it takes to make resilience measurable, continuous and intelligent. If today's discussion has got you thinking about how you can strengthen your own organizational resilience, please reach out to us at Commvault. Take a look at the website. Thanks for listening to Strive.