Video: Fortifying Your Unstructured Data: Cyber Resilience Across On-Prem and Cloud | Duration: 1572s | Summary: Fortifying Your Unstructured Data: Cyber Resilience Across On-Prem and Cloud | Chapters: Unstructured Data Explosion (26.654999s), Unstructured Data Challenges (229.25s), Cloud Protection Challenges (519.93s), Rubrik's NAS Solution (872.90497s), Cloud Data Protection (1051.1799s), Conclusion and Case Study (1200.865s)
Transcript for "Fortifying Your Unstructured Data: Cyber Resilience Across On-Prem and Cloud": Hello, and welcome to our session, fortifying your unstructured data cyber resilience across on prem and cloud. My name is Bill Dutier. I look after the cloud and security go to market, within our Rubik X team, and I'm joined today by Stuart Davidson. Stuart, do you wanna introduce yourself? Yeah. Sure. Hi, Bill. Hi, everyone. I'm Stuart Davidson. Been with Rubik for five years. I am a specialist systems engineer for unstructured data. And then prior to Rubik, I've worked on both customer on the customer side, both public and private sector. Thanks, Jue. Really pleased that you could join us here today. So, look, we're we're talking about unstructured data and why you need unstructured data protection. We're seeing an unstructured data explosion in our customers. Petabyte scale environments are becoming normal. There's huge, growth across every industry that we talk to. You know, the investment bank UBS anticipates that total data volume will increase 10 x this decade, reaching 660 zettabytes by the year 2030. That's a 129 gig per person, on Earth for every person on this planet. And by the end of this decade, you know, that that's a a lot of data. The, research from IDC also says that as soon as this year, the end of this year, we will see that 90% of all data will be unstructured. So there's many types of unstructured data platforms. You can see traditional NAS platforms like Dell PowerScale, the NAS, platforms from the likes of Pure, Qumulo, and NetApp as well. You also see file shares such as OneDrive in Microsoft three sixty five or g Drive, in the Google world. And then there's also cloud data, so, AWS s three, Microsoft Blob Storage. But what is unstructured data? What do people use it for? What's the criticality? Stu, would you mind giving us a quick overview of, like, the verticalization of of this data, some a couple of examples? Sure. So my my background is health care. So I call them all ologies, like pathology, oncology, cardiology. All these packs images and MRI scans are fully digitized. Medical devices generate more detailed images. And then outside of health care, everyone and everything is pretty much generating unstructured data, word documents, spreadsheets, presentations. My wearable tech is generating data. The sensors in my fridge are generating data, your car, heating system, EDA media, and then the latest kid on the block, AI, are all generating tons of unstructured data like never before. Thanks, Stu. And, you know, the data protection landscape has changed too. We're talking about fortifying your unstructured data. The data protection landscape has changed from twenty, thirty years ago from being simple backup and recovery through to cyber threat protection today. So ensuring that your data can survive a malicious attack, a ransomware attack, or even just an accidental, insider threat issue. Now the impact of loss or compromise is severe. It results in downtime to these businesses, and the loss of data can be critical, especially if some of this data involves intellectual property and, critical for for those businesses holding that data. Stu, you work closely with customers. What's keeping them up at night, with regards to protecting the unstructured data? Yeah. Of course, no no customers are exactly the same, but there are definitely common challenges that are universal. And I'm gonna say that the main one is the size of unstructured data, and that's size is relatively dependent on the customer. Some struggle with a couple of 100 terabytes, while others are, you know, tens of petabyte scale. And it's not just the size of the data that becomes challenging. So unstructured data now ranges from hundreds of millions or even billions of files. And I think once you get to that scale, the traditional tools that IT teams use, we've actually our own, and they're not able to cope with that scale. It's like a never ending backup loop, with increasing backup windows to a cluster we have in the in the finance industry, which heavily regulated, they've been trying to use MDMP backups to protect. A data set is footprint is north of 10 petabytes. Those MDMP jobs have been queued up for over six months. So in spite of adding more resources, like tape drives, libraries, etcetera, the reality is these jobs are never gonna catch up when they're finished, just because it's NDMP is not designed to handle the scale. All the challenges that I see, again, that worry kind of storage admins is when it comes to scanning the file system to understand what's changed, what needs to be backed up. So these scans potentially take that long that the, the RPO is unachievable. So if I need a daily backup, the scan, the backup needs to be really performing. It can't drag on for days. And then on that vein, kind of, the real challenge is some customers genuinely, don't back it up. They've deemed that the dataset is too large, and they've convinced that there's there's no solution that can handle this scale. So they do run with the risk of zero backup and potential data loss. I think by far and away, the most common method is snap and replicate. We see how that that's where a snap is taken, held short term locally, potentially replicated to a second NAS, provided by the same vendor, and sometimes sent to even a third NAS, again, by the same vendor. And from a disaster recovery perspective, this is, this is what customers go with. This snap and replicate has its own challenges or limitations, if you like. So if I'm replicating data and I get hit with a a ransomware event, I get data is encrypted on my primary and also on my secondary. So it's now got two copies of encrypted data. Logic dictates from a cost perspective, if I have two or three times the amount of data, costs increase, and that's not just for the nows itself, but this is also environmental factors, power, cooling, floor space. And then don't forget the IT admins, this stuff all has to be fed and watered, so management also increases. Just in terms of the the snap and replicate method and with a focus on cybersecurity, this is relying upon the NAS vendors code base here. If those credentials are compromised, what's the attack we're gonna do with those snaps? Potentially, not good. And, again, if you're using multiple vendors, the multiple tools to learn different ways to manage and report on that. And from a recovery standpoint, complexity is is very rarely a good thing. Just from a I'm sure our audience is aware, but the NAS snapshots are also vendor proprietary, so you can only restore snapshot data back to the same NAS. So if you have plans to change that NAS at some point in the future, how are you gonna handle those all backups, especially if you're regulating, you gotta keep still for seven or ten or twenty years potentially. And largely the same challenges are true for unstructured data in the cloud. So in both AWS s three and Azure Blob, they provide native file versioning and potentially replication. Neither is true backup, lacks any data reduction, and the data potentially has to be stored on the same expensive tier, the same region, or within the same account, which may not meet the business requirements. And then promise I'm finishing up, the the other big challenge that I see is compliance. The new regulations that are coming to force have more teeth, so it's less than things like DORA, PRA, the CAFE in The UK, HIPAA. It's less of a checkbox exercise now, from an audit process. You need to actually prove that you can restore a minimum viable business or a minimum viable hospital and restore it through the lens of a cyber attack. And, like, to finish off, I would say a ransomware or cyber attack is probably the nightmare scenario that every admin, that would be everyone's worst nightmare. Wow. I mean, quite a lot to consider when it comes to protecting that unstructured data. So, you know, if I was to to also talk a little bit more about the cloud, it's no different to what you've outlined in terms of NVMP, Snap, and Replicate. A lot of customers are finding that the cloud, has an option to replicate their, backups using the native tools. Most times for things like Blob Storage and s three, they're just using the native replication. And if there was an account compromise or a credential compromise, it would mean that those, backups, those those protected copies are not going to survive, in that scenario. And, you know, as I said at the outset of this session, fortifying that unstructured data means ensuring that it can also be tamper proof when it comes to a credential compromise or a ransomware incident. And the other thing that we haven't covered in a huge amount of detail here, but compliance requirements not being met. So, Stu, you mentioned that customer that's got, you know, months of of acute jobs for backups. A lot of these, regulated customers have to keep data for five, six, seven years and even longer, and those requirements are not being met. So, you know, that's also a concern in the cloud because the cost can sometimes drive a decision to say, hey. I can't actually keep this data for that long because it's cost prohibitive. So, you know, there needs to be a more efficient way to do that. And, you know, Stuart, if I was gonna ask you to summarize the capabilities that organizations actually need to address these challenges, like, if you could give us a, like, shopping list of three, four, five things that customers really need to be looking for when it comes to protecting this, you know, growing, issue of exploding NAS and cloud unstructured data, what would that be? So I'll try and give you a brief, simplicity. So complexity is the enemy. For me, something that's easy to set up and manage is is paramount. We've talked about problems with scale and depth. So a solution that is purpose built for this, scale and built to protect that stuff at at speed. And ideally with a proven track record, so, like, you know, someone would reference customers either in the same industry or within the same region. Doing more with less as well, so one of the commonalities is consolidation. So So from the IT admin's perspective, like, the ideal solution is a single solution that can manage all our unstructured data regardless of where it's located. For me, again, with some of the customers I work with, automation, And that means different things to different people, but that could be simply automatically discovering your shares and exports on the Nas and applying a policy so that they're lassoed with a with an SLA in terms of protect every hour, every day, and and keep that retain that data for the relevant period. But, also, as we talk about cloud and automation, when it gets to hundreds or thousands or tens of thousands of storage accounts, it's it's physically not possible for a human to go through and protect all that. So the ability to have a full API and and manage this in kind of a DevOps environment. Probably the last few things for me, an index. So, again, if we go to the the old school methods, if I need to restore just a file or a folder, having an index of that data so I don't have to mount multiple snapshots and just browse, search for that file, that folder, that path, and then restore it, kind of solving the needle in a haystack problem. And vendor agnostic restores, so, again, really relevant in the regulated industries where you've got to keep data for a long time. The ability to be able to rescore that back to any NAS target, or even the cloud if needed in two, three, five, ten years' time. I'd also look out for additional capabilities. What else does the solution deliver? Is it future proof? Are we gonna migrate to a different system in a year or two to make sure it's covered? And then probably last and maybe most important is cost reduction and cost avoidance. So how can this solution help reduce storage costs and potentially reduce operational overhead as well? I I think, all of those points are are really, relevant, but the one thing that I see quite a lot from our cloud customers is that migration between platforms. So, yes, you've got customers who may have multiple NAS environments in the data center. So you might have a Pure, you might have a Qumulo, you might have a Dell PowerScale or or, maybe a NetApp. And then we see customers moving that data across the cloud, and the data itself becomes a lot more transient. So having a single protection mechanism across those environments over time is really important, especially if you have that long term retention, that regulated requirement to keep that data for a long time. But the other thing that I find is is quite relevant there. So, you know, looking at the speed, the automation, the integration across multiple platforms, the cost efficiency, but then it's that index. Like, we can't underplay that index. Having an index, having a a kind of a way that you can simplify the recovery capability is critical in today's cyber threat landscape world. If you don't have an index, if you can't search through that data and ensure that you can minimize not only the time it takes to look for the data, but the time it takes to recover that data with a granular recovery capability, then it's gonna extend the downtime. It's gonna extend the the recovery time, which obviously impacts the the business. And, you know, when we look at the types of data, having outages on this data and this critical unstructured data can really stop business operations. So that's that's, you know, paramount is to keep the business running is having that index as well. So let's talk about how at Rubrik we've solved those challenges and and what makes us unique. Stu, we haven't always had a solution to address these challenges at petabyte scale or billions of file density. In fact, it's evolved for us over the last few years as the problem has become more significant in, our customers, and we've invested in this area heavily to ensure that we can meet those needs. Do you wanna talk a little bit about how, you know, how Rubrik is different here, particularly in our solution across, you know, NAS as well as cloud? Yeah. Sure. So in fact, we have invested heavily. The the the the product name that the solution we talked about today, NAS Cloud Direct or NCD for short, this is what this is the solution we use to protect that petabyte unstructured data wherever it's stored. And what I'll be I'll kind of break it down into three capabilities. So first up, it's a data scanner specifically designed to scan faster than flash. The scan engine is part of the secret sauce, so we have custom NFS, SMD, and s three clients that will scan hundreds of thousands of files per second. The second is the data mover. So this is built to move unstructured data at line speed no matter what, capable of supporting billions of files and that, again, petabytes of data. And then you talked to enable the data index, which is key. So that we provide a real time searchable index for all data, and that gives us granular file level or bulk export recovery capabilities. That index also gives us a few additional things as well. So we can, track file data by age last modified so you can quickly, across an estate behind the cloud or on prem, identify hot, warm, cool, or cold data. And this will able to, you know, enable you to locate and identify stale or obsolete data and then potentially free up that data, archive it to less expensive storage. I think the the other thing from from my point of view is those three components, so the scan, the movie, and that's all these operations run-in parallel. So each individual thread is monitored for latency, and we auto throttle that, meaning we'll move the data as fast as possible without impacting production storage. And, again, to help us with that performance, we operate first full and an incremental forever approach. So once the first fulls are done, we only ever need to protect new or changed files. No periodic falls or synthetic falls or any of that sort of stuff. And, yeah, if I can kind of finish off a little bit, the NCD is purely a SaaS based offering, so there's no hardware required. Nubrick hosts the management plane, the index, and automatically updates all the components. So there's no feeder and feeding and watering required here. For for on prem customers specifically, you deploy one or more virtual machines, so the NCD image effectively. These VMs are stateless, so backup data's never stored on them. Jobs also resume if the VMs are powered off. You can recreate them in minutes. These VMs write data directly onto pretty much any NFS or s three target, so including the public cloud, but directly to the most cost effective tiers, and so the likes of Deep Glacier, Azure archive, or Ruby's home Cloud Vault, or a combination of those targets. And we support object level meetability. So this ensures that that backup data cannot be compromised and is located away from those production stand SANs that that code base. Just the the unstructured data in the cloud, so this is also protected with NCD, but it's delivered in as purely as a SaaS service. So the data mover is delivered as a cloud native Kubernetes automatically. So there's actually nothing customers need to deploy themselves to kinda set expectations. This isn't my first rodeo with NCD. I would say once the firewall rules are set up, probably about thirty minutes is usually more than enough to have this configured and and and at the races. And that that in itself helps customers to protect at speed, protect the depth of the data, to ensure that we can scale to petabytes and be cost effective because, you know, that incremental forever capability is also incremental forever in terms of the storage. And, of course, you know, within the platform, we are very efficient with how we encapsulate the data as well. Now one of the other things I think is is really fundamental is the ability to couple that protection mechanism and that index with the ability to detect ransomware. So adding on top Rubik's Security Cloud security capabilities to look for anomalous changes in the files and to look for levels of entropy, to search for suspicious levels of encryption, but also couple that with sensitive data monitoring and data security posture management in the cloud helps customers to ensure that they have a single capability to not only protect that data, but to understand, is it sensitive? Is it at risk? Has it been tampered with? And that allows customers to also, ensure that they could fast forward any recovery if there is a compromise. So, you know, I think one of the the interesting, stories we have, in our customers, University of Reading actually uses Rubik NASDAQ Direct as they transform their research data platform, which was a growing challenge. They moved towards a platform with Nutanix files, and they had some challenges. They had a lack of data visibility and control, vast amounts of fragmented NAS backups, across lots of files and folders make data management and monitoring difficult. There was insufficient, ability to identify the blast radius of an attack if that was to happen. They had no tools. They couldn't deploy, tools like, endpoint detection and response tools that you would traditionally see in file services. They couldn't do that on their new Nutanix Files platform. And when it came to recovery, things will be very manual. So, you know, using snapshots was one option, but it just wouldn't give them the ability to, separate the production data and the backup data, and it just wouldn't survive intact. So Nelstel Direct enables the University of Reading to effectively protect that data at scale, identify if there was any compromise or any tampering with the data, to also couple that with the ability to test recoveries and to ensure that not only are they meeting their recovery point objectives, I. E. Their backup windows and their backup objectives, but also being able to test how long it takes to recover that data. They achieved a 50% faster file recovery with targeted and rapid restoration of critical data. They also streamlined and automated their NAS protection. So as this data is growing, customers don't grow their teams. Their human capital doesn't grow at the scale of this data growth. And on top of that, they reduced their cloud storage footprint by 70%, really improving the cost efficiency, and the scalability of of this platform. So so thanks to you for joining us today, and, thank you everyone for joining our session on fortifying your unstructured data, sub resilience across on prem and cloud.