Video: Fortifying Your Unstructured Data: Cyber Resilience Across On-Prem and Cloud | Duration: 1336s | Summary: Fortifying Your Unstructured Data: Cyber Resilience Across On-Prem and Cloud | Chapters: Introducing Unstructured Data (23.615s), Unstructured Data Growth (139.175s), Unstructured Data Challenges (219.505s), Data Protection Considerations (616.83s), Rubrik's NAS Solution (869.765s), Conclusion and Benefits (1213.13s)
Transcript for "Fortifying Your Unstructured Data: Cyber Resilience Across On-Prem and Cloud": Hello, and welcome to our session fortifying your unstructured data sub resilience across on prem and cloud. My name is Bill Dutier. I look after the cloud and security go to market, within our Rubrik X team. And I'm joined today by Stuart Davidson. Stuart, do want to introduce yourself? Yeah, sure. Hi, Bill. Hi, everyone. I'm Stuart Davidson. Been with Rubrik over five years. I am a specialist systems engineer for unstructured data. And then prior to Rubrik, I've worked on both customer and the customer side, both public and private sector. Thanks, Stu. Really pleased that you could join us here today. So look, we're talking about unstructured data and why you need unstructured data protection. We're seeing unstructured data explosion in our customers. Petabyte scale environments are becoming normal. There's huge growth across every industry that we talk to. Know, the investment bank UBS anticipates that total data volume will increase 10x this decade, reaching six sixty zettabytes by the year 2030. That's 129 gig per person on earth for every person on this planet. And by the end of this decade, you know, that's a lot of data. The research from IDC also says that as soon as this year, the end of this year, we will see that 90% of all data will be unstructured. So there's many types of unstructured data platforms. You can see traditional NAS platforms like Dell Power Scale, the NAS platforms from the likes of Pure, Cumulo and NetApp as well. You also see file shares such as OneDrive in Microsoft three sixty five or G Drive in the Google world and then there's also cloud data, so AWS S3, Microsoft Blob Storage. But what is unstructured data? What do people use it for? What's the criticality? Stu, would you mind giving us a quick overview of like the verticalization of this data? Some a couple of examples. Sure. So my background is healthcare. So I call them all ologies like pathology, oncology, cardiology. All these PACS images and MRI scans are fully digitized. Medical devices generate more detailed images and then outside of healthcare, everyone and everything is pretty much generating unstructured data, Word documents, spreadsheets, presentations. Wearable tech is generating data. The sensors in my fridge are generating data. Your car, heating system, EDA, media, and then the latest kid on the block, AI are all generating tons of unstructured data like never before. Thanks, Stu. And you know the data protection landscape has changed too. We're talking about fortifying your unstructured data. The data protection landscape has changed from twenty, thirty years ago from being simple backup and recovery through to cyber threat protection today. So ensuring that your data can survive a malicious attack, a ransomware attack, or even just an accidental insider threat issue. Now, the impact of loss or compromise is severe. It results in downtime to these businesses and the loss of data can be critical, especially as some of this data involves intellectual property and, critical for for those businesses holding that data. Stu, you work closely with customers. What's keeping them up at night, with regards to protecting the unstructured data? Yeah. Of course, no, customers are exactly the same, but there are definitely common challenges that universal. And I'm gonna say that the the main one is the size of unstructured data, and that's size is relative depending on the customer. Some struggle with a couple of 100 terabytes while others are, you know, tens of petabyte scale. And it's not just the size of the data that becomes challenging. So unstructured data now ranges from hundreds of millions or even billions of files. And I think once you get to that scale, the traditional tools that IT teams use, we've actually outgrown them. They're not able to cope with that scale. It's like a never ending backup loop, with increasing backup windows to cluster we have in the in the finance industry, which heavily regulated. They've been trying to use NVMe backups to protect a data set whose footprint is north of 10 petabytes. Those NVMe jobs have been queued up for over six months. So in spite of adding more resources, like tape drives, libraries, etc, the reality is these jobs are never going to catch up on the finish just because NVMP is not designed to handle the scale. Other challenges that I see, again, that worry kind of storage admins is when it comes to scanning the file system to understand what's changed, what needs to be backed up. These scans potentially take that long that the RPO is unachievable. So if I need a daily backup, the scan and the backup needs to be really performing, it can't drag on for days. And then on that vein, kind of the real challenges, some customers genuinely don't back it up. They've deemed that the data set is too large and they've convinced that there's no solution that can handle this scale. So they do run with the risk of zero backup and potential data loss. I think by far and away the most common method is snap and replicate that we see out there. That's where a snap is taken, held short term locally, potentially replicated to a second NAS, provided by the same vendor and sometimes sent to even a third NAS again by the same vendor. And from a disaster recovery perspective, this is, this is what customers go with. This snap and replicate has its own challenges or limitations if you like. So if I'm replicating data and I get hit with a ransomware event, I get data is encrypted on my primary and also on my secondary. So it's now got two copies of encrypted data. Logic dictates from a cost perspective, if I have two or three times the amount of data, costs increase and that's not just for the NAS itself, but this is also environmental factors, power cooling, floor space. And then don't forget the IT admins. This stuff all has to be fed and watered, so management, also increases. Just in terms of the the snap and replicate method and with a focus on cybersecurity, this is relying upon the NAS vendors code base here. If those credentials are compromised, what's the attack you're gonna do with those snaps? Potentially not good. And again, if you're using multiple vendors, the multiple tools to learn different ways to manage and report on that. And from a recovery standpoint, complexity is very rarely a good thing. Just from a I'm sure our audience is aware, but the NAS snapshots are also vendor proprietary. So you can only restore, snapshot data back to the same NAS. So if you have plans to change that NAS at some point in the future, how are you gonna handle those old backups? Especially if you're regulated, you gotta keep stuff for seven or ten or twenty years potentially. And largely, the same challenges are true for unstructured data in the cloud. So in both AWS s three and Azure Blob, they provide native file versioning and potentially replication. Neither is true backup, lacks any data reduction, and the data potentially has to be stored on the same expensive tier, the same region or within the same account, which may not meet the business requirements. And then promise I'm finishing up, other big challenge that I see is compliance. The new regulations that are coming to force have more teeth, so it's less than things like DORA, PRA, the CATH here in The UK, HIPAA. It's less of a checkbox exercise now, from an audit process. You need to actually prove that you can restore a minimum viable business or a minimum viable hospital. And restoring through the lens of a cyber attack and, like, to finish up, I would say a ransomware or cyber attack is probably the nightmare scenario that every admin, that would be everyone's worst nightmare. Well, I mean, a lot to consider when it comes to protecting that unstructured data. So, you know, if I was to also talk a little bit more about the cloud, it's no different to what you've outlined in terms of NVMP, Snap and Replicate. A lot of customers are finding that the cloud has an option to replicate their backups using the native tools. Most times for things like Blob Storage and S3, they're just using the native replication. And if there was an account compromise or a credential compromise, it would mean that those backups, those protected copies are not going to survive in that scenario. And, you know, as I said at the outset of this session, fortifying that unstructured data means ensuring that it can also be tamper proof when it comes to a credential compromise or a ransomware incident. And the other thing that we haven't covered in a huge amount of detail here, the compliance requirements not being met. So Stu, you mentioned that customer that's got months of acute jobs for backups. A lot of these regulated customers have to keep data for five, six, seven years and even longer. And there's a quantum not being met. So, that's also a concern in the cloud because the cost can sometimes drive a decision to say, Hey, I can't actually keep this data for that long because it's cost prohibitive. So, you know, there needs to be a more efficient way to do that. And, you know, Stuart, if I was going to ask you to summarize the capabilities that organizations actually need to address these challenges, like if you could give us a like shopping list of three, four, five things that customers really need to be looking for when it comes to protecting this growing issue of exploding NAS and cloud unstructured data, what would that be? So I'll try and keep it brief. Simplicity. So complexity is the enemy. For me, that's easy to set up and manage is paramount. We've talked about the problems with scale and depth. So a solution that is purpose built for this scale and built to protect that stuff at speed and ideally with a proven track record. You know, someone with reference customers either in the same industry or within the same region. Doing more with less as well. So one of the common assets is consolidation. So from an IT admins perspective, the ideal solution is a single solution that can manage all unstructured data regardless of where it's located. For me, again, with some of the customers I work with, automation, and that means different things to different people, but that could be simply automatically discovering your shares and exports on the NASS and applying a policy so that they're less suited with an SLA in terms of protect every hour, every day and retain that data for the relevant period. But also, as we talk about cloud and automation, when it gets to hundreds or thousands or tens of thousands of storage accounts, it's physically not possible for a human to go through and protect all that. So the ability to have a full API and manage this in kind of a DevOps environment. Probably the last few things for me, an index. So again, if we go to the old school methods, if I need to restore just a file or a folder, having an index at that data so I don't have to mount multiple snapshots and just browse, search for that file, that folder, that path, and then restore it, kind of solving the needle in a haystack problem. And vendor agnostic restores. So, again, really relevant in the regulated industries where you've got to keep data for a long time. The ability to be able to restore that back to any NAS target, or even the cloud if needed in two, three, five, ten years time. I'd also look out for additional capabilities. What else does the solution deliver? Is it future proof? Are we going to migrate to a different system in a year or two, make sure it's covered? And then probably last and maybe most important is cost reduction and cost avoidance. So how can this solution help reduce storage costs and potentially reduce operational overhead as well? I think all those points are really relevant, but the one thing that I see quite a lot from our cloud customers is that migration between platforms. So yes, you've got customers who may have multiple NAS environments in the data center. So you might have a Pure, you might have a Cumulo, you might have a Dell Power Scale or maybe a NetApp. And then we see customers moving that data across the cloud and the data itself becomes a lot more transient. So having a single protection mechanism across those environments over time is really important, especially if you have that long term retention, that regulated requirements, keep that data for a long time. But the other thing that I find is is quite relevant there. So, you know, looking at the speed, the automation, the integration across multiple platforms, the cost efficiency, but then it's that index. Like, we can't underplay that index. Having an index, having a kind of a way that you can simplify the recovery capability is critical in today's cyber threat landscape world. If you don't have an index, if you can't search through that data and ensure that you can minimize not only the time it takes to look for the data, but the time it takes to recover that data with a granular recovery capability, then it's going to extend the downtime, it's going to extend the recovery time, which obviously impacts the business. And, you know, when we look at the types of data, having outages on this data and this critical unstructured data can really stop business operations. So that's paramount, to keep the business running, is having that index as well. So let's talk about how at Rubrik we've solved those challenges and what makes us unique. Steve, we haven't always had a solution to address these challenges at petabyte scale or billions of file density. In fact, it's evolved for us over the last few years as the problem has become more significant in, our customers, and we've invested in this area heavily to ensure that we can meet those needs. Do you want to talk a little bit about how know how Rudrick is different here, particularly in our solution across, you know, NAS as well as cloud? Yeah, sure. So if I we have invested kind of the the the the product name, the the solution we talked about, today, NAS Cloud Direct or NCD for short. And this is what this is the solution we use to protect that petabyte unstructured data wherever it's stored. And what I'll do, I'll kind of break it down into three capabilities. So first up, it's a data scanner specifically designed to scan faster than flash. The scan engine is part of the secret source. We have custom NFS, SMB, and S3 clients that will scan hundreds of thousands of files per second. The second is the data mover. This is built to move unstructured data at line speed no matter what, capable of supporting billions of files and, petabytes of data. And then you touched on it about the data index, which is key. That we provide a real time searchable index for all data, and that gives us granular file level or bulk export recovery capabilities. That index also gives us a few additional things as well. So we can, track file data by age last modified so you can quickly cross a an estate behind the cloud or in prem, identify hot, warm, cool, or cold data. And this will able to you know, enable you to locate and identify stale or obsolete data and then potentially, free up that data, archive it to less expensive storage. I think the other thing from my point of view is those three components, so the scan, the move, the index, all these operations run-in parallel. So each individual thread is monitored for latency and we auto throttle that, meaning we'll move the data as fast as possible without impacting production storage. And again, to help us with that performance, we operate first fall and an incremental forever approach. So once the first falls are done, we only ever need to protect new or changed files. No periodic falls or synthetic falls or any of that sort of stuff. And I haven't kind of finish off a little bit. The NCD is purely a SaaS based offering, so there's no hardware required. Nubrick hosts the management plane, the index, and automatically updates all the components, so there's no feeding and water required here. For on prem customers specifically, you deploy one or more virtual machines to the NCD image effectively. These VMs are stateless, so backup data is never stored on them. Jobs auto resume if the VMs are powered off. You can recreate them in minutes. These VMs write data directly onto pretty much any NFS or s three target, so including the public cloud, but directly to the most cost effective tiers. And so the likes of Deep Glacier, Azure Archive, or Ruby's own Cloud Vault, or a combination of those targets. And we support object level immutability. So this ensures that that backup data cannot be compromised and is located away from those production stands that code base. Just the unstructured data in the cloud, so this is also protected with NCD, but it's delivered in it as purely as a SaaS service. So the data mover is delivered as a cloud native Kubernetes automatically. So there's actually nothing customers need to deploy themselves. To kind of set expectations, this is my first rodeo with NCD. I would say once the firewall rules are set up, probably about thirty minutes is usually more than enough to have this configured at the races. That in itself helps customers to protect at speed, protect the depth of the data, to ensure that we can scale to petabytes and be cost effective because that incremental forever capability is also incremental forever in terms of the storage. And of course, within the platform, we are very efficient with how we encapsulate the data as well. Now, one of the other things I think is really fundamental is the ability to couple that protection mechanism and that index with the ability to detect ransomware. So adding on top Rubrik Security Cloud security capabilities to look for anomalous changes in files and to look for levels of entropy to search for suspicious levels of encryption, but also couple that with sensitive data monitoring and data security posture management in the cloud helps customers to ensure that they have a single capability to not only protect that data, but to understand, is it sensitive? Is it at risk? Has it been tampered with? And that allows customers to also ensure that they could fast forward any recovery if there is a compromise. So, I think one of the interesting stories we have in our customers, University of Reading actually uses Rubik NAS Cloud Direct as they transform their research data platform, which was a growing challenge. They moved towards a platform with Nutanix files and they've had some challenges. They had a lack of data visibility and control, vast amounts of fragmented NAS backups, across lots of files and folders make data management and monitoring difficult. There was insufficient, ability to identify the blast radius of an attack if that was to happen. They had no tools. They couldn't deploy, tools like, endpoint detection and response tools that you would traditionally see in file services. They couldn't do that on their new Nutanix Files platform. And when it came to recovery, things will be very manual. So, using snapshots was one option, but it just wouldn't give them the ability to separate the production data and the backup data, and it just wouldn't survive intact. So Nasdaq Direct enables the University of Reading to effectively protect that data at scale, identify if there was any compromise or any tampering with the data, to also couple that with the ability to test recoveries and to ensure that not only are they meeting their recovery point objectives, I. E. Their backup windows and their backup objectives, but also being able to test how long it takes to recover that data. They achieved a 50% faster file recovery with targeted and rapid restoration of critical data. They also streamlined and automated their NAS protection. So as this data is growing, customers don't grow their teams. Their human capital doesn't grow at the scale of this data growth. And on top of that, it reduced their cloud storage footprint by 70%, really improving the cost efficiency, and the scalability of this platform. So thanks to you for joining us today. And, thank you everyone for joining our session on fortifying your unstructured data, sub resilience across on prem and cloud.