Video: Predibase and Rubrik: Accelerating the Enterprise AI Journey | Duration: 3426s | Summary: Predibase and Rubrik: Accelerating the Enterprise AI Journey | Chapters: Welcome and Introduction (0s), Introducing AI Conversation (0s), Democratizing AI Access (51.250591206179436s), GenAI Adoption Challenges (190.1405912061794s), POC to Production (296.9405912061794s), Open Source Models (566.0055912061794s), Model Selection Strategy (813.6805912061795s), Advancing Agentic AI (1213.4456912061794s), Prediabase Integration Benefits (1703.9555912061794s), Customer Use Cases (2492.2855912061796s), Future AI Trends (2928.770791206179s), Concluding AI Discussion (3237.2506912061795s)
Transcript for "Predibase and Rubrik: Accelerating the Enterprise AI Journey": Hi, everybody. I'm Anneka Gupta. I'm the chief product officer here at Rubrik. Hi, everyone. I'm Devvret Rishi. I'm one of the cofounders and CEO for Predibase. Really excited to be here today. We're really excited to have a conversation with you all about where AI is going, help you learn a little bit more about what Pretabase and Rubrik are doing together. As Chase mentioned earlier, if you have questions throughout the session, please add them into the q and a tab as we will take live questions both during the conversation, as well as at the end if we have time. Great. So to get started, Devvret, I'm so excited that you're joining us here today. You are at the forefront of working on GenAI, and I was curious if you could share with us to begin with, how did you find yourself working on cutting edge GenAI technology? Yeah. Definitely. It was, a little bit of an inadvertent journey in some ways. I cofounded Predibase in 2021, and for me, the kind of initial mission was to be able to democratize access to AI. The reason that really, stuck with me is because prior to Predibase, I worked in machine learning and AI for probably about a decade, both in terms of grad school and then as a product manager at Google. I worked on Google Research, productionizing machine learning applications, as well as, being the first PM for Kaggle, a data science and machine learning community. One of the incredible things I saw at the time was the growth of what at the what we called citizen data scientists. Essentially, people that didn't have classical machine learning training, but were starting to adopt the technology at heavier and heavier rates. I saw this on Kaggle where we were maybe less than a million users when I joined and grew to be over 10,000,000 users by the time I left. And I think I really saw a large part of the opportunity for where this technology could go. Surprisingly, though, while citizen data science was actually growing and a lot of folks were, I would say, ML curious, the actual applications instead of enterprises were still pretty limited. And one of the things that a lot of folks had cited was the lack of access to ML expertise. And so I decided to partner with my cofounders, who'd been building machine learning infrastructure technology at Uber. Kind of the, different end of the of the spectrum, really sophisticated technology for training deep learning models and then deploying deep learning models, including open source projects like Ludwig and Horovant that come out of that. And so we decided to start the company in 2021 really oriented on this idea of democratizing access to AI and specifically this new technology deep learning that we thought was really powerful. The way I like to tell the story is in 2023, ChatGPT and LLMs democratized AI even better than we were doing at the point. And so we just like to focus fully on one aspect of deep learning, which is really large language models in chat of AI, and help customers be able to not only get democratized access towards them, but help solve the last mile problem that we are really see seeing in a lot of enterprise applications. And so that really was the journey, for us over the last four and a half years, and I'm really excited for what the future of the journey is gonna look like for us as well. That's amazing. It's, I think everyone that works in GenAI, like, found their way there, obviously, in a circuitous manner because no one could have predicted that this is what this like, what would happen over the past five years happened. So, it's really interesting to hear that. So you talked to a lot of organizations today that are somewhere on their journey of leveraging Gen AI. What would you say surprises you most about the state of GenAI adoption today, especially in large enterprises? Yeah. I'll say that there's probably two things that surprise us. Today, we work with a number of different customers from leading tech companies here based in Silicon Valley all the way to large Fortune 200 enterprises. And so we really see kind of the entire spectrum of four different customers come in. One of the first things that surprises me is that, you know, I I live in San Francisco Anneka, and I I kinda see the ubiquity of AI here, in Silicon Valley. But I think when I talk to some of the larger enterprises, it there's a really real security concern that I think blocks a lot of these enterprises from being able to actually adopt Gen AI. One of the most surprising things that I kind of hear for me is when the large enterprise, that I speak to kind of talks about how maybe chat GPT or, g t four was blocked internally for a year or longer, because of concerns around sensitive data, leakages, and others that might not prevent them from being able to go forward. The reason I see that being as, you know, really surprising is I think a lot of the companies that I work with have identified the genre of AI as can be core to maintaining a long term competitive advantage and differentiation. And so for us, we really wanna make sure that these companies are able to onboard onto that tech as quickly as possible so they they can start to build that competitive advantage on top of a platform like ours. That's it. That's kind of, you know, surprise number one. I think surprise number two is, the gap between POC and production. I know for a long time as people have experimented with different technologies, people have often talked about how the the need to graduate from POC to production often ends up being a stumbling block. But I think that this has really been, extenuated in the generative AI era. And one of the, I think, main reasons for it is kind of embodied in a one of my favorite customer quotes that I hear, which is that generalized intelligence might be great, but I don't need my point of sale system to recite French poetry. Pretty much what I consistently hear is people are able to build all these amazing demos quite quickly and, you know, build a proof of concepts very, very easily. But there's a gap before they actually are able to make it into production. And, you know, I think Gartner actually has a study where they kind of cited that between POC and production, 80% of the projects end up failing or taking more than six months. And so that is probably the second thing that when we started the journey in particular, we found to be really surprising and where we started to really orient a lot of our products to be able to help on. Yeah. That really resonates with me because I've been talking to our customers at Rubrik, about how they're approaching GenAI or just AI in general, and a lot of the similar themes have emerged in our in those conversations as well. There's a huge concern around overexposure of sensitive data. For instance, like someone in the HR team getting access to finance data that they shouldn't or someone in engineering getting access to HR data that they shouldn't. And this whole idea that you have around democratized access, like, there's the con of that is, like, is it too democratized and are people getting access to stuff that they're they shouldn't? The other thing that that, we you know, also even internal experiences, like, concerns around, well, is my, like, confidential IP going into some third party that I don't know what they're doing with it and how that data is going to be used? And that's, that's something that we have heard a lot, as well. So it's those things as well as in, like, how do you actually go from prototype to production, quantify the top line, bottom line results. Like, all these things are are really challenging. And I can share that even here at Rubrik, we struggle with those same things. And we're trying to figure all of this out as we're looking at what is the potential of GenAI to change the way that we all do work and the way that we create, and and the way that we innovate quickly in our space. Yeah. Exactly. And I think that's really well said. For me, I think seeing a lot of the positive impact to, Jenny, I can have both, I'd say, as a consumer, but in particular on enterprise use cases in terms of being able to drive additional value, save costs, and mitigate risk. I am really kind of passionate and excited about the idea of getting led of some of those roadblocks so that people can start to actually, you know, see this value directly in in mass. And, you know, in large part inspired by some of the customers we work with today that have already been able to see that value. Yeah. Totally. And, I mean, I I see it just for ourselves here at Rubrik. Like, we enable Gemini, for the entire organization and then just the grassroots ways that people are figuring out how to leverage Gemini to get more productivity. It's really incredible, and it's really inspirational. I think the more like, I'm a huge believer in the the more that we like, that enterprises can safely enable AI for their organization, and give these tools to people while having the right guardrails around both the security, but then the quality and cost aspect as well, the better off we'll be because the best ideas are not gonna necessarily come top down. They're gonna come bottoms up as well. I don't know if that's something that you see as well when you're talking to customers. Definitely. In fact, with AI, I think that what we tend to see is that, a lot of organizations have kind of taken the let a million flowers bloom approach where it becomes very inexpensive to test for any individual developer, and you see, like, number of different use cases. The number of, I think, Fortune 500 companies I talked to that say I don't have one or two generative AI use cases. I have 50 plus that I'm actually working on right now is, really amazing. And I think, you know, I do resonate a little bit with the the pain that might introduce for a CIA or CSO, but now I have to wrangle with the fact that there's AI developers both new to AI as well as, you know, existing in that space that now have access to this technology and maybe are looking to be able to apply them in ways that weren't, you know, previously locked down. Yeah. Yeah. Totally makes sense. Just wanna give a reminder to everyone out in the audience. If you have a question, you you have an expert here with Devon, so it'd be great, if you shared some of the questions that you have top of mind, and we'll we'll answer those. So good. Moving on, what are some of the unexpected headaches that teams encounter when they try to deploy off the shelf commercial models like OpenAI within a serious enterprise environment? Yeah. It's a great question. And I'll start off by saying that I think, using an off the shelf model is, in my opinion, like, an excellent way to get started. When you're able to go ahead and actually very quickly be able to, see some value, I think it really helps you consolidate your gains. Where we see people struggle with using an off the shelf model today is when they're trying to go beyond that step of I built an initial proof of concept or demo and into the step of I'm actually going out into meaningful scale and production. And, you know, I think that there's there's a set of challenges that we've really seen. Everything from latency to restrictions on rate limits, cost, especially at a large scale with some of the largest frontier models, and, of course, quality. You know, the ability of that model to not just be 70% accurate, but something that you need 90% plus on. So we've seen a lot of different pain points as people try and take the same initial setup that they sort of had working for that demo for their boss or for the internal use case and are looking to be able to scale that up, you know, significantly. To me, we oftentimes bucket it into three different areas, that we can look at, helping our customers with. The first is, reducing cost. The second is improving call, quality, and the third is reducing latency. The fourth theory that oftentimes comes up as well, is being able to actually take some ownership of that model. So this idea that, I want the model deployed inside of my environment, without necessarily having to send data across my four walls over to a third party external service, like OpenAI or any of the other models. And so to us, these are kind of like the primary value propositions that we really play in. And I just wanna maybe, take a sneak peek to, I think, like, actually go to the fact that a lot of the ability to play this value proposition has really been driven in large part because of the open source community. And, you know, one of the most amazing trends that I think we've seen is the shrinking gap between open source and closed models. I hosted a dinner with about 15 or so AI leaders, about a year ago. And after the dinner, we spent some time talking about these, Jira applications, and somebody came up to me and said, I want to believe. I wanna believe in the future of open source models, but I'm not sure that they're quite there yet from a quality standpoint. And, you know, I don't blame them about when we made the decision to go all in on open source models, one of the leading models was GPTJ, which is not a great open source model. But what's happened ever since then is, you know, with the investments from teams at Mistral, Quan, Llama, and then, of course, the things that really, I think, broke a lot of ground with DeepSeek earlier this year. We've seen these open source models catch up, and I think not only get to the point where they're competitive with frontier models, but in some cases, in some use cases, really even exceed that basis. And so that's been really compelling. And the reason this matters, I think, is just coming back to the unexpected headaches, that people run into. Well, they tend to be around cost, latency, ownership, and quality. And what we see is, you know, for the first three, being able to deploy smaller models, open source models inside of your environment with full control really help a lot of customers be able to get over that piece. And then on the quality side, we've obviously built a larger investment and the ability to help customize, tailor into models towards your individual task and data. And so those are kind of the the main types of, I think, challenges I've seen as well as maybe the biggest tailwind that's helped customers be able to address those challenges. What do you like, when you talk to customers and and when you evaluate the kinds of use cases where, foundation models make sense versus, using open source, like, what do you typically see, and how do you kind of recommend people approach their use case and matching it to the right model for them? It's a great question. So the first thing I'll go ahead and say is, one of those common questions I initially get asked is, what model should I use for my task? And it's a really hard question to answer a priority without actually having done a little bit of the experimentation. But luckily, at this point, we've done enough experimentation in a duration where I think we've been able to go ahead and detect some patterns. So, typically, one of the things that I'll recommend upfront is as a customer is building out their first production use case, use the most capable model that you have access to. Now the most capable model might be something that's a little bit slower or very expensive that you wouldn't be able to go out to scale, but helps you validate, is this a task that generative AI can actually do a decent job at solving in the first place? So, you know, one example from a customer that we work with might be doing, like, automated document classification or extraction of key terms inside of a document. The first step before choosing a model is getting access to the most capable model, whether that's a large frontier model from one of the leading labs or a great open source model. That helps you actually be able to identify which task you're looking at or, you know, which kind of solution form factor you might be looking for in that. But the second step, I think, is actually as you think about what the production scale could look like. And there, I think for customers, the main kind of, considerations that you would wanna be able to make is also around cost and latency. And so as you think about this, one thing we work a lot with our customers on is just getting a forecast for what is the level of volume that we might anticipate seeing, for this use case over the coming six or twelve months. And how does that rationalize with the ROI you might wanna be able to see? I think one of the most compelling things that, a lot of customers might be unaware of is the level of difference in cost savings that you might actually be able to get from being able to choose a smaller model tuned towards your data, than using a larger model. And so with a lot of our customers, we see savings of sixty, seventy, 80% plus, you know, in comparison to using what they might start off with as the most capable model with OpenAI. And so I think that there's one aspect of this, which is making sure you start with something that's really great and then back out towards something that's really cost effective or latency effective for what you're looking to be able to do. And then I think there's a lot of really good literature. And once you're looking to make that second decision of, okay, what's the right small model maybe for my use case? And, you know, we actually at Predibase published something called the, Predibase fine tuning index where we've benchmarked, you know, dozens of open source LLMs across 30 different, tasks. Everything from, just to give you a sense, like, everything from doing things like, entity recognition to content generation, code generation, legal clause classification, and others. And you can actually see how different models on our leaderboard have performed on these different tasks. So if you're, looking for maybe, you know, a very quick shortcut and you have a task that looks like one of the ones that we've benchmarked on, we do maintain the, fine tuning index that will give you a quick way to be able to get a sense of which open source models might do particularly well here. And maybe, one question that, some folks might have is how much technical expertise do you need to have in order to leverage fine tuning? It's a great question. I think the honest answer is, at least a year and a half ago or two years ago, the the technical expertise in bar was relatively high. I think today, platforms like ours have actually drastically lowered that. The way that I think about it is if you're an AI engineer comfortable iteration that you might be doing with an OpenAI API, you're actually probably well, ready to do fine tuning as well. And there's a couple of reasons that I think that's the case. The reason that fine tuning was viewed as this really cumbersome activity, maybe, you know, two years ago, is number one, it was expensive. Fine tuning a model, oftentimes when I talk to customers, they're worried that might cost them hundreds of thousands or millions of dollars. But that's not the case. Today, most customers can fine tune models for $20.50, 100, or a few $100 itself. And so there's really kind of a significant cost saving that's come in from fine tuning. The second is just the level of accessibility. If you were fine tuning maybe a few years ago, you had to not only fine tune the model and figure out the right parameters and the weights and how to be able to run that on your GPU instance and get a GPU instance in the first place, which can be a bit rare. You got to actually then figure out how do you wanna serve and deploy these models. You know, one part of it is fine tuning a model, but that another part of it is actually being able to go ahead and deploy that fine tune model, into production. And so this is, you know, an area where I think as a platform, we really made a a huge leap in being able to, not only offer end to end fine tuning that's accessible for customers, but also be able to make sure that you don't have to worry about the massive amount of infrastructure that goes into the pieces of steps after you train a model. So right out of the box, you know, one of the things that I think we help provide for a lot of folks is, no code fine tuning, where you can actually get started even just directly in our UI to start to train models using an any number of the different, techniques that we might have. And we've seen a number of customers really get comfortable and get started with with something like this, very early on. And I think this allows you to be able to iterate quite quickly. And the thing I'm particularly proud of is building kind of an integrated tuning and serving stack, which means once you've tuned, you were immediately able to go ahead and serve and deploy that model as well. And so this allows you to be able to make sure that you don't have to worry about, you know, the conventional infrastructure headaches that have to go with things like auto scaling or post deployment monitoring, and others. So, you know, maybe the short answer to your question is I think it was reasonably. As I think it the fine tuning most people expect is gonna be more work or harder than it actually, will be. And I think a large part of it is platforms like ours have really kind of lowered the barrier to entry to not only doing the training, but also the deployment, of those two involves. That's great. So there was a question from the audience that I think will resonate with a lot of, the the audience here, which is a lot of the places where, organizations are starting are leveraging things like Microsoft three u 65 Copilot within the enterprise Microsoft three sixty five license. So what thoughts come to mind for you around the best ways to leverage us, get started, and then maybe where to build from there? Yeah. Absolutely. I think one of the reasons that Copilot are, so popular today in, like, the Office three sixty five one as an example is it provides a really familiar interface towards how how you can actually look to, essentially get, like, a chat sheet BT style question answering flow on top of data that might otherwise exist inside of Microsoft Office three sixty five. And I think that's a I think that's, like, a fantastic tool and gateway until where a lot of generative AI actually starts to get used. So to me, I think in particular for search and retrieval and knowledge tasks, being able to use Office three sixty five Copilot, in particular, maybe in a chat setting, is a really, productive way to be able to solve a lot of the internal search style use cases. I think where we see or or where I kinda see the trend going next after people start to adopt these co pilots, most co pilot flows I see today are question answering flows. But if I think a lot about where is the kind of long term enduring value for generative AI going to, come in, I do actually believe very heavily in Ajentic AI. And I think Ajentic AI needs to be defined, but in my view tends to mean three things. The first is, you know, it typically chains multiple LMM calls. The second is it gets access to tools, so you can call functions and tools. And then the third is that, it's typically used to be able to automate some work. And this is where I think, like, the next generation of Copilot plus some build at your own are really gonna be focused on, which is how do you take, what was kind of conventionally built on top of, like, a RAG technology, maybe in terms of retrieval augmented generation for being able to do search and discovery internal, for your internal data. And now start to build an automation flow where your AI might actually be taking some action on top, of the of on top of the data it has access to and on top of the tools you've given it access to as well. And so to me, the rise of, I think, Adjunctic AI is gonna be where a lot of customers will naturally go next after, they are able to do some of the initial Copilot like flows. That's actually very similar to a use case that we've seen with one of our customers, which maybe I'll just share briefly. So one of our customers, who we've been really fortunate to work with is, Marsh McLennan, a Global Fortune 200 company who initially built, started building a, a Copilot very similar to maybe what Microsoft Office three sixty five Copilot, also offers for a lot of customers out of the box. It was a way to be able to assist their, 100,000 employees internally, be able to, do knowledge work. So the first thing that this agent system did was it was able to actually just, answer questions that people might have on their internal systems. I think that's a great use case that you might be able to start off, here. But where I've seen it actually kind of advance afterwards is they started to build a number of tools entire, inside the same agentic flow. And these tools can actually automate work, like coming up with architecture diagrams, or being able to extract information from docs. And so what we've actually seen kinda as a natural evolution is you went from a more knowledgeable people to an agenda k I flow. And we've actually done a a really nice webinar where they speak a little bit about how they built that entire, routing model and the ability to build this agent model directly on top of Credit Base that's accessible, as well on our YouTube channel. So how big of a deal do you think AI agents are? Like, if chat, kind of the search and question answer maybe is giving people, like, 10 and 10 to 20 productivity gains, what do you think the possibilities are with the Genentech AI? Yeah. I think that, you know, if we're seeing, let's say, percentage, gains from, information retrieval use cases, I think Adjuntic AI needs to be more, narrowly tailored in terms of what it's used for. You can kind of throw chat and search at anything. With an agent application, you really wanna make sure you understand kind of the guardrails of where is where should it be used, what do you feel comfortable with an agent doing fully end to end, and what might you need kind of some more, in the loop approval for. However, I think, like, the overall impact is gonna be, like, if it's, you know, 20% for chat, it'll be, like, a factor, two, three, five x when we talk about agent AI. And the reason for that is I think that what we've already started to see is these agents are able to automate workflows that would have naturally taken you know, in March and plan's case, I think they conservatively estimate a million employee offers hours. That's a really dramatic amount of time saving and economic impact when you think about, like, what the underlying, kind of result of that, can look like. So the the impact for agents is gonna be absolutely massive. But I do think that there's a lot of, things you need to watch out for when you're building an agent application. It's a little bit less straightforward than maybe building initial chat application. Yeah. That makes sense. And how do you think about like, I I think one of the the big differences, obviously, between AgenTic AI versus traditional RPA solutions is that AgenTic AI can adapt on the fly, and it's nondeterministic in its approach. What are the pros and cons do you think of the fact that just anything that inherently relies on LLMs, like, is is a probabilistic approach versus deterministic? Yeah. I think that that is actually one of the things that makes a lot of agent AI systems today still feel a little bit brittle. So, you know, I think that one of the realities is that agentic AI systems tend to, really require high accuracy for two reasons. The first is a lot a lot of these systems are actually, looking to be able to make a decision or action. So it's not just, maybe giving you back some information, but it might actually be taking an action on your behalf. And so if you're taking actions, maybe a 70% accuracy isn't quite, quite satisfying because if you are taking the wrong action a third of the time, you know, it could actually become really damaging to reverse your overall system. But the second thing that we see is that agent AI workflows tend to be a little bit more complex than, standard, you know, single pawn to all on workflows. And one of the reasons is agent workflows tend to chain multiple AI calls together, in order to be able to complete an action. So, you know, with, the the use case example from Marsh as one example, the first step will be understand what the user wants. Do they want to extract information or document or write a write an architecture diagram or something else? Then the second step will be get the inputs for the user as an example. Okay. It is information from document. What information? Then the third step might be actually completing that. So you, and that's, like, maybe one very simple flow. We see many kind of flows where you actually have maybe ten, fifteen, 20 plus calls before you're going into that. I think this makes it really critical to focus on two things. The first is you need to be able to tune for quality. And the reason is that if, you know, you let's imagine you can get a model that's 90% accurate out of the box at each step. Once you're compounding this over multiple steps, say five steps, you're you're dropping dramatically in, like, the overall accuracy that you're able to see from the system. And then the second thing that we often see is, latency and speed become particularly important in these agent applications because, you're oftentimes treating this agent as something that's working maybe in real time, and maybe you're comfortable waiting a few seconds for a model to give you a response. But if you're waiting a few seconds times five steps or 10 steps, it can get a lot, get lot longer as well. So this is where we see, again, with agent applications, this need to be able to actually build more purpose built and also tune models as being a particularly compelling area. Awesome. And there was a question from the audience about the Marsh use case. Is that was that done through Predibase? Yeah. It's a great question. So the Marsh use case is actually something that I think, followed the pattern that I was discussing a little bit earlier, where the very first instance of the use case that they built was using OpenAI. And then what they started to do is actually migrate pieces of that, AgenTek workflow over to Predibase. And it was exactly kind of that initial prototype to production and postproduction optimization workflow that we saw as being particularly useful. So what, they initially saw was while they were using some of the OpenAI models to be able to query, and be able to do things like routing or intent classification, they were choosing which model they used on the basis of cost as well, as, like, they're trying to balance cost and accuracy in some way, and they were choosing a less expensive but also a little bit less capable model, overall. One of the first things that I think, really resonated for Marsh was that they could actually get you know, at the time, it was GPT four level quality and GPT 3.5 prices, by tuning a smaller language model to be able to do what that large model was able to do. And so I think this was a great example for, where a customer started off building towards OpenAI and built kind of an initial workflow using that, and then, was able to actually migrate those calls over to Predibase to be able to actually get to more resilient kind of long term production solution that improved accuracy as kind of, like, the primary reason for what they were looking to be able to do. And for anyone that's, you know, curious to be able to dive into this, at any kind of further level, We actually do have as one of our featured customer case studies, and there's a great, case study that's been written about how they actually started to do this as well as a video where you'll be able to see some of the, Gen AI engineers who sit within IT at Marsh, actually work through building out this agent application. Okay. We have another question around model context protocol. What are the most significant security and data governance challenges companies should anticipate when creating connectors to internal proprietary data sources, and what best practices does Predibase recommend to mitigate these risks? Yeah. It's a great question. I think that with MCP, and model context protocol maybe as a quick review for the audience, you can think about MCP as essentially creating an API layer that LLMs know how to speak that allows you to understand, different data sources, and data syncs. So places you might both read data from and then be able to write out, you know, LLM responses towards. So it's kind of a attempt to be able to construct a universal, API interface that LLMs understand across a number of applications that you may know and love today. Salesforce, Slack, many other folks, anyone can create an MCP server essentially for an application. It's a really powerful technology that I think has a lot of, future, applicability. And I think as a question kind of correctly identifies, when you have a very powerful technology, there can also be a lot of concern about what does that mean from a security standpoint. To me, one of the things that I think stands out the most is permissioning with MCP. And so once you have the ability to go ahead and connect across multiple different MCP servers and the ability to go ahead and move from different sources and syncs, I think one of the, biggest challenges is, like, how do you know what kind of the level of permission for each individual user across each individual data source and sync should be? And I think this is actually a very, very challenging problem. So if, you know, for example, customers, like, querying Salesforce data and then being able to go ahead and write that Salesforce data results back into, you know, a a workflow or something that's automated in Slack, What access to which records does that customer have access to in Salesforce, and should they be able to share that within a certain set of channels in Slack? Just a very small example. So so I think permissioning is one of the most interesting challenges that MCP, essentially introduces. And in terms of best practices, I think that the frank answer here is I think we're gonna see a lot of these actually develop over time, because I think when I think about MCP, it's a newer technology still getting kind of, I think, a larger adoption. And it's I it's oftentimes not just the read it's not just the read behavior that you have to be worried about. It's actually also the write behavior where you might be able to do something destructive, as well with the kind of the overall technology. So to me, I think that's, the area that's gonna be the most interesting to consider from a cybersecurity angle for MCP. So, one question that, came up was how, how do you get started testing with AgenTic AI? How do you get started testing with AgenTic AI is I think that with a lot of things, getting started is first by identifying a use case. And so this, I think, I often encourage, folks to do by choosing something that would be really simple to get started off with, but you'd be able to go ahead and actually prove value, very, very quickly. A pretty common getting started use case is something like, classification. So maybe what I'll do is walk through one additional customer case study and just walk you through how I think I could see that and actually starting to develop into, you know, like agent AI. So one of our customers, is Checkr. Checkr is an e verification background check company. It automates millions of background checks, using AI. And, what they you what you can think about the task as really being able to do, again, some of the core workflows we see in agent, workflows, which is like automation, classification of documents, and extracting out information from documents. Where this, where Checker and the ML engineers, they really got started was by using OpenAI and prompting OpenAI to be able to start to see how it was actually doing on this initial journey. And so what we they were able to observe, and they had a actual classification accuracy metric to be able to do this. But what they were able to observe was that they were getting accuracy anywhere from mid to high seventies to low 80, just using GPT four at the time out of the box. It was actually more expensive than they'd be able to go into production with. It was also more, it was slower than what they would ideally like. But those weren't even the main issues. The main issue was background check automation is really Checkers core business, and so 80% just wasn't good enough for what they were looking for. One of the really compelling things that they were able to do though was they got started building kind of this initial experimentation using, GPT four and JetF AI. But they were able to create, you know, a fine tuning dataset. So they had examples from their previous business of what were the correct classes and what were the correct information extracted. They had thousands of these examples. They were able to fine tune a smaller model using Predibase. And they kind of almost immediately saw a significant gain in accuracy getting past 90% for where they're looking to be on their quality threshold. And so to me, getting started with Adjunctic AI is very similar to getting started with just, let's say, any AI use case with one exception, which is, like, get started, use a more capable model, and then, see how you're doing and then be able to fine tune or choose a smaller model once you're thinking about your production applications. I mentioned that there's one exception, which is I think having a good evaluation is in particular even more important when you're thinking about agent applications. Because agent applications tend to be so quality sensitive, you wanna make sure you can actually understand what does good look like, when you're deploying an agent application system. So in Checkr's case, one of the evaluations that they were using was classification accuracy as an example. And, similarly, I think whatever your agent application is looking to be able to do the deliveries and end value proposition, you wanna make sure you can measure as you're doing with the agent system. Yeah. That's a great tip. So as of June 25, Rubrik and Predibase have entered an agreement to join forces. And for Rubrik customers who aren't yet familiar with prediabase and vice versa. Can you articulate how our joint work is gonna help address a lot of the GenAI adoption challenges that we've been talking about throughout this webinar? Yeah. Definitely. And first of all, I just wanna say I'm really excited, I think, about the prospect of what we would be able to build together. And that excitement really, I think, comes to that so understanding some of the key challenges that people may run into today. So, what I feel kind of most strongly about in terms of helping, folks with is the fact that a lot of people are experimenting with generative AI pilots, but they're not necessarily making it out over into production. And I think there's, essentially, like a two by two of, like, different challenges for why we think that might be. On the left hand side is something that I think we see quite often with a lot of the product based customers that we speak with, which is that the quality of the models was good enough for my pilot, but it's not really a production solution. Or maybe quality was actually okay, but the efficiency of the model, which we'll say is cost and latency, is not what I actually need to be able to go into at scale production deployment. So I think this actually comprises a lot of types of settings where you have solved some of the initial challenges of, like, can I even build a and, you know, can I build something that works on top of my data? And you're looking to be able to get across the next set of challenges, which is what is this gonna look like at scale? But there's actually a fundamental set of challenges that even happen, sometimes before someone is at the point where they're ready to be able to use something like Predibase out of the box, which is, well, what if you actually have a restriction in how you're able to use data with your model in the first place? So whereas the challenges I think on the left hand side oftentimes deal with kind of the nitty gritty of, I I wanna optimize my workflow. The challenges on the right hand side oftentimes just see as blockers to be able to get started with, you know, like, in the first place. Concerns about I think, Anneka, you had a great example of, like, you know, what if HR data is leaked to the wrong area? And, like, who actually has the exposure, the access towards it? So to me, I think these are the types of challenges and bottlenecks we see. Sometimes immediately getting out the door challenges with data and data access, and sometimes, you know, things that allow you to crawl but not walk or run, before you actually have the quality and cost latency. I think where we see a lot of the, actual future potential for us to be able to create a really strong joint solution that's oriented on kind of the value propositions of being secure, specialized, and scalable is the fact that I think Predibase naturally solves a lot of the challenges on the left. We work on quality and efficiency consistently and constantly today in terms of helping people have the best post training stack to be able to tune models for those tasks and then be able to turbocharge their efficiency for cost and latency. And Rubrik, I think, has one of the best, foundations to be able to solve the challenges on the right. Because Rubrik not only has a lot of the understanding of the data that is actually there within the backup, but also the security and the permissioning for that data. So that you need in order to be able to understand what data can be accessed by which users, which data is sensitive and needs to be suppressed before it goes into an ALM, and otherwise. To me, when I look at general like, enterprise AI applications, the most forward looking ones are seeing some of the challenges on the left where they're looking at things like quality cost latency. A lot of enterprises, frankly, are are are are going to run into those challenges once they're able to false solve some of the initial data access challenges. And to me, I think, you know, one of the things that's particularly compelling is the ability for us to be able to solve this as, like, a cohesive solution at the end for our enterprise customers. Yeah. I couldn't agree more. There was a question around as a, for Rubrik customers, what does this partnership presently mean? And I think the great news here is that we are trying to holistically solve challenges that organizations are facing in terms of going from prototype to production with GenAI applications. So if you're an active Rubrik customer and you're wherever you are in your journey of trying to leverage AI internally, we would love to partner with you and help you accelerate that journey. In addition, I think some of the intangibles, that you'll see from this partnership is with this amazing team of people that, Devvret has built, within the company. Also, some of the ways that we're embedding generative AI within our products, will accelerate, for things like Ruby, our Gen AI assistant for Rubrik Security Cloud. And in other areas, we'll look for opportunities where we can really bring productivity gains to you as you're operating the Rubrik platform. Exactly. I think the biggest thing I would think about for any current, Rubrik customer in terms of what the partnership means is we'd love to talk to you and figure out how we might be able to help solve the problems. Amazing. So, we talked about a few customer use cases with Marsh McLennan and and Checker. I was wondering if there were some other use cases that you could share from customers that really help, showcase what are the like, what is the real life, like, in real world impact that, Predibase is having for large enterprises? Yeah. Absolutely. I think, you know, Marsh is a great example of an agent to get application in a large enterprise, and and Checkr as well as one that we shared. One other that I think, is a really great example for the kinds of real world impact that we're able to see is, actually a group of folks that, we are fortunate to have on stage giving a talk, at one of our predefined conferences in just the last couple of weeks, which is, Tinder part of Match Group. And so for folks that are probably familiar, Tinder's, you know, an online dating service where, you are they're looking to be able to do profile matching in particular at scale. One of the really interesting use cases that they have, they have a number of interesting use cases, you know, within trust and safety, which is obviously top of mind for a lot of folks that are on service. But also in terms of just being able to provide the best possible, service for what they're looking to be able to do. So, you know, we talked about how Marsh has an agent AI system where they're looking at being able to do, like, fulfillment of actions, checkers doing document classification extraction. Tinder, I think, does something, just as interesting as all of these, which is they provide natural language compatibility summaries, for a lot of, for what they're looking to be able to do powered by generative AI solutions. Now one of the things that I think was really fascinating with, this particular use case, which, again, you know, we're really fortunate that they were able to, give a talk to be able to detail exactly the challenges that went into building this, is the incredible scale that they need to be able to access and be able to run on. And the need to be able to do this privately inside of, you know, the four walls of their virtual private cloud inside of their instances in AWS. Now when you look to be able to solve this challenge, you run into a number of different issues. The first is being able to really optimize out on deployments to minimize cost, improve latency is actually quite a challenging problem if you're looking to be able to solve it yourself. But even beyond that, you know, the need to be able to solve all this high volume production use cases also means you need a production serving platform. You know, at this point, it's not okay if your model goes down for a few hours in the day. It's not okay if your auto scaling is blocked as well based on the kind of the input of the traffic. And so what we actually tend to see with these types of customers is they're very serious about the production needs for their application, and the ability for that to be something that they can actually rely on from an infrastructure standpoint. So to I I think they're a great example of maybe, like, one additional customer use case that I'd love to be able to highlight just, here for us. And, really, I think one of my, favorite examples of being able to operate at scale and needing a trusted platform that can actually operate at that scale. For anyone that I think I know that, customer use cases tends to be really top of mind. And for anyone that I think is, curious about what are the other types of workflows that people have built on top of the platform and how exactly did they do it, we're fortunate to work with a really positive and strong set of customers that not only have brailed and tested our technology with some of the most challenging use case that exist, but have actually gone on record and maybe spoken a little bit about it as well. And so you can actually see a lot of these on prediabase.com/customers. You can see a number of different use cases from different customers that exist. Oftentimes, we have videos, with the customer themselves telling the story as well as, you know, the actual case studies themselves. And so it'll give you a nice look at different use cases going everything from high throughput volume and applications to, cogeneration, to marginal finance, agentic AI system, and others. This is great. One question that comes from this is, I think there's a lot of common misconceptions that you need a massive amount of data to do fine tuning. So can you demystify for the group, like, how much data do you really need to be able to get started with fine tuning? It's a really good question. And I'll just start off by saying that I think data is the the hardest part of fine tuning today. And one thing that I think is really, encouraging, though, is that when I started in machine learning, you know, maybe over a decade ago, one of the most common things people said is 90% of machine learning is, cleaning up data, and the other 10% is complaining about cleaning up data. So it really used to be about this entire data challenge and bottleneck. But, it's become something that's become a lot more accessible today, and I'll I'll I'll give you three reasons. The first is just to give you a direct answer on, like, how much data is required. The correct answer is it depends, but no customer ever likes when I tell them that. And so what it'll actually suggest is that, we see people get started with fine tuning with anything from order of hundreds of rows or examples of data, but where we see the best results is typically thousands. It's kinda like, you know, the ballpark you can think about. But there's there's two things that I think made this, in particular really a lot a lot more accessible from a data side than maybe traditional model training had been a decade ago. The first is we see, model distillation workflows become a lot more popular and common. So model distillation is basically the act of taking a larger model to be able to do a task and then teach that towards a smaller model. So I mentioned how I think a little over a year ago when we were speaking with Marsh, one of the things that resonated was GPT four level quality at GPT 3.5 prices. And, you know, we see the same type of workflow exists with open source models as well, which is take the largest model, something you'd never put into high volume production, but use it to label data to create a dataset that you can then tune and create a smaller model with. And so that tends to be one thing that helps solve a lot of the data bottlenecks and challenges that customers might have. And then the second thing, is, a feature that we launched, earlier this year called reinforcement fine tuning, which also reduces the, data, the data load that's needed to get started. With reinforcement fine tuning, you can get started with as few as a dozen examples. And instead of having more data, what you do is you actually write, what are called reward functions generally refer to as Rubrik, no pun intended, which essentially, just help you score the model's outputs. And so you can say, you know, give it three plus points if it's the if the code that the model generated compiles and subtract two points if it uses too many import statements as, you know, a one example. And so, that's just a little bit in terms of the guidance towards data. As it's been true kind of the history of machine learning, quality trumps, quantity of data, you know, pretty consistently. And so I think that one of the things that you often wanna think about is getting access to the best or high quality data that you might be able to do because it helps really drive the performance of your model. That makes a ton of sense. The old moniker of a garbage in garbage out still very much applies here. Yeah. So I wanted to wrap up with, a more of, like, a bit of a fun question and thinking about the future. What trend are you most excited to use AI for in your own life, personal or professional, in the coming twelve months? Yeah. It's a great question. I think that let me let me start just one one step before that, which is, like, what is the trend that I'm excited about? And then I'll I'll talk a little bit about, like, where I'm excited about that from from, like, my life, as well. So from a trend in AI over the next twelve months, the thing that I'm most excited about I have the first thing I'll say is, like, if you'd asked me this question two years ago, I probably said I'm excited about the rise of, like, open source models because I see a lot of the goodness that would exist in terms of ownership cost, latency quality, but, like, it's you know, frankly, it wasn't there a few years ago. I think it is there now. So I'll move on to, you know, maybe the next, trend that I kind of see, which is right now, a lot of models and workflows tend to be, what I might consider one shot. And I'm overloading that term, but what I mean by one shot is you prompt them all, it gives you an answer. But where I see a lot of the future of this development is an area where you actually have a continuous improvement loop that exists with your models. And I think that this is hard to build and orchestrate today, but the best and coming edge companies are starting to do it. And I think in a few years, we'll be a lot more ubiquitous. And when I talk about, like, these continuous improvement loops, I think what I'm really referring to is this idea that these models are constantly taking in inputs and getting outputs. And, what we'd like to be able to see is these models actually learn how to improve the outputs that they're given over time. And this can be done in a few different ways. The first that can be done is, like, you collect some of the inputs and outputs and you, you know, as a human, tell the model what were the actual responses that were good or bad, and then you have the model adjust. But we're actually also increasingly seeing areas where LMs can self critique and act as judges where they're able to improve on those modeling workflows directly too. And so that's the real trend, that I think I can see in terms of, like, where, you know, or or that I'm most excited about in terms of, like, coming forward in the next twelve months is, like, this idea of the flywheel, essentially fact. And then I think in terms of, you know, what am I personally just most excited about using in, in in personal and professional life, I think that, you know, one of the things that, ends up, I think, taking a lot of time for me is just doing, like like, both planning exercises, because there is oftentimes a lot of deep research that's required. Think about planning a trip as an example. And whereas I don't necessarily want the AI to book the trip for me, I think the fact that I can essentially get access to 10 or a 100, you know, the equivalent of 10 or a 100 people fanning out and searching all the options. I'm somebody who likes to optimize and make sure I've, like, you know, looked under every single lock, which takes a lot of time. So to me, I think one of the things that I'm really excited about is, like, these some of these planning agent workflows where you give it a task. It breaks down whether that's, you know, a trip or or a birthday or an event, and it actually goes ahead and figures out, here's the 10 different options that you'd like are gonna be most suited for you in lodging and then travel and others, and then it helps you kind of fulfill that. That to me is probably the, kind of use case that I'm looking forward to being able to use personally. I think I would, I would love to use that planning AI as well. My my answer to that question is, I feel like there's just so many logistics and, like, planning, but just also just tasks that I'm doing on a daily basis. I'm sure a lot of people on this call can relate. It's like we're working full time jobs that are pretty intense. We have families. Like, I have a three year old at home. Like, I there's just so much stuff going on, and I just want, like, a agentic assistant that's gonna do a lot of my logistics from creating my grocery list to meal planning to travel planning to scheduling doctor's appointments and and doing that all for me with the preferences, that I have. And that world will will be amazing because I'll be able to spend so much more time on this, like, either with family or doing things that are actually, like, very impactful, and thoughtful versus just managing some justice. You know what's amazing, Anneka? It's like, I worked on the Google Assistant back in 2017. And so, you know, it feels almost like Asian history now, but that was very much the vision that we had at the time was what if you had an assistant that could create your grocery list to help you recommend what you make for dinner, you know, coordinate the pickups that you might need for your kids from school or the after school activities. And it was like this amazing vision that I think we had great videos for. And then the technology at the time was really just, you know, basically kind of a lookup and then an answer, in some ways, but I was never able to do kind of more sophisticated thinking. So, I think it's amazing that now we actually exist in a place where the tech is caught up to the vision. And I think we're gonna see a lot of these really incredible workflows get built in the next twelve to twenty four months. Yeah. I'm incredibly optimistic about it. I think it's gonna be life changing for everyone, not just people that work in in technology. I'm gonna have a couple more questions and then wrap us up. So there was a question. Do you think the quality of data warehouses will diminish with the errors introduced by AgenTek AI, or is it improving in the early adopters? The quality of data warehouses. So I think one of the things that you might think about with the quality of data warehouses or at least if I were interpreting this question, one of the things that I think about is like, hey. A lot of the data warehouse customers are querying them quite consistently. And with those queries, you're both doing data engineering and pipelining tasks. So oftentimes, you might be just building a report. And, these reports tend to be something that can be a little bit fragile because either maybe you wrote the query wrong or something changed in your upstream tables that you didn't even know. So the same table you ran for your report every month, all of a sudden the report's broken. I'm sure everyone on this call has seen this at some point or the other. I think we're gonna go through, like, a, like, a a parabola shaped journey here where I think, like, the very first thing that's gonna happen, frankly, is that quality will go down for a little pea short period of time. And that's because we're gonna, like, build and release some of these agents. And those agents are gonna get a lot of things right, and they're gonna get, you know, a decent chunk of things wrong initially. But I think where we're gonna see is, like, we're gonna use where those things get, get long to be able to actually adjust course from those models. And I think from there, we're gonna see actually the agent systems develop a lot better things than kind of the duct tape and wire that I often see, put into these data engineering pipelines today. So imagine, like, the ability to automatically remediate some of these types of pipelining issues that have to report upstream. I think that's definitely something we're gonna see, over the next year. Amazing. Okay. So there's a question I think is a good one to wrap up on. I'm not a customer yet, but I wanna explore and test drive Predibase, and the agentic AI capabilities. What do I do? Well, we have an answer for you here. Please shoot an email, to Louie Ferreira on our team, and we'll, he will respond and get you started. And we can explore together, whether what what you need to solve for your use case, and we'll be excited to have that, consultative journey with you. Devvret, thank you again so much for having this conversation. I've spent so many hours with you, and yet every time we chat, I feel like I learn more and more. So I hope everyone here also, learned a lot from, from what he shared. We also, look again in the docs tab if you want references to some of the materials we mentioned. There was also many in the chat, and we can share that as a follow-up as well. And, again, if you just wanna get started or even just have a conversation about where you are in your AI journey, even if you don't know where to get started, please shoot us an email, and we would love to have that conversation with you. Thank you all so much for joining us today. Thank you, Monica, for the great questions, and thanks everyone to the audience. Look forward to working with you all..