00:01
Here we are. Good morning, good afternoon, everybody. Thank you for tuning in to another Ask Us Everything. I'm Don Poorman, Technical Evangelist for Pure Storage. You've seen me as the host on this thing now for a while. I think we're up to our sixth or seventh episode.
00:18
Time flies when you're having fun. Future note for anybody who's interested, these Ask Us Everythings in about a month are gonna go to about two per month, because had such great uptake from everybody participating, and really, really driving us to bring a lot more material through this forum to you guys.
00:37
So, another quick note, if you notice dressed, I'm gonna blow the internet up now, 'cause I don't know if you can see it, but we gave these Lego sets out last year. It's kinda hard to see, but we gave Lego sets out, and we got to dress like the Legos in the set, so it was kinda fun to get dressed up for that. But anyway, thank you again for joining us.
00:57
Today we're gonna talk about everything object at Pure Storage. Now, before you say, "Oh, Pure does object?" Yes, we do. And the good part is I'm joined today by some real powerhouses when it comes to unstructured data, specifically object. I've got Karthik on one side, who's of Product Management for FlashBlade and
01:17
Object here at Pure, and I also have one of the smartest guys I know, Justin Emerson, Field Solutions Architect, on today. And these guys are going to blow your mind on where we're going with object and what our vision is for it. There is a huge, huge future for object, and it's important that we as Pure get it
01:39
right in the sense of how we know it's going to go down. So, guys, without further ado, Karthik, I'm gonna turn it over to you. Start, introduce yourself and, and get this ball rolling. Thanks, Don. Hello everybody. My name is Karthik Srinivasan.
01:53
I'm a Director of Product Management. I look at FlashBlade growth as well as object storage strategy. Great to be here talking about, what object storage, work we do at Pure, why we do the things we do very uniquely, where we see the world evolving. So look forward to a great conversation.
02:12
Keep the questions, coming in, and we'll, we'll answer everything that we can. Justin, over to you. Thanks, Karthik. Hi, I'm Justin Emerson. I've been here at Pure for almost six years. I'm a Principal Field Solutions Architect focused on unstructured data, so that's file
02:27
and object. But, over, over the last year, a lot of my time is spent talking to customers about, object storage and, and our, play in that space. Awesome. Guys, again, thank you so much for joining me today.
02:40
This is such a hot topic for us because Not just because of unstructured data growth, but, but how it's growing, where it's growing, and why it's growing, right? And, and Justin, you have done this before, and I've always enjoyed the you to step us through the, the nature of object, why it exists, why it's the future. Yeah. Because it's been around for a while, but I
03:02
don't think anybody has been able to tell a story like you do as it relates to why its relevance is so important. Yeah. So I think what Before we talk about, like, what object storage is, it's important to talk about what why it was created and why it's different from file, because a lot of people lump file and object storage
03:20
somewhat, you know, rightfully, as unstructured data, but, file and object are different things. And I think to understand why object is different, we have to talk about sort of where file came from so that we can understand why object exists at all. And if you think about it, file, files and folders, and the desktop, and the inbox, and the outbox, these sort of like computer
03:41
concepts, they're all analogs or, or digital analogies for these physical pieces of the 1980s office desk worker, and those don't exist anymore in the modern context, right? The idea of, you know, I've got a, a desk with a desktop, and an inbox and an outbox, and a f- and a file, files and folders in a drawer. Th- that's like a, a, a relic of a bygone era, right?
04:08
So when we think about file storage, right, it's based around paradigms that have their origins in the 1980s, right? File Like, the NFS protocol and the SMB protocol are about as old as I am. Mm-hmm. And I'm getting, getting a lot more gray over here if it puts it, puts it in the right time, time period.
04:26
So, what object storage as sort of a concept was decided to w- what was, was built to do was to solve for some of the low-level assumptions and promises that file storage, specifically stuff like POSIX makes. Yeah. Yeah. Where, you know, early on when we were still figuring out how to do storage and how to make storage more user-friendly, where it's
04:54
like, hey, instead of, like, this, you know, part on the tape to this other part on the tape or this sector to this sector, let's put some sort of, like, directory somewhere so that you know, like, what data this is, and give it, like, a name. So that's where, like, file storage has its origins. But these standards and sort of general, tenets of file storage really
05:20
became incompatible with the kind of scale and distribution or, or cloud kind of, scale over multiple data centers and multiple availability zones and multiple regions. A lot of the things that file promises, things like sort of instantaneous consistency, the ability to do, like, locking between multiple people, having operations that can do to multiple things at once that are not
05:49
those kinds of things are somewhat challenging when you're trying to build, say, for example, a storage service, maybe like a simple storage service, that works across multiple, you know, virtualized y- sort of, fil-You know, disintermediated, cloud storage. And so object storage predates the cloud, but the cloud is really what accelerated the
06:20
adoption of object storage. And so, like, object is a fundamentally different thing. It's still unstructured data, but the core problem it was trying to solve was the kinds of promises that file made were incompatible with the kind of scale and, and, and cloud-like experience that, you know, Hyperscalers wanted to deliver.
06:40
And so the de facto standard for that became, what, AWS built around their simple storage service, also called S3. Yeah, and I think you bring up a very good point, right? And we've talked about this, and we talk it all the time at Pure, that when you look at legacy infrastructure, right, you look at legacy storage and things like that, there
07:02
are lots of attempts in that space to try and continue to make it relevant despite the nature of storage changing so much in recent years, that i- it's actually better to just say, "Look, we need to evolve to a standard that matches what is needed today"- Yeah as opposed to force fitting. It's not to say that file shares are gonna die, and they're, they're, they're
07:24
is the new thing you've gotta do, but there will be a point, right, that there's an inflection point where you say, "Is the juice worth the squeeze still for the way that I'm trying to make this work, and I just need to change my mentality?" Right. Some of the low-level promises that file get more and more difficult to deliver on as your scale increases, as the, the, the level of distribution increases.
07:47
And so what you're seeing in the industry is more and more applications transitioning from, say, I want a file share, to I want an object bucket. Great example of this is in the data protection space. Yeah. In the cyber resilience space, where most of, all, you know, all of our partners in the data protection space, they're all, pushing very
08:10
heavily on object storage targets as go for their, where, where they're their, their data protection copies, and that's because it solves all sorts of scale problems for them, it solves permission problems. It, it makes the whole operation easier to configure, and it gives those applications more, control over what happens because object protocols
08:35
are, much more feature-full than file protocols. Yeah. And a great example of that is versioning. Object storage has a, a concept of versioning built in. File doesn't. You just make another file and you add V4.- Right dot date dot final dot- And it's on you final dot actually final dot BTC.
08:51
It, it's on you to, it's on you to manage the metadata, right, yeah. Exactly. Right. Yeah. And so, like, there's a- all this extra fidelity that you get in terms of metadata and, and, and, other things in, in object storage that are really useful to apps. And so I agree with you. I don't think file is gonna disappear, but I do think that over time we're seeing this more
09:11
and more, is that as applications are refactored, as applications are, made more cloud-native, they're all trying to think about, how can I leverage object storage both to make my application simpler, and fundamentally, how If, my customer is running that application in the cloud, how do I make it less expensive? Because object storage is one of the, the cheaper, storage tiers that you can have in,
09:34
in the public cloud. Where- Yeah where people run into challenges is that for a lot of, storage has not been very performant. Right. And that's changing very quickly, and I think that's where, one of the areas where Pure's been, at the, the vanguard of, delivering
09:49
high-performance object. Yeah, and Karthik, I'm sure you're gonna dig into that one, obviously. Karthik, can you throw up the slide that, that gives the landscape of the object? And while you do that, Justin, you know, for people out there who've not heard your metaphor before, give us the metaphor of the parking lot and- Yeah and, you know, it gives
10:10
a very good graphical view or visual of what the main difference is, and people It's an aha moment almost. Yeah. So when I, when I talk about the difference between block, file, and object storage, they're all sort of different levels of abstraction away from the actual data, right? So if you think about a plot of land, that's kind of like block storage.
10:31
I'm giving somebody a volume. The only thing that I as a storage platform are, is responsible for is metaphorically, a sinkhole doesn't show up and suck up whatever anyone's built on there. Yeah. But everything that gets built on top of that plot of land is up to, at that point, what I'm serving it out to.
10:49
So I'm gonna send it to a server, and that's gonna format it with the file system, that's gonna run applications on it, whatever. So they're gonna build something on that plot of land. File storage is kind of like paving it and putting a parking lot there. And if you If it's a big one, it's a parking structure with multiple floors.
11:06
But you go in there, you park your car, you have to remember what floor is it on. That's like, what directory is it in. You gotta remember what your space number is. That's what the name is. And there's only so many ways in and out of the parking lot, and so you're
11:22
of the parking lot, you're providing, you know, a, a higher level service. But now you're responsible for arbitration, like things going in and out. You're responsible for more things in the stack. Object storage kind of takes that to the next level, which is, you're in one of those newfangled, robot parking structures, where you drive your car onto a sled, it disappears
11:45
inside this big, building, and you get a little valet tag. And when you want your car back, you go put it in. That's kind of what object storage does. It's a completely flat namespace. There's no organizing of things inside there.
11:57
You can do things with prefixes, but just fundamentally, like, part of the name of the object. So you don't have a situation where it's like, oh, I've got, like, subfolders with their own metadata. It's like, it's one big flat namespace. And all operations are atomic, which is really important.
12:13
It means that when I do one thing, I'm one thing to one object, and that operation is to the whole object. So I can either overwrite the whole object, I can write a new version of the object, I can write a completely different object, but I don't have things where I'm, like, doing part of something and not something else.
12:30
Um-And objects themselves, like, don't get edited. They only, you know, get new, new objects or new versions of objects. And so these changes in how we promise how storage works is one of the things that unlocks that scale. It allows us to make it more portable between all sorts of different things.
12:50
File migrations, as I've talked to many customers, are terribly painful. Yeah. Yep. But object migration's far less so because object is really just, you know, it's a, it's an HTTP call. It's basically a web browser- Mm-hmm getting a bunch of data. And there's so much you can do in between those, that, that client and that server in
13:08
that, in that aspect, that it makes it much easier to do things like migrations. Yeah, I've always wondered about that because the nature of object, right, is that the, y- you know, when you look at file, right, it's, it's the data blocks and then it's the metadata, right? The metadata's being managed by a directory or whatever.
13:25
And but object, my understanding is the attributes of the data itself are contained with the object, right? So there's no dependency- Right on some central metadata service to make sure that the security is applied and all of that. Is that, is that a fair assumption with that portability comment?
13:41
Yeah. Like, like, an object itself carries whole bunch of other attributes besides just the payload. In a, in a file- Right context you have the file name, you have its path, have things like ACLs or, or attributes, but those are extremely limited. Whereas in object you can have an arbitrary number of custom metadata fields, and some of
14:01
those metadata data fields are things checksum, which are enormously important for data integrity amongst- Absolutely you know, other things. Yeah. So that's kind of where we are and, and I think object storage really is the, the fastest growing part of, the unstructured data landscape as more and more applications move towards it, as it gets higher and higher
14:21
performance, we're going to see more things do it. And I think that's probably a good opportunity for Karthik, for you to Pure's vision for object storage going forward and how are we playing in this Before you do, Karthik, before you do, though, sorry to cut you off. Before we get there, Justin, you and I need to talk about the two bullets on
14:43
unlimited scaling and metadata, 'cause you and I were close to the, the big mark with FlashBlade. I'll let you Flex on it, though, with the numbers we did produce, because that is a significant number when you talk scale. Yeah. Yeah, like one of the things that's really hard to do is, you know, have a file or a folder, or sorry, a folder with billions of files, but we can build object buckets with
15:04
trillions of objects, right? So this says unlimited scaling, billions of objects. We, we ran a test and there was a blog that, that, was written, late last year about how we did this test, but we were able to you know, 3.8 trillion objects in a bucket- Yeah. Yeah for as a, as part of a test for a
15:24
customer as a proof point of, yeah, we didn't have an object scaling limit in our FlashBlade platform. And the only reason it wasn't more than 3.8 is that we ran out of time and had to stop, but it could've kept going. Yeah. And by the way, the guy that did the testing just chimed in on chat.
15:38
Hey, Russell, how are you, man? Best dressed guy up here, by the way, and did the three trillion thing, so. Yeah. Yeah, that's, that's an idea of the scale, right? Three trillion objects.
15:48
You just have to say that and sit back and, and we had to stop 'cause they needed the FlashBlade for something else. All right, Karthik, we sidestepped there, I apologize, but it's now your turn really to get up here and give us an idea of what the object vision is at Pure, because it obviously supports not just our ability to evolve as the times evolve, but it definitely
16:09
Justin's point which is it's gotta be the fastest growing form of unstructured data out there, so our vision matters. So it's up to you now, Karthik. Hey, thanks. Thanks, Don. Great points by Justin.
16:21
I, I think, I do wanna just, touch upon a couple of things and then I will switch to what Pure is, doing from a longer term vision perspective, right? I think the unlimited scaling part of it, especially when it comes to on-premises or, you know, non-cloud, deployments, is not a given, right? If you think about it, the cloud, can do a lot of things because it's all under the covers.
16:47
You don't get to see it. When you build, say, an on-premises appliance-based, you know, capability or even a software-defined capability, the scaling portions of it have to be architected in. You can't actually just guarantee that hey, scale will just automatically happen just because it's object storage, right?
17:08
We know of many, many, challenges with, with, with, with non-Pure, environments where they'll have restrictions on number of objects in a bucket and so on and so forth. I, I just wanna just make sure that people understand that, that the three trillion objects, along with it, you know, with, very little performance, you know, impact, needs to be kind of built in into the architecture for us to be able to serve these
17:32
massive scales. I think there were questions on AI/ML and immutability and so on, and we'll get to that. But, but really, we have to consider so many different things as part of kind of building the, the capabilities, to offer not just, just seamless portability and, and native S3 support, but also making sure that it's built for workloads of the future, whether it's AI,
17:56
cyber resilience, all of those things. We can always come back to these questions and, and have a discussion, but let me just tell you a little bit about where Pure is going, right? What is our vision, right? I mean, Justin spoke about, you know, what are the great things about object storage and why
18:11
it makes it so, unique, in, in today's, time because that's the fastest growing, you know, data footprint that we see in enterprise environments. Let me see if I can actually Are you able to see my screen? Yeah. So, so now this is our vision. Our vision really is to ensure that-When we build these modern applications, AI/ML
18:34
environments, right? They're able to harness the data seamlessly wherever it is, right? Whether it's at, your data center core, which means providing you with very massive scale, highly resilient environments. You know, in many environments, what we have seen is, customers wanting to mimic what AWS
18:53
does with regions and availability zones. So that's what we're building towards. We have most of the pieces for that puzzle already in, in place. But data is just not restricted to data centers alone, right? You know, data is now growing everywhere, whether it's, you know,
19:08
outside the data centers. Data centers themselves are becoming more, I would say, disaggregated simply because of power and space constraints, so they're also pushing out. So, you know, you have these distributed pools of data that need to be orchestrated, right? A-and so we are-- we introduced, Flash you know, object storage on our other great
19:28
platform, which is FlashArray, which is a unified, block file and object storage. The best block storage platform, the best unified platform, and now we're introducing object on, on FlashArray as well. And, and the reason for that is, in many environments, what we've noticed is that, hey, you have power and space constraints, you wanna consolidate workloads.
19:50
You don't need as much performance as you'd need in your core environments. Perhaps you're just using smaller apps, maybe these are container-based apps, small image repositories, you know, test and dev environments. You know, so the FlashArray from a profile perspective, from a performance perspective, and a capacity perspective is perfect for that, right?
20:12
And so, we've introduced FlashArray Object. You know, it's the 1.0 product that's just coming out that's gonna be GA, in a couple of months. So that will serve kinda like what we call the near edge needs. And then the FlashBlade, which is our massive, you know, environment, will
20:29
the core needs with, with, with the highly resilient multizone, multi, multiregion capability and, and various versions of, of resiliency. You know, we're building things like strong consistency using synchronous replication. So for in-region, highly available and strongly consistent semantics. And of course, you know, you know, to other regions, from an async
20:56
that you, you can do like multiregion environments. Now, the last piece of the puzzle is, of course, the cloud. And cloud is a little bit in the future, but really the way we are thinking about is how do you ensure that, we can seamlessly operate across this edge core cloud paradigm, right? How do you work with, local buckets in AWS S3, which means we need to
21:18
kinda control the metadata. How do you ensure that that is, you know, is, that people or at least the, the core data centers and the edge data centers can seamlessly talk across. Those are things that are still being worked on, but this is really our, our vision. Wherever there is data, wherever there is, object data, we wanna be present there.
21:37
And you know, we've start-- we started with the core. We started with in-region, multiregion semantics. Now we're tackling the edge, and then pretty soon we'll be tackling the cloud. I'll pause here. I'm happy to take any questions. Happy to double-click and DeepReduce dive into, specific areas. Yeah.
21:54
Let's, let's dig into some of the Q&A that's out there, obviously, because you, you did mention S3, so that's obviously the API that, that we're using, which would make complete sense. But going backwards into what you talked earlier, optimizing for AI training and what-- at the core, because we would be talking about FlashBlade object if we're
22:13
talking anything AI because of scale out, what, you know, what are the advantages there? Cause I know we have RDMA and some other things, if you could talk into that a little bit. Yeah. Totally. Let's talk about the AI environment itself, right?
22:28
And Justin, feel free to jump in, right. You know, we think about it in two, two ways. One is there's of course, training workloads, and then there's gonna be both of these have very different IO profiles. Training workloads, you know, especially when you're doing checkpointing, you know, lots of
22:46
GPUs writing. If there's bursty massive amounts of data, requires a lot of throughput, right? Inferencing is like, hey, you're, you're doing smaller pieces of data. It's, it's, you know, while, while training is largely from a, a, from a checkpointing perspective, heavily on the write side, inferencing, on the flip side is, is heavily
23:06
on the read side, right? But much smaller amounts of data latency matters because you wanna kinda serve these, responses very quickly, whatever the application may be. Yeah. So, let's take that, right? Massive amounts of throughput required to kinda serve that.
23:23
FlashBlade is perfect for that. FlashBlade comes in multiple flavors, as you-- as many of you know. There's of course the FlashBlade//S platform, which comes in two flavors, which is the S200 and the S500. The S500 is our workhorse for really massive, highly, parallel environments,
23:42
heavy metadata, IOPS as well, right? So that's the, that's the, that's the FlashBlade//S. And of course, last year we introduced FlashBlade//EXA, which is a disaggregated platform that is meant for, you know, these AI/ML types of workloads, right? FlashBlade was built for throughput.
24:00
It's got a bladed architecture with the FlashBlade//EXA. Now the metadata is kinda disaggregated from the data nodes themselves, so now it allows for massive parallelism, and direct access to the data nodes through things pNFS and RDMA, whether it's S3 over RDMA or NFS over RDMA. All of those things enable you to have massive throughput, you know, directly to the data
24:22
nodes, on the X side, and then even on FlashBlade//S, you know, because with bladed architecture, you can actually have massive throughput. So it's great for, you know, you know, training types of workloads. You know, the way we segregate it is like, hey, up to about thousand GPUs, FlashBlade//S will do fine.
24:41
Greater than that, FlashBlade X is your product, right? So that's, that's on kinda like the training environments. And then when it comes to, inferencing, it's, it's mostly around read and, and, you know, low latency.The FlashBlade stack is super optimized for IO, right? Up and down the chain, right? You know, th- that's because we control
24:59
everything from, when the IO lands all the way down to FlashArray, right? Because we, we, we own the Flash modules, we build the Flash modules ourselves, directly from NAND. We have o- optimized the IO path, so, it's in significantly lower latency compared to anything else that is out there, and that is great for inference types of environments, right?
25:22
And so that's how the f- the architecture really, serves AI/ML workloads. Justin, you wanna add anything to that? Yeah, I think sort of just in general, I, you know, I find the, the AI space really, really fascinating here because, you know, I, I, I was working, you know, wi- with, with, you know, trying to help customers with DGXs and stuff back in 2018.
25:43
And what I sort of observed happened was the when AI workloads were just sort of starting, everybody was just using file kind of as the de facto. It just sort of file ended up accidentally being the thing that AI kind of got, AI workloads kind of got built on, and it wasn't because of, like, any good reason. It was just sort of like, well, I started with one system and I put all the things in a
26:07
folder, and then I needed to find a shared folder because now I have two DGXs. Yep. And as more and more time is And, and by the way, that was perfectly reasonable to, to think of at the time because most people did not consider object to be a high performance protocol at all. That's changed, and so you're starting to see more and more momentum in the industry around
26:26
a push towards object for AI workloads for all the reasons that Karthik talked about. But, you know, if you're thinking about huge training data sets, or you're thinking about, from an inference standpoint, right, at, you know, recording every single i- i- inference response that you've ever made, that's all about scale, and being able to do things without having to run into the same limitations of, of files and, and folders and,
26:51
you know, directory limits, is pretty appealing in these spaces. It's just that there's had to be some underlying infrastructure plumbing work done in order to bring object up to sort of the same potential, that file had. Because file's been, you know, in the high performance HPC space, things like parallel file systems have been in place for, you know, 40 plus years. Yeah.
27:13
And so, it's just now that I think the, the industry broadly is finding, ways to bring object storage up to a similar level of performance and capability so that AI workloads can then also take advantage of that scale piece. Yeah, and a question for you about that. So, and, and this came in on the chat.
27:32
So just for our edification, what object protocols are we supporting today on the different platforms? You know, RDMA, all that stuff. What's, what's the rundown on that? Right. So, we support, S3 as the data protocol. So, uh- Yeah.
27:49
And we have a fairly extensive, highly compatible S3 implementation, for most of the things that make sense for an on-prem object store to do. Some things don't because they're sort of only relevant in the cloud context. We do almost all of those. Now the, the sort of lower level protocol below that, which is sort of like the
28:12
transport protocol, I think, which is what you're getting at- Yeah most, most, S3 customers are using S3 with TLS, which is, which is encryption. And so for example, all of our sizing tools and various stuff assume encrypted workloads, because we find the majority of customers are using that. But Karthik, recently we, released, S3 over RDMA as well, and that was very
28:35
to support some of these AI workloads. Can you talk more about how that functions and, and why that's cool? Yeah. So, so let's talk through that, right? I mean, S3 over RDMA, was a joint effort between Pure and, some of our GPU community, NVIDIA specifically.
28:52
It's not part of the traditional S3 spec, in, in the sense that hey, you know, it's not been, you know AWS has not made, made that change. What we did was we worked with, NVIDIA to try and, you know, ensure that, when you have these training instances, you could go through, a specialized client, which has the NVIDIA libraries, and then go to kind of like the data nodes, right?
29:21
So the way it works is, you know, in traditional IO path, you, you know, the, the, the first call would go to the metadata layer, which says, "Okay, hey, where is my data? Here's my authentication, authorization," of those things get taken care of. And then it says, "Okay, get me the data," and it goes through the same IO path, right?
29:42
What we have done is with S3 over RDMA from a transport standpoint is the data disaggregated and separated out from the metadata path. The initial authentication and still goes to the metadata layers, but then what we do simply is just say, "Hey, open up a channel, an RDMA channel," that says, "Okay, now you have acc- direct access to the data.
30:03
Just take this data and push it into your memory, on the GPU side, you know, bypassing a whole bunch of, you know, you know, you know, compute related layers and layers." That way you have, you know, the access is much quicker, a lot more throughput. And so we have seen massive improvements in basically the throughput itself. You know, we can, we can go up to 300 GB/s on a five chassis FlashBlade system.
30:31
This is still very early stage. We've just launched it. It We just launched it in December, and we're starting to work with customers and, and trying to kind of like deploy this in environments. And so there's they're super excited about this thing and, and we will continue to do
30:44
that through the course of this year. That's awesome. Yeah, and y- not, not to throw you a curve ball, Karthik, but we did get some Q&A in related to object on FlashArray. Um- Yep, yep one, one of the questions I thought was really, really interesting is, y- you know, we put object on FlashArray, whereas, you know, we had object on scale out, but with
31:04
FlashArray, you know, we have the-Only two controllers, you know. So if a controller fails in the middle of transactions with objects, so, you know, we're, we're starting to talk about integrity of transactions here. It, it's probably easier handled on the FlashBlade side. What, what does it look like on the FlashArray side?
31:21
Coz obviously that's a big growth point for us, so those kinds of things are gonna come up. Yeah. So I think, you know, if, if, if something fails during, you know, a particular transaction, what'll happen is the second controller will kick in. If there is a timeout, we anticipate that most of the object applications will have some sort
31:41
of timesou- timeout built in, and many of these operations are Let's say it's a read or a write operation, you know, there will be some sort of a retry. We anticipate that the applications will have that kind of retry logic built in. And so that's something that we would, we would anticipate happening, and whether it's on the FlashBlade or a FlashArray.
31:58
The FlashBlade, of course, has multi-blades, and then so other blades can kick in and serve the data, or the serve the IO in any way. The same thing will happen on the FlashArray side as well. Yeah. Justin, feel free to jump in and add anything you want. Yeah. On, on FlashArray, right, the, the, the architecture is that both
32:16
same underlying media, and the same And the underlying media is where all the stuff gets committed to. So if a controller fails, the second controller, which is in, in sort of a secondary or a standby mode, is immediately able to reference all the data that's just been written. So we don't acknowledge the data back to the cl- or we don't acknowledge the, the, the
32:34
completion of a write to the client, whether that's on S3 file or on a, a block storage with, with FlashArray. We don't acknowledge that back to the client until it's been committed in a persistent way to the first, sort of part of the, the data pipeline, which is our NVRAM. And so once it's in NVRAM, both controllers can see it, it's just whichever one's active
32:55
is gonna look at it. Very similar to how it works on, on FlashBlade, just in a distributed way. But in both cases, right, we're not gonna acknowledge that the write is complete to the client until it's been safely persisted, means if the controller dies before the data is persisted, then the client's never receive the acknowledge and it's gonna
33:15
do a retry. And in the case of object, object already rides on top of TCP, which means if you don't get the TCP ACK back, then you're already going to have the logic built into that layer of the stack to do, you know, re know, resets and, and so forth. So, in, in that way, object is actually quite robust, especially also because part of
33:38
committing the object is saying, "Here's the checksum for this payload that I just sent you. Make sure that it matches this when you commit it." and so if something gets corrupted in flight, we actually know that before we ever commit it, and so when we respond back to the client, not at the TCP level, but at the S3 level, we'll say, "We didn't commit this because this data that you sent doesn't match the checksum that you also
34:01
sent, so something happened. Send, send it to me again." Yeah. And I saw that, somebody typed in on the chat, they said they force their developers do three retries before they consider it a failure. So- Yeah there's obviously, you know, resiliency built into the
34:14
actual IO process as well. All right, so All right, I'm gonna throw you a curve ball, Justin. Okay. So, and, and this is gonna actually relate to something on the slide that, that's up there already. You know, somebody's asking, you know, "What do you do when I need petabytes of object
34:30
storage, but only TiB of block storage?" So, you know- Yeah obviously he's probably thinking in the vernacular of a FlashArray. But I think there's something to be said here for the idea that, well, now that object is on FlashArray, especially when you're Pure Fusion, there's a bigger picture needed- Yeah to extend your object to petabytes, Pure Fusion and FlashBlade, and there's
34:55
a, there's a magical picture there I think that could be painted, right? Yeah. That, that's where I When I saw that one in the, in the Q&A, that's where I was gonna go with it. The So, like, let's say for example, you start with a really small environment and you've got, you know, a single FlashArray, and you've got your, you know, maybe the block aspects.
35:09
I think the question was specifically in, in relation to JFrog Artifactory, something I'm, I've been, I've There was a JFrog question in there that's- Yeah that's got a similar need, right? Yeah. So, like, you can start with a single FlashArray.
35:21
You can start with, hey, here's our, our block volume for this particular thing. Here's our object store for the sort of bulk data repository. Now over time, so FlashArray is a scale-up platform, meaning, you know, if I start with X20, I can upgrade to an X50, I can upgrade to an X70. I can then go all the way to an XL 170, or maybe it's a FlashArray//C and it can go all
35:41
the way to a C90. But you're gonna hit a limit of scale at some point, right? And either you can make a decision of I'm gonna do, go with two FlashArrays at this point, or you can say, "Hey, maybe what I should do is move this object portion of the workload to a FlashBlade." So in the case of if someone has a small requirement for, for,
36:01
for block and a larger requirement for object, you know, it makes sense to use the right tool for those two different jobs. You know, if I'm a carpenter, I don't only have one saw and try and figure out how I use that one saw for everything necessarily. I'm not a carpenter, but that would be imagined how, how, how it would work.
36:21
And so really it's a question of scale. And, and long term, as, as Karthik's shown here on the screen, our intention is to make it so that moving that data is going to be, is going to be easy between those kinds of platforms, which is how something like Fusion in front of these multiple systems can make it appear like one single, pool of data, one single endpoint for automation, and so that I-
36:46
you know, y- you can look at, hey, how do I use the right tool for the job, but also how do I not add complexity to my environment by doing that? So th- that's how, that's how I'd probably address that.Yeah, it makes sense. I mean, a- and Pure Fusion is kind of the thing that ties it all together, right? If you need petabytes suddenly, you can extend into FlashBlade, and I'm sure there's some
37:09
kind of really good story in there of how Evergreen//One can help you do that too as far as- Absolutely. Yeah changing, right? Like, like if you have, if you have a, a Evergreen//One, contract and you're like, "Hey, I need to do less of this thing but more of this other thing," that's one of the advantages of being in a, in a truly STaaS service model, is, you know, if, if I'm in the
37:30
cloud and I'm like, "Hey, I need less of the block service and more of the object service," that's just, you know, changing your subscription. And Evergreen//One gives you a similar capability for on-prem infrastructure as well. It really opens up that flexibility for you. Yeah, it's almost like a magic knob, right?
37:45
Like, I've always dreamt of that. Like, oh, I need more block today. Oh, I need more- Yeah object today. At least you can just kinda use it as a, as a shock absorber almost. All right, Karthik, let's come back to you real quickly 'cause there have been some
37:59
technical things that came in on the Q&A. First thing to address is the cybersecurity stuff around this, right? Cause object, I assume, can be covered by SafeMode Snapshots, but let's have you talk that through 'cause I know there's some people asking about that. Yeah. So, I mean, as, you know, we don't do really
38:17
snapshots in object storage. You know, objects has versions, you know. This is our version on an individual, you know, individual object you'll have versions. And I think I, I just I do see the question on hey, you know, immutability and, locking and so on and so forth.
38:35
We already support, object lock on FlashBlade. It is already certified with Veeam, as an example. And it's already certified with other, you know, backup vendors as well. And so we already support that immutability. Object, objects themselves are immutable by nature.
38:55
You know, when you write an object, you're always creating a new copy or a new version o- of the object. Let's say you're, you're saying, "Hey, this, this is my dog, Foo," or whatever. Even if you create another image of the same, with the same name, it creates a new version of the object, right? Except it's, it's put on a stack, and what you
39:14
CSI is the latest version that sits on top of the stack. You can lock individual objects. You know, that's what that object lock capability allows you to do. And we support it on FlashBlade, and we will eventually support it on FlashArray as well. When we start out with FlashArray, you know, we'll, we'll just do the basic
39:32
we go GA, but then we will start overlaying other capabilities soon. What about, some of the more core things? Cause I, I saw this on the Q&A, I leaned in, the, the deduplication in large objects. Is you know, 'cause do we do garbage collection? Is it real time? Is it in line?
39:48
You know, how are we doing that? It Yeah. So, let me just, let me talk through a couple of things, when it comes to deduplication and, and all of those things, right? We already support, compression within FlashBlade. Compression is always turned on, right?
40:08
So- Yeah I, I'm, I'm gonna talk about data reduction in general. I'm gonna use compression and, and, oth- techniques to kind of get to that, right? Ultimately, that's what matters. So we already support compression. All of the performance numbers that any Pure person, gives you, you know, for, for you and
40:26
your organization, is all inclusive of all the overhead that compression takes, right? You know, which means that we've accounted for all of that. Let's say we promise you, "Hey, this particular system, this particular configuration can do 100 GB/s," that is including the fact that hey, compression is always running in the background, right?
40:45
So that's step one, right? We're also introducing what's called DeepReduce, which is our version of taking a look at, you know, we don't deduplication because it's not deduplication. It's, it's a, it's a different technique. Deduplication uses exact block, exact block matching.
41:02
We use what is called similarity-based approach. And we've just launched it. In fact, it, it just went live, you know, in the last month. And that allows you to have, additional capabilities. Let's say, for example, you have two video files, right, that come in and they look
41:19
almost the same. You know, there's a little bit of a what will happen is if the block, you know, we will, we will first, you know, break it up into smaller parts and then we will match to see is, hey, are these similar enough? You know, and then so can we now then compress the, the ones that are ve- the same, right? And so these techniques are, what we call, a,
41:42
as part of our DeepReduce that we've just offered. And so that'll enable you to have significant data reduction. What we've seen is, ef- e- effectively, if you have files and, image storage and so on and so forth, you know, we can, we can go up to north of five is to one. In cases of, say, backup, backup environments, you know, 2.5 plus, you
42:04
know, and so on and so forth. So all of those things are accounted for. Hope that answers that question. Yeah. So we're squeezing more out of it, obviously, and DeepReduce is a great example of that. Right. A- and I, I, I think it bears to have a conversation too, because people are still
42:19
asking latency questions and things like that, you know, especially with mixed Right. What, what are your thoughts as it relates to object and how it is eventually, I think it's eventually heading this way, for, like, the Zero Move Tiering thing, right? Cause- Yeah, yeah. Yeah you, there's such value behind Zero Move
42:39
Tiering for a file, it better be there for object as well, right? Totally, totally. Let's talk through that, right? But before we go there, I do wanna just, one more thing to the previous question, right? I think there was a question on whether th- you know, there's gonna be any performance impact on deduplication, right?
42:53
Yes. Yes, there- The way we do our, the way we do our, you know, our, our DeepReduce capability is-Uh, we check for but we do a lot of things post-process, right? So effectively that, the moment the data is written, it gets acknowledged. There's gonna be some li- limited impact on performance, but not a whole lot.
43:11
So it's, it's not a concern at all for us. Let's go Let's take a look at your, your question on, hey, economics and, you know, the, the, the need to use, perhaps multiple storage classes. Like, for example, if you take a look at the cloud, the cloud has multiple storage classes. You have an S3 storage class, you have intelligent tiering storage class, you have
43:30
Glacier, you know, all of those things that enable you to, ensure that, you know, as data ages out, it is kept in, in a location, at the best possible cost, right? We introduced a technology called Zero Move Tiering last year, specifically for the file workloads. Effectively, it, it helps you create storage classes on the same array, where
43:54
not moving the data around because it's all-flash. We are now redirecting the performance for the hottest data, wherever it is. And then when it, when it ages out, we don't give it any performance at all. You know, you know, it's very minimal, so if you get access it, you know, it, it will be served at, you know, at a much lower level, right?
44:13
So that enables you to have two things. One, TCO for your entire environment. Yeah. So it's a great story that way. But also, you know, very low performance impact because you're not, like, constantly moving the data from a disk tier to a flash tier and back and forth.
44:28
And, and you avoid all that messiness that comes with it. Because the data's right there in the media, and all we have to do is just figure out how to redirect the performance to that particular data set the moment you Which gives you a perfect place for placement for talking about, you know, different workloads hitting the same FlashBlade, right?
44:46
It's- Right. Y- y- you don't know what next workload you're gonna bring on, so you need to have a tool like this in your bag to help balance it out to some degree. Absolutely. Absolutely. It enables, you know, FlashBlade, for example, was always built for, you know, of multiple workloads, whether it's file or object. Yeah.
45:04
FlashArray is gonna do the same thing too. So we have block workloads, and, and file workloads, and object workloads as well. I mean, the, the footprints and the performance profiles are gonna be very different. One is gonna be smaller, more scale-up oriented types workloads.
45:18
You know, FlashBlade is gonna be scale out and high throughput, those kind of workloads, right? Yeah. And as Justin mentioned, you know, the, y- you know, with Pure Fusion and the way we orchestrate all of these systems, you know, it makes sense for you to consider both platforms, you know. You pick the platform that works for your
45:35
workload, and Pure Fusion will orchestrate everything, and your workloads will enjoy the best of multiple different, underlying hardware platforms that we provide. Yeah. Or, or, or at a certain point, the, if you give it enough information, we'll, we'll get to a point where Fusion will make that decision for you based upon, w- upon those criteria, right? So rather than having to choose a platform,
45:55
right, you can let the, the intelligent, you know, the, the Enterprise Data Cloud make that decision. Oh, right, yeah. Yeah. Yeah, that's the perfect part about this, right? Is that w- we're not just evolving the service to say, "Me too." We're actually
46:10
in the bigger picture that we see with, you know, the Enterprise Data Cloud experience, the whole platform. Go to one place and have it smart enough to let you to put the data put kind of thing, which is phenomenal. Okay, Justin, I've been waiting to ask this question of you because it I've been watching
46:30
this one for a while. Okay. So we did have somebody come in and say, "How do I mile- migrate from file to object? What do I have to think about?" Yeah. Because obviously y- you and I have talked about this, like, i- it's a completely different paradigm, so- Right you can't take the bones of the old thing
46:48
over with you, right? Yeah. You've got a great story for that. It, it's best to not think about how do I move a bunch of data from file to object, because fundamentally, your data is there because something's using it, right? So it's better to think about this in terms of the application stack.
47:04
So for example, if I'm moving my, data protection system from f- you know, w- today it's writing to a file share and I want to move it to object, then the engine that moves that data is not gonna be most likely an administrator writing a script. Right. It's going to be orchestrated by whatever that application is.
47:25
Because the formats are different, the way that they're accessed are different, so it's not gonna be something as simple as, like, you know, a robocopy from, from file to object, right? It's gonna have to have some kind of, application-driven, process, and many applications that are going through this transition have, you know, are,
47:49
are, are cognizant of that. So, in, in the data protection space, it's, "Hey, I'm gonna do an OX copy of my stuff from this one to this one." Right. In the case of, you know, what we mentioned earlier, JFrog Artifactory, there's a way to convert your file store, which is where all the artifacts are. There's a way to convert that from, from a, a,
48:06
you know, a file-based one to, to an object-based one, and the application very often is the one that's, that's driving that mechanism. And I think that's how it should be because, it, just like it doesn't necessarily make sense to take an on-premises application in a bunch of VMs and then just put it in the cloud in a bunch of VMs- Yeah, sure it doesn't necessarily make sense to just take a bunch of
48:29
data that's sitting on file and put it in object. There needs to be some refactoring that's involved because fundamentally, file and object are different things. If they were the same, we wouldn't need two of them. We, we wouldn't need two different paradigms.
48:42
And so that's why I think that, you know, the migration from file to object is not a storage conversation, it's an application And as more and more applications make that change, like we've seen, you know, a, a major shift in terms of, how things like, data warehouses work with, with, with, with-Uh, technologies like Parquet and Iceberg, we've been moving towards object storage.
49:08
Object storage is, is gonna be powering the next generation of structured data, which is wild. Yeah. Yeah. but, but also, like, for example, you know, when I first came to Pure, one of the things we worked with customers on was how to move their Splunk environments from, you know, a traditional block-based to an
49:27
version called Smart Store. Which was- Smart Store, yeah you know, gener- which was, which was written for the cloud. But it works on-prem, equally well when you have a, a, a suitable object store for, for it to, for it to access. And I think your point is well taken as it relates to the, the growth of object is a lot
49:46
I- is based a lot in machine-produced data, right? Applications doing it. Yeah. Now, you know, let's, let's turn it a little bit though that because humans like their shared folders and all of that stuff. So when do you think you'll see humans interacting with object in ways that they can
50:05
relate to starting to make the shift? Well, I don't necessarily think that you're gonna get to a point where, you know, somebody leverages S3 browser on their desktop in order to pull files- That's what goes to my head, right? Like when I- I don't think that's- When I hear that I don't think that's gonna be the point
50:20
of it. But let me show you where, where object storage would be relevant, right? Instead of, You know, we here at Pure, we don't have, you know, a file share access for our home directory. Right. We have Google Drive.
50:33
Right. Where do you think all the da- application data that Google Drive is storing, where do you think that sits? It sits on object storage- Object fundamentally, right? So there'll be an application like OneDrive or Dropbox or the Google Drive application for your laptop or whatever, and that's gonna be the thing that interacts with object storage.
50:56
So when you think about, "Well, how am I gonna migrate my home directories from file to object?" You're not thinking about that. You, you think, "How do I migrate my file shares to Google Drive?" Google Drive, right. And that is an application migration. I'm migrating from file shares to an app, and I'm not, you know, writing a script to take a
51:16
bunch of my files and push them into GCP objects. Yeah. Yeah. Like, I'm copying them over and letting the application handle that conversion process because the application fundamentally is the one that's interacting with it. Object storage, S- S3 especially, is really designed for machines and programs to, to, to talk to a- and s- and store and retrieve data, whereas files and folders
51:43
are very much a concept that is relevant to humans, right? But, like, I, I, I talk about if you think about, like, the, the f- the file drawer, like, imagine you pulling a file drawer out goes back two miles because it's got a folders in it right? Right, right. That metaphor breaks down pretty quickly, so
52:00
you need to do something else. And the metaphor holds when you're like, you know, you pick up the phone, you call the valet, you're like, "Look, man, I, I need this one thing. Can you get it for me?" And it comes to you in a, in a- Right performant way, right? So, okay, so we're, we're down into our last five minutes here, and we do
52:14
have a few more questions. Karthik, I'm gonna throw them over to you. I don't know if you've been seeing them, but some people- Sure have been asking about page space reclamation in the FlashArray side of object, if you wanna talk about that. And then we can probably tie some things off on the other side with, some final performance thoughts and things along those lines.
52:32
Yeah. So the way the space reclamation on FlashArray will work is very similar to how it does for file systems, and, and file data. So if you're gonna delete an object, garbage collection will kick in at some point in time. And so, you know, that mechanism, because we, we use the underlying substrate very similarly, I mean, it's, these are all
52:52
citizens- Yeah but the underlying substrate is the same, so all of those things work, very similarly. Now, I, you know, you said there were some questions on performance. Would you like me to talk about- Yeah, there, there's one about max la- yeah, about max latency guarantees on FlashArray as well, "cause" we're, we're getting a lot of
53:08
questions related to object on FlashArray, which is neat. Yeah, yeah. So, so just to be clear, right? I mean, like, you know, the early versions of the FlashArray object are meant for smaller applications, you know, low performance type environments to begin with, and then of course, you know, and we will continue to improve that over time, right?
53:27
I think, you know, I, I'll give you a sense of kinda like, you know, the FlashBlade object store is the more mature one and, you know, fine-tuned completely over many, and we have seen even That's, that's not even a built, that's not even a system that's built for latency. That's a system that's built for massive throughput. But even in those environments, we do see low
53:48
double-digit millisecond latency or, you know, single-digit millisecond latency objects, right? This is in actual customer environments, right? I'm just giving you some anecdotal information to give you a sense of the platform. FlashArray, you know, we're still working through it. The product is not GA yet, and we'll share
54:07
more details as we get through, some of the initial rounds, as we get closer. Yeah. Yeah. That's- And I wanted to chime in there as well on the, the reclamation piece. So, so like I Karthik's absolutely right. Like, we have one of the lowest time to first byte, which is kind of how we
54:22
in the object space. Like, latency usually covers the entire transaction, right? So if you're in the object space pushing around objects that are a GB in size, the latency of that transaction is very large. Mm-hmm. But what you're really trying to think about is, how quickly do I start
54:38
getting that information? And that's what time to first byte means, and we have a very low time to first byte. A FlashBlade, a FlashArray, it, you know, is gonna get there, as well in a, in a similar way. But something that really is relevant to, to object storage. So I mentioned I've spent the last year
54:53
talking to a lot of very large customers about their on-prem object challenges that-Legacy on-prem object stores have is they are not terribly good at deleting and cleaning up data. Yeah. And part of that is because there are all these layers of indirection that many of these solutions, which are sort of software defined
55:14
in general have, which is, "Hey, I've the object." Okay, the object's been for deletion. That's the, the highest level of logical data. Okay. Well, now I go down and I've got, okay, now I've got the nodes where that data lives, and I gotta go mark that data for deletion. Okay, well then now they're gonna go do that.
55:30
Well, that data is actually sitting on a file system. Maybe it's like X4, maybe it's XFS, maybe it's something e- whatever it is. Okay, now I gotta mark that for deletion. Okay, now I gotta go do that thing. And then that might be sitting on SSDs, themselves have their own garbage collection
55:44
and reclamation process that they have to run. Right. And the more they do that, the more they have to do, you know, rewrites and garbage collection processes, and burn program Array cycles and so forth. So this is really where it gets down to, like, one of the big advantages that we object store with the, Kontxtual data all the way down to the physical data, is
56:07
that when someone deletes an object, we cut through all of those different layers because we've collapsed all of those different layers. When you delete that object or mark that object for deletion, when we go to reclaim it, it's being reclaimed at all of those different layers of indirection, which remain simultaneously.
56:25
So you're not going through this game of telephone where it's like, "Delete. Now you delete. Now you delete." It's like when this it gets done all the way through. Automatic, right. We garbage collect once, both at the media level, at the metadata level, and at the logical level, all simultaneously.
56:42
So we've actually had customers, there was a particular CDN customer in Japan where a CDN, like a content delivery network, workload, is enormously high churn. Yeah. Because you're like, "I'm, I, I gotta do something else," right? If I'm, if I'm doing a video, right, the latest, episode of whatever it is, that's gonna change maybe every week, right?
57:03
So I'm constantly o- turning over data. Well, they bought half as much capacity with a FlashBlade as they did with their solution because they could delete data so much faster. They didn't have to have a huge buffer of capacity there because it was gonna take Forever to just reclaim all of that data, and that's true both for, for FlashArray and for FlashBlade.
57:25
The way that we've architected our solutions is end-to-end visibility, from the logical data to the physical. And I think that really underscores the entire theme of what we wanted to set we're at the end. I, I, I think what you guys have laid out, especially that last point you made, Justin, is that, you know, as objects becomes more
57:48
prevalent and becomes used a lot more, these things that you just described with garbage collection and everything are going to become more prevalent. And existing legacy systems may not have thought of that or may have just written it off as, "Oh, it's fine." Whereas we are looking at it and we're saying, "No, we've got a much more efficient way to do this," when you get to scale in 3.8 trillion
58:13
these things are really gonna matter. So I, I think it's a great way to finish what our vision is as it relates to forward-looking in how object storage works on our platform, and also how it really is aggregated together to provide the vision that you see in the slide there with the whole idea of a hub, and then there's spokes to the edge, and we've got the right array for each one.
58:38
So, Justin, Karthik, thank you very much for joining us on the Ask Us Everything. Karthik, any final thoughts? And then I'll flash it over to, Justin. Yeah. So, you know, I think we've had a great conversation. Lots of great questions on, what, what both FlashArray and FlashBlade can do.
58:56
We're more than happy I know we, we, we may not have gotten to every, every question out there, you know, we're more than happy to kind of answer that, through other forums. Please reach out to us. I will tell you this, right? Our vision is to ensure that the stack, especially the protocol stack, remains consistent across whether it's FlashArray or FlashBlade.
59:14
So any capability that you see on FlashBlade today, over time will, will make its way into FlashArray and vice versa, right? You know, we wanna go where the, the data is and the form factors, you know, will dictate, you know, the environment will detect- dictate the form factors that we position the product in, and then Pure Fusion will orchestrate everything.
59:34
So things like SDS and so on and so forth, all of those things will, will, will have commonality across both. And so, you know, and, and one last point before I hand it over to you guys is, you know, FlashBlade's a, FlashBlade and FlashBlade, FlashBlade//S, great for multiple use cases, whether it's very small environments to very large environments.
59:54
Whether it's your, Rapid Restore, we have some large customers using FlashBlade//S for Rapid Restore for their backup needs and restore needs, especially in ransomware type environments, to AI training and, and, you know, high performance analytics and quantitative trading and so on and so forth. So, you know, you know, and, and great it, it's a packaged environment so you
01:00:16
have to worry about the management, the scaling, the non-disruptive upgrades, everything that comes with every Pure product, right? So keep that in mind. And of course, FlashArray, FlashArray object, the 1.0 version is gonna go out in a couple of months as, generally available. And, and we'll continue to march towards maturing that stack as well.
01:00:34
Awesome. And with that, Don, I'll hand it back to you and, and Justin- Yeah, Justin, it's, it You drive it home, man. It's up to you. Yeah. A- a- I know we're a bit over time, so really quick, if, if you do have those questions Karthik said, please feel free to ask them. Definitely leverage the community site, which
01:00:47
is, purecommunity.purestorage.com, I believe. Yep, there's a link, there's a link in the chat that Nicole- Yep put up there for everybody. Exactly. So we'll, we'll see you there. Thanks very much to everybody for coming.
01:00:57
Yeah, guys. So thank you very much again. It's a big, big topic, and I really, appreciate everybody tuning in to yet Ask Us Everything. I promised these guys we would get a lot of questions and answers, and we sure did, so there's obviously a lot of interest. And thank you again for tuning in on this.
01:01:12
And meet us out on Community, and we'll see you at the next Ask Us Everything, which will cover databases, all things databases. Thanks again, guys. We'll see you. Have a good day. Cheers, everyone. Thank you. Take care. Bye-bye