Skip to Content
Find dismissed updates here
Edit My Preferences
51:43 Webinar

Bring your hardest questions Building an AI Factory Live with FlashStack

This session cuts through the complexity with a live, practitioner-led walkthrough of the Everpure and Cisco FlashStack® for AI CVD—a proven reference architecture purpose-built for the NVIDIA AI Factory.
This webinar first aired on April 21, 2026
The first 5 minute(s) of our recorded Webinars are open; however, if you are enjoying them, we’ll ask for a little information to finish watching.
Click to View Transcript
00:06
Hello, everyone. We are so happy to have you with us today. So Melody and I were at NVIDIA GTC in March. We were both at the Everpure booth. Was anyone in the audience joining us today at NVIDIA GTC as well?
00:23
Let us know in the chat. Let us know where you're, where you're, tuning in from as well. Would love to, would love to see that. All right, so, we talked to over 1,000 people, when we were in the booth, at NVIDIA GTC this year, from an incredible variety of industries and roles.
00:44
Melody, what's your hot take from NVIDIA GTC this year? You're right, Erin. There was so many great things at GTC, and one hot take was definitely inference. It's eating the world. Jensen came out and said NVIDIA is the inference king, and the $1 trillion
01:07
compute demand projection through 2027 is being absolutely driven by inference. I mean, you, you can't not notice this. It's really not training. It's that meaningful shift from even just a year ago, I think. I'm sure you're noticing it too, and it reframes the whole infrastructure conversation.
01:32
It's not just about peak GPUs throughout anymore, it's about that sustained low latency, high concurrency serving at scale. Really at scale. Yeah. Absolutely. So let me, let me put this in perspective. I, I was standing there waiting for the booth to open and just kind of watching
01:56
wake up, and suddenly this toddler runs by, totally at full speed, down the center aisle, completely unpredictable. You know, as toddlers are, right? I mean, you and I have both had toddlers. Yeah. You know what I'm talking about, right?
02:12
Exactly. Yeah. It was totally unexpected, not something I was expecting to see at GTC. No. So, so of course I had to go investigate, and I turn down the aisle, and it's AWS getting set up for the day. They're, they're prepping their robot for demo, and you can see it come to life.
02:35
The movement is deliberate. It's articulated, not like rigid, blocky. It actually feels natural. Wow. And it struck me, that robot at that moment, it's on the clock.
02:51
It's not learning. It's performing. It's doing its job. It's taking in everything around it, camera feeds, lidar, touch. I mean, the guy was actually pushing it, like trying to push it over, trying
03:06
to knock it off balance. That was part of their demo, like showing how it moves in that moment, and processing that came through its pre-trained brain- and turning it into action in real time. That movement right there, adjust this, pick up that, because you would never want
03:30
it learning from scratch in that moment. That would be like trying to figure out how to drive when you're already in the middle of an intersection, right? Totally. It has to rely on what it already knows, and that's really the distinction. There's a time for training, and there's a time for inference, and when you're in
03:53
production, when it matters, you need that system to respond instantly, reliably, based on everything it's already learned, which is conveniently exactly what we're talking through today. So how about you, Erin? What was your take? Agents. It's all about agents, and I think that actually is, you know It, it, it sort of
04:17
builds on exactly what you are saying, right? As we move from training to inference, what, what are we doing with inference, right? We're trying to help agents become autonomous. So it's all about agents, which is sort of that natural next step. Jensen said every software company of the future will be agentic,
04:39
and I think that's true. Agents have massive potential to automate so many tasks. So we saw at GTC, ServiceNow spoke about using agents to triage, route, and resolve IT and HR tickets, without a human touching them. Roche and Eli Lilly both spoke about using agents to run experiments, making it way
05:03
faster to discover new drugs. It's kind of incredible to think about what's possible, but also at the same It's a ton of work, and you have to get so many things right at the infrastructure level, like security and governance, and the way you build and optimize for AI is completely different.
05:27
Yeah, absolutely. What I found interesting was how much of that agentic conversation immediately turned into a data conversation. Yes. Because Yeah, absolutely. Agents aren't just running models.
05:39
They're retrieving, reasoning, acting on live data, which means your data pipeline has to be fast and, and fresh, honestly reliable, and an agent that's working off sta- stale data is really just worse than no agent at all. Yes. You are absolutely right.All right, let's get into it.
06:07
I'm Erin Stevens, senior AI product marketing manager here at Everpure. You've now met Melody Zacharias, technical evangelist director, who because she's the one who actually understands how all of this works I think I'm just here to make the diagrams make sense. And I'm beyond thankful for that.
06:30
So here's the format for today. We want to take you through the ins and outs of what it takes to build an AI factory and an AI data platform. Sounds a little ambitious, but, but we're going to, we're going to get there. When we were getting ready for this session, I had this idea that maybe we could deploy a
06:52
FlashStack for an AI factory live, and basically I was told that it was a bit silly, because it's pretty straightforward and a bit boring to actually deploy, especially with that new Cisco validated design reference architecture makes it super straightforward. So instead, we're going to walk you through the architecture diagrams layer by layer so you can really understand all the components that go into our AI factory design
07:23
with Cisco and NVIDIA, and we really want your questions as we go. I see some of you are already letting us know where you're dialing in from in the chat. So Oregon, Colorado, Maryland. Love it. I'm in California today, but normally in
07:40
Eastern Washington. So it's, it's great to, great to have you all with us and, and participating. So please do, as we go, add your questions, in the Q&A panel and, we will get to those as we're walking through, as we're walking through those architecture diagrams. So the Q&A is open right now.
08:02
First off, let's talk about the problem we are trying to solve with an AI factory. We hear this pattern constantly. An organization proves that AI works in a pilot. The model performs, the use case is real, all your stakeholders are excited, but then you try to move to production and hit a wall.
08:21
A lot of these pilots are built successfully maybe in the public cloud but are too expensive to scale. So then you start looking at AI clouds or GPU-as-a-service providers like STN or BeyondPL. Those are both Everpure customers using our, our technology, behind the scenes, and they've
08:44
built these AI factories that you can essentially go and rent capacity from them. Or the other alternative you have is you can start looking at building your own AI factory. And And, and that's really where FlashStack becomes a strong option, because it's a full stack validated solution.
09:08
You're not spending months stitching together infrastructure before you can even start. Exactly. You, you can really begin with a deployment for a specific workload, and then scale as demand grows, adding compute, maybe some storage, or expanding that environment without having to re-architect.
09:32
The key is that you're scaling within a proven architecture, not rebuilding your stack every time your AI workload evolves. Cause we know that happens all the time, right, Erin? Yeah, and it's painful, and y- you know, this is exactly what we want to show you today. By the end of this session, you'll have a clear picture of what a validated
09:53
production-ready AI infrastructure stack actually looks like. Now, before we get into this, I want to give just a little bit of background for what an AI factory is and why it matters. Melody made me promise I'll keep this short, though. Yeah. We, we don't wanna get bogged down in slides.
10:15
Yeah, exactly. That's right. Before we get into AI factories and data pipelines, I want to start somewhere a little unexpected. Back in the 1800s, the Industrial Revolution wasn't really about smokestacks At its core, it was about a new kind of a machine, one that took an input, heat, and
10:36
turned it into an output, work. Sounds a little familiar. That simple transformation rewired the global economy. Entire industries were built around it. Entire cities were planned around where the machines lived.
10:52
We're living through this same kind of moment right now. A new machine has arrived. The input this time is electricity. The output, well, that's the question, isn't it? People call this the data revolution, but I think that undersells it.
11:09
What these machines actually produce is something closer to intelligence, and the buildings we're putting them in aren't really data centers anymore. They're closer to factories. And you may have heard this term used before. At the 2025 NVIDIA GTC conference, Jensen Huang, CEO of NVIDIA, introduced this concept
11:30
as a new paradigm for building data centers in the age of AI, and it's a concept that has absolutely stuck. So if intelligence is the output, how do we measure it? Well, NVIDIA has introduced this concept of tokens. Think of tok- a token as roughly a word or a piece of a word, the unit an AI model reads
11:52
and generates. It's like a little bit of intelligence. Tokens are to AI what packets were to the internet, and they're increasingly being measured as a quotient. Tokens per dollar, tokens per second, tokens per watt. That's the language of AI ROI, that's, that's being
12:10
discussed.So Erin, just a quick question to clarify here. Mm-hmm. When a CIO asks what does this AI initiative cost, are they actually asking in tokens now? Yeah. Increasingly yes, at least in the infrastructure conversation.
12:29
Cost per million tokens is showing up in planning docs the same way that cost per compute hour did a decade ago. Okay, so that's interesting conceptually, but why does this matter for today's audience? Thank you for keeping me honest, Melody. That is a great question to the point.
12:49
ESG Research found that organizations further along in their AI/ML and analytics maturity are two and a half times more likely to outpace their peers on customer satisfaction. They see 350% higher growth in revenue per employee, and are launching 46% more products than their peers. So AI really is the competitive edge that your business needs.
13:15
Now of course, the catch is it's moving fast. On the left side here, you see what's happening at the model layer. Models are getting smaller and edge ready, more specialized, grounded in your data through RAG, and on the reasoning side, more capable every month. And Evergreen//One of those shifts puts on the infrastructure layer on the right.
13:42
Okay, so it honestly feels like every month there's a new groundbreaking announcement in AI, so this doesn't surprise me. It does. Does, does it surprise you? Yeah, no. That's I, I mean, I feel it, and I'm sure that everyone here watching feels it too, right?
13:59
And it, again, it just puts more and more pressure on the existing systems. And so new deployment paradigms are emerging to meet these challenges, and that brings us to the AI Factory. So this is NVIDIA's new paradigm to optimize data centers for token output with purpose-built, built full stack AI systems.
14:23
Electricity and data go in, intelligence comes out. Compute, network s- and software, many of which are provided by NVIDIA, they're what make the factory run. Data storage is unique. NVIDIA relies on partners like Everpure provide, and it's the layer that determines
14:42
whether the rest of the factory actually runs. So data infrastructure is critical to AI success, but traditional storage systems weren't really designed for any of this. The demands of AI really strain existing storage systems at every stage of the AI life cycle, and that's really where Everpure comes in.
15:09
Everpure is built differently and is perfectly suited to power your AI factory. It's one unified platform built to handle every stage of your AI life cycle with three integrated AI solutions: FlashStack for AI as the converged factory stack, Data Stream, currently in beta, as the data engineering and AI readiness layer, and FlashBlade//EXA for large scale training and inference.
15:36
All right. Melody, without further ado, I'll hand it over to you to walk us through how FlashStack for AI works. Awesome. All right. This is the fun part. Let's build this. All right. I'm gonna bring up this architecture layer by
15:56
layer so you can see how each piece relates to one another, not just the final solution. This is the Cisco validated design. Everything you see has been tested, validated, and spec'd out. So let's start at the top. Here at the top of the stack is the AI software and orchestration layer.
16:21
NVIDIA AI Enterprise provides the AI software stack, including NIM Microsystems, m- sorry, NIM microservices for streamlined model deployment and inference, along with NVIDIA blueprints that accelerate common workflows like RAG pipelines and multi-model AI. On top of that, Red Hat OpenShift AI serves as the Kubernetes-native control plane
16:50
orchestrating AI workloads across the environment. Portworx by Everpure extends this with enterprise grade data services for Kubernetes, delivering persistent storage, data mobility, and consistent data management across on-prem and cloud environments, which is critical for stateful AI workloads. So Melody, when you say validated as an integrated stack, what does that actually mean
17:20
for a team deploying this? What are they not having to do? And that's key, right? That's why validated designs are so important. It saves them weeks of compatibility testing mostly.
17:34
The classic failure mode is you, you get your GPUs, you pick your software stack, and then you spend six weeks figuring out why your storage driver is tanking throughout at scale. With the CVD, that work is already done. You're deploying against a tested configuration, not discovering edge cases in production.
17:58
So yeah, this For example, this is exactly where s- where teams usually get burned. I've, I've lived through this. Back when I was working in, in banking, we were deploying a new SAN for a DR environment, and at the bank-Um, on paper everything looked totally compatible.
18:21
Servers, HPAs, switches, storage arrays. But what we didn't realize until we into testing was that the HPE firmware and driver combination wasn't actually certified with the array's multi-pathing stack. So under light loads everything looked fine, but as soon as we pushed real DR failover testing, IIO multi-paths and started seeing, path flapping, and when we
18:51
started to see this, we got inconsistent throughput, and at that point we weren't sure where the issues were. Was it the SAN, the fabric, the host config? We lost days just isolating it, and then the fix wasn't a patch. We actually had to order a new HB We had to order new HPAs that were on the storage vendor's supported compatibility matrix, swap them into multiple hosts and retest everything.
19:21
Ugh. That alone Yeah, exactly. Yes. That alone set us back, like, two weeks. Mm-hmm. So when we say validated integrated stack, what that really means is you're not discovering those kinds of incompatibilities under load in your DR test or, worse, in production.
19:41
That work is already done for you. Yeah. It's the, discovering edge cases in version that's the expensive version. Yes. It's the very expensive version. Yeah. That's how I see it.
19:57
And, you know, sometimes it's even just that extra little bit of time. Nobody likes to read long documents, but if it saves you days or weeks, sometimes worth it. Yeah. Absolutely. And honestly, now that we have AI, take that long document, run it through an AI and say, What do I need? Yeah.
20:18
And sometimes then you don't even have to read the document. Yeah, for sure. For sure. Right? So all right, so this next layer, the Cisco UCS Accelerated Compute, selection depends entirely on the workload, and I wanna be deliberate here because the GPU choice depends entirely on what you're going to be doing.
20:41
This is where a lot of infrastructure decisions go sideways. Because for inference, running trained models against live data, whether that's RAG, chatbots, semantic search, or agentic workflows, you're typically using NVIDIA RTX class GPUs, like the RTX Pro 6000. These are optimized for inference efficiency, delivering strong performance for real-time
21:07
response throughput and cost per token. Right? Back to that token. Mm-hmm. You can start with a single X400P node for edge or departmental use cases, and scale out to multiple GPU configurations for enterprise-grade model serving.
21:25
And importantly, you're not over-provisioning expensive training GPUs for inference workloads. Now, and for training though, that's a different conversation, right? Absolutely, and completely different. For training and fine-tuning, whether that's customizing models on your own data,
21:46
RLHF, or full pre-training, you move to high-end GPUs, like the H200 and the Blackwell generation. H200 gives you around 140 GB of HBM3E memory, which is critical for handling larger models and data sets in memory. And then with Blackwell GPUs, you're pushing even further on performance
22:11
That's what Jensen was highlighting at GTC. These systems are connected with NVMe- NVLink, enabling high-speed GPU-to-GPU communication, which is essential for scaling distributed training. Here it is less about cost efficiency and more about maximizing memory bandwidth and GPU-to-GPU throughput to reduce training time.
22:36
So- Two completely different situations. Yeah. So what's the most common mistake that you see when customers are speccing this layer? Well, honest answer, one of the most common is treating inference and training as the same problem, and they're absolutely not. They're completely different things.
22:55
So using training GPUs for inference workload, it works, but it's like using a semi-truck to do grocery runs. Technically capable, but widely inefficient for what you actually need. You're overpaying on cost, burning more power, and not really utilizing the GPU the way it was designed. I mean, in fairness, I'd also like a Porsche
23:20
Cayenne for grocery runs. Um- Sure but I have been told that the math doesn't math there, so. So, so true. Also the BlueField-3 DPUs, these offload infrastructure services like networking storage and security processing from the CPU, so the system can keep GPUs
23:46
I see teams skip them to save cost and then struggle with inconsistent performance or lower than expected GPU utilization because the CPUs become the bottleneck. At scale, that offload really matters. Yep. So then we get into some networking and, and Cisco Nexus 9000 provides the Ethernet-based AI fabric.
24:10
This is a dual fabric design which really does matter. You have a back end or east-west fabric dedicated to GPU-to-GPU communication. That's where your distributed training traffic lives, things like-All reduce. And then the front-end traffic for everything else, storage, access, management, and user traffic.
24:35
The back-end fabric is a high bandwidth, low latency 400, GB network using RDMA over converged Ethernet. It's designed to be lossless with congestion control mechanisms like PFC, ECN to tra- to keep traffic moving efficiency under load. And there's And it's also built as a non-blocking fabric, so you're not
25:06
bottlenecks as you scale out GPU nodes. Yeah, and network is re- really I feel people underestimate in the AI infrastructure conversation, KI- kind of like data storage, right? It's always compute, compute. Yes. And then, and then they build a training
25:25
cluster, kick off a distributed job, GPU utilization crater because the bottleneck isn't compute, it's the network. So all reduce operations get backed up, gradients can't synchronize fast enough, and every GPU in the cluster is sitting idle waiting on the network. So you're paying for that compute, but you're not using it at that point.
25:50
So at scale, your network can really become a limiter if your AI system, system, not your GPUs. Mm-hmm. So that can be an issue. At the foundation of all of this is your Everpure data platform, FlashBlade. It's This is the unified file and object storage, so you're supporting both
26:20
high-performance file access for active training data and storage for things like model artifacts and pipelines. It's designed for high throughput parallel workloads because in AI/ML you're not just storing data, you're feeding GPUs continuously. And, and then you rightsize it on the workload.
26:44
The S200 is typically used for inference style deployments in this range of up to 480 terabytes with low latency retrieval you need fast, low latency access to models and embeddings. The S500 is where you move into large trainings in the range of 960 TiB and above, handling massive data sets and sustained throughput across distributed jobs.
27:12
And then EXA is for the largest Your huge environments. It's for multi-petabytes, supporting the biggest models and the most data intensive workloads. A lot of systems can't hit peak performance for even a moment.
27:29
The problem is they can't sustain it, so you can get dips in data delivery and GPUs end up waiting. That's what we do differently. We deliver data consistently so those GPUs stay busy. It's not just about peak performance, it's about sustained throughput.
27:50
Keeping GPUs fed constantly is what drives real utilization. Yeah, and this is the part I always want to make sure lands, because data storage is the last thing that people think about when they're building an AI infrastructure, and it's actually The first thing that breaks down. Yeah. Especially in a good architecture.
28:13
Yeah, every single time. Yeah. Yeah, every single time. Yep. You can have the best GPU clusters in the world, but if you can't feed them data consistently, those GPUs are just gonna sit idle, burning costs. Yeah. And FlashBlade//S is really built to keep the
28:34
system fed. It delivers that throughput consistently, concurrently, and it's there to match how modern AI workloads actually behave. It scales linearly, and as your AI footprint grows, your s- performance scales right alongside it. Yep. And as we move to this last layer that I'll
28:59
have Melody, wrap us up with, please do make sure that you, put your questions in the Q&A panel. We can pause for a moment after this walkthrough before we get into the next one and answer your questions. I see one coming through in the chat, so we'll get to that. So Melody, I'll have you move to layer five.
29:20
So this last layer is management and automation. You've got Cisco intersite management and compute, and infrastructure lifecycle, Pure1 providing visibility into the storage layer, layer, and Ansible and Terraform automation.
29:44
But the point isn't the tools, it's that you have a single operational model across the stack. Yeah, none of this five different consoles open at once, type situation. Yeah. No five console situation. We can't have that. Mm-hmm.
30:03
And more importantly, you're not managing this environment manually. The automation hooks into your existing CI/CD, and MLOps pipelines, so provisioning, scaling, even updates can be handled programmatically. Because once you get past a single cluster, manual operations just don't scale. This is what lets you operate AI infrastructure the same way you
30:28
operate modern applications. So for example-Let's, let's just say a team needs to scale training environments from 32 GPUs to 128. Instead of manually provisioning servers, configuring networks, and attaching storage and validating everything, you're triggering that through automation.
30:50
The infrastructure comes up in a known good configuration, integrated into the existing pipeline, and ready for workloads without weeks of coordination across teams. Okay. That sounds amazing, doesn't it? Yeah, that sounds great. And, and there you have it.
31:08
It's the, the full stack. And we did have a question come in from Mark. In FlashStack network is internal within the hosts and between hosts. Is there something I'm not understanding regarding network? Steven jumped in to answer and said, "Hi, Mark.
31:26
The backend network is how the GPUs communicate across the dedicated fabric. The other front-end network is more public-facing for access to users, admin, and storage." So we have that one answered. I don't see any other open questions. So I will have us close this one out.
31:46
Melody, what does this mean in practice from a deployment standpoint? That it's a matter of, following the, the guidelines in the CBD. It's quite easy to deploy. Yeah. It's I mean, it's just like everything else that we do at Everpure.
32:11
The deployment is super simple. It's, Our, our regular deployments are usually the size of, a card, and FlashStack is not much different. Yeah. It's a super easy deployment. That's, that's why we didn't do this webinar on a deployment, because it was so
32:33
simple it was gonna be boring. We didn't- Yeah wanna bore you. Yep, exactly. Instead, we wanna make sure that you understand what's happening behind the scenes. I think sometimes too, you know, you, y- you see a demo and it's, you know, moving through a console and, and deploying, or you see, you know, you see that giant reference
32:52
architecture, which is great information, but can be daunting. I think it helps to have a picture, right, just to be able to understand what's happening for the components in the backend. So hopefully, hopefully this is, this is working for, for folks. All right, so, now I just wanna connect what Meli- Melody just showed you to the broader
33:13
NVIDIA ecosystem. The NVIDIA AI data platform, which we are going to, to go through next, is NVIDIA's customizable reference architecture designed to integrate NVIDIA accelerated computing, networking, and software into enterprise storage systems. And here at Everpure, we are working to deliver this with a new service called Data
33:39
Stream, currently in beta. Melody, do you wanna tell us what this means practically? So this is a part of, sorry. We're gonna go through this, through the Data Stream stack now? Yes. Yes.
34:06
I'm moving us, moving us to, the, the data platform portion, as we, we started running a little over on time on, on FlashStack, so. Ah, sorry. I've, I've moved you, I've moved you along. Was I taking too long? Was I having too much fun on that stack?
34:23
I think we were having too much fun. I think we need to plan, much longer for our webinars. All right. So let I'll, I'll, we'll dive right in then to, into this stack. Let me- Perfect let me dive into Let me, let me di- let me dive into this, AI stack then.
34:52
Perfect. Let's bring up the left side of that Excellent. Because that's really the reality of the Enterprise Data Cloud. We have here the, S3 object storage, NFS file storage, data warehousing, and different, all sorts of different formats, different teams. And the key points were none of these were
35:19
built with AI consumption in mind. Yeah. This is usually where people realize it's not that they don't have data, it's that their data is fragmented, siloed, and incompatible with AI workflows. E- exactly. And this is where 90% of unstructured data
35:40
really, really matters. Because unstructured data isn't just bigger, it's harder. Yeah. You know? The Kontxtual is rich. It's actually, not often not indexed, not labeled, not vectorized, not connected to anything.
35:56
But even so, it's, With the best GPUs and validated architectures, like what we gonna show you today, the models are only good as the data and what they can get access to, and that's entirely And right now, most of that data is effectively invi- invisible to AI, wouldn't you say? Yeah. And, and that's the gap, isn't it, right?
36:24
You know, there's so much focus on compute. Many, many folks, have solved the, the compute question or standardized infra, but the data pipeline, really getting from raw data to AI-ready data, doesn't exist in most enterprises.Exactly. And , oh, sorry.
36:46
That's exactly what this middle layer is about. Turning this, all of this, into something that AI can actually use. Mm-hmm. 'Cause most enterprises d- don't fail at AI because of their models or because they don't have enough GPUs. They, they fail because their data isn't ready or never made it
37:08
to the pipeline. Yeah, exactly. Would you say? Yep. Yeah. So think of Data Stream as the high-performance engine under the hood of your entire operation. It's a single SKU, fully supported AI-native powerhouse
37:29
that takes the grunt work out of the equation, sort of- Yeah by automating from messy raw data to refined AI-ready asset. You know, just making things simple. Instead of juggling a dozen different tools, you get a streamlined five-stage process that handles everything, and we're gonna walk through exactly how each one of these gets you
37:56
to that finish line. So let's start with ingest. Forget the headache of building and maintaining messy web of custom connectors for every single system you own. Data Stream simplifies that chaos by pulling in data from any source, be it object, file, database, or live stream, through one API.
38:22
Everything is normalized into a single sleek pipeline right out of the gate, stop playing IT mechanic- and start actually using your data. Which is usually how these custom pipelines are described in the postmortem. We just need a custom connector for each source, right? And then there are 12 sources, and then someone leaves the team.
38:49
Yeah. Exactly. Isn't that always how it happens? Yes. Yeah. So this is where we tackle the curate phase, the heavy lifting of cleaning deduplication deduplicating and enriching your data with metadata to keep everything organized.
39:16
While most teams burn huge amounts of manual effort and caffeine, I know I've done a lot of caffeine, trying to scrub their data sets by hand, we've made the whole process automated and consistent, so you can get clean clean-room quality without human error. Yeah. And then transform. This is where we take the heavy lifting off your plate by converting content into formats
39:47
AI can actually wrap its brain around. Mm-hmm. Handling all the messy chunking, parsing, normalizing in the background. It's the essential bridge that transforms a dense PDF full of legal jargon into a streamlined structured format that your model can instantly retrieve and actually
40:08
reason over. Yeah, the step everyone skips until the model starts hallucinating and they have to figure out why. Yes, absolutely. That's a great point. Exactly, because garbage in, garbage out, just with more GPUs involved as we add GPUs to everything these days. Yeah.
40:30
Mo- most expensive garbage disposal in enterprise IT right now. Yeah, I would agree with that. This vectorizing is the most important, I would think. It's This is, this is a moment in the data truly becomes AI-native. Mm-hmm. And we're seeing it even in relational
40:58
databases now. That's how important it's becoming. We're generating embeddings and indexing them for deep semantic retrieval, moving WEKA beyond the limitations of old-school keyword searches. Mm-hmm. Because this is natively integrated with NVIDIA, NeMo Retriever, and you're stuck standing
41:21
up and stitching together a separate vector database. It's all backed right into the workflow for a seamless high-performance finish. Yeah, and this is the part I always want to demystify because vectorization sounds intimidating, and it's actually just we're creating a way to find things based on meaning, not just an exact word match.
41:48
Right. You wanna be able to ask, "What are our policies around customer data?" And get the right answer, not just documents that contain those exact words. Yeah, exactly. So finally, we're delivering AI-ready data straight to the GPU layer in real time, and we're doing it continuously.
42:10
As fresh data hits the system, it flows through the pipeline automatically, ensuring your AI applications are always grinding against live current information rather than yesterday's stale snapshots. And the continuously part is important, right? Because the alternative is a batch job that someone has to run.
42:30
And monitor- Mm-hmm and restart when it fails- Mm-hmm and explain to someone why the data in the RAG application is two days out of date. Yep. Yeah. Absolutely.So two properties worth calling out explicitly. In-place processing, you're not copying everything to a new location first, and
42:53
continuous operation, not a scheduled batch. You're also scaling on data needs, not compute needs. You're not paying for extra compute to run your data pipeline. Yeah. What's, what's the thing that teams most consistently underestimate when they try to build this themselves?
43:21
Typically the maintenance. The pipeline, the pipeline works on data one on, sorry, on day one. Mm-hmm. But six months later you have three new data sources. Someone changed a schema upstream, or your embeddings are stale.
43:39
The ongoing, ongoing work of keeping a DIY pipeline healthy is the part nobody plans for. Yeah. Yep. And then on the very bottom is, here on this diagram, is the NVIDIA AI data platform, a validated reference architecture for running AI at scale.
44:10
This is the- So- Sorry. Go ahead. I'm sorry. Go ahead. This is the full stack GPU compute with Blackwell, BlueField-3 DPUs, NVIDIA AI Enterprise, and NIM microservices. FlashBlade//S with Portworx provides, the storage foundation, and NeMo
44:31
Retriever with, C- CUVS handles retrieval and vector search between the data layers and the models. This isn't a collection of parts. It's a tested, integrated system designed to work together. So this is the same validated stack looking at in the FlashStack diagram, just
44:52
from the data pipeline perspective. Yeah, exactly. Yeah, exactly. The infrastructure doesn't change. The GPUs are the same. The storage layer is the same.
45:04
The NVIDIA software stack is the same. What changes is the data. Without Data Stream, only a small fraction of the enterprise data ever makes it into this pipeline. With it, you have a continuous flow of AI-ready data feeding
45:19
directly into the platform. Mm-hmm. And so then on the right is what comes out the other side, right? This is the- Yeah AI/ML use cases. So what comes out the other side are the AI use cases everyone is trying to build. RAG pipelines, real-time inference, fine-tuning on domain-specific
45:53
data, systematic semantic search, sorry, agentic AI, multi-model, multi-modal AI, across text, images, and video. And these aren't aspirational. They're what is vali- what this validated stack is designed to support today. Yeah. And the agentic one is the one I want to flag
46:19
specifically given what we talked about in the GTC section, because agents are only as good as the data that they can access. A well-architected data pipeline, is what actually makes agents reliable. Yeah. That's totally true. Exactly right. An agent retrieving from incomplete, stale, or
46:43
unvalidated data, even worse, will give you an answer. It'll just be wrong, and confidently wrong. That, that- Yeah. Yeah. And that's so much a harder problem to solve, not ha- or than not having an answer at all.
46:59
It- Yeah and way worse. Exactly. You have to be able to, to trust, especially as we start to think about, letting agents loose to make their own decisions, right? That, that data quality is critical. And this data infrastructure is critical to having, AI agents that,
47:17
that you can actually trust. All right. So, Melody, oh, my gosh, that was oh, I thought I had a slide here that said Q&A, but that's okay. We'll, we'll leave this one up. Melody, that was amazing.
47:35
Are you exhausted? It was a lot to go through. Well, I was starting to lose my voice, but that's okay. Yeah. We're almost done. It's okay. I'll let you drink some water. I, I, I'm not, I'm not that hardcore.
47:47
Okay. So, this is the moment that we have left for questions. So please do, put your questions in the, in the chat. I'll start with Melody. How do you size a FlashStack deployment for a specific inference workload?
48:05
What are the key variables there? Ah, that's a good question actually. We kind of touched on this in, in the talk, but sizing is important to start with the workload be- not the infrastructure itself. Mm-hmm. So for inference, the key variables are, are model side- size, but concurrency,
48:28
latency targets, and data access patterns. Mm-hmm. So, like, a small model with low current- concurrency and relaxed latency looks very different from a multicloud real-time application serving thousands of requests per second. Yeah. So you, you kind of have to go based on what
48:50
kind of workload it is that you have. From there, you work backwards and look at GPU requirements, 'cause those are driven by throughput and latency SLAs.Well, your storage and your network are driven by how quickly you need to feed those models with data. If you're doing RAG, retrieval latency and vector search performance become just as
49:14
important as raw compute. So the advantage of FlashStack is that it's a validated architecture. So instead of guessing how components will behave together, you're sizing within a system that's already been tested for these kinds of workloads. So that makes it a whole lot easier, and then it's just a matter of sizing within what your
49:36
needs are. Fantastic. Fantastic. Well, for any more questions, please send those in and, and we'll, we'll, we'll try to answer in the chat and, and reach back out to you. Thank you, Melody.
49:51
Genuinely, this is way more fun when the person next to you actually knows how all of this works, especially when it's you, Melody. And, thank you, everyone for your time, your questions. Really appreciated, really appreciated the engagement and, and the, the questions today were, were great.
50:15
So the recording will go out to all the registered attendees. If you want to go deeper, if you want to map your use cases to an AI pod configuration, talk to, talk about Data Stream for a environment, we have an amazing solutions team that specializes in AI. You can connect with them, via Everpuredata.com.
50:38
And also watch for our IDC report in May. We, we are going to be releasing some joint primary research, that, looks at, results from 1300 AI, infrastructure practitioners, so please stay tuned for that. And of course, we do have Pure//Accelerate coming in, up in June in Las Vegas, so register for that.
51:05
Come and see us. And as always, join, join the conversation on the Everpure community. Again, a lot of our solutions folks are active on there, so if you wanna continue this conversation, it's a great place to go. Oh, there's my question slide.
51:22
Well, we'll skip that. Thank you so much for joining us today. You have been a fantastic audience, and we'll see you next time.
  • Artificial Intelligence
  • FlashStack
  • NVIDIA
  • Cisco
  • Expert-led Demos
  • FlashBlade

Erin Stevens

Senior AI Product Marketing Manager, Everpure

Melody Zacharias

Technical Evangelist Director, Everpure

Most organizations racing to build AI infrastructure are assembling point solutions that create hidden bottlenecks before a single model ever trains. This session cuts through the complexity with a live, practitioner-led walkthrough of the Everpure and Cisco FlashStack® for AI CVD—a proven reference architecture purpose-built for the NVIDIA AI Factory. 

Experts from Everpure will show you exactly how FlashStack eliminates the guesswork from AI infrastructure deployment, from storage and networking to compute and data pipeline readiness. 

Key takeaways include: 

  • Why a CVD-backed reference architecture reduces deployment risk and accelerates time-to-AI
  • How FlashStack integrates with NVIDIA technologies to support training, inference, and agentic workloads at scale
  • Perspective on what enterprises get wrong when standing up AI infrastructure—and how to avoid the most costly mistakes
04/2026
Everpure FlashArray//X: Mission-critical Performance
Pack more IOPS, ultra consistent latency, and greater scale into a smaller footprint for your mission-critical workloads with Everpure®️ FlashArray//X™️.
Data Sheet
4 pages
Continue Watching

* indicates a required field.

We hope you found this preview valuable. To continue watching this video please provide your information below.

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Your Browser Is No Longer Supported!

Older browsers often represent security risks. In order to deliver the best possible experience when using our site, please update to any of these latest browsers.

Personalize for Me
Steps Complete!
1
2
3
Continue where you left off
Personalize your Everpure experience
Select a challenge, or skip and build your own use case.
Future-proof virtualization strategies

Storage options for all your needs

Enable AI projects at any scale

High-performance storage for data pipelines, training, and inferencing

Protect against data loss

Cyber resilience solutions that defend your data

Reduce cost of cloud operations

Cost-efficient storage for Azure, AWS, and private clouds

Accelerate applications and database performance

Low-latency storage for application performance

Reduce data center power and space usage

Resource-efficient storage to improve data center utilization

Confirm your outcome priorities
Your scenario prioritizes the selected outcomes. You can modify or choose next to confirm.
Primary
Reduce My Storage Costs
Lower hardware and operational spend.
Primary
Strengthen Cyber Resilience
Detect, protect against, and recover from ransomware.
Primary
Simplify Governance and Compliance
Easy-to-use policy rules, settings, and templates.
Primary
Deliver Workflow Automation
Eliminate error-prone manual tasks.
Primary
Use Less Power and Space
Smaller footprint, lower power consumption.
Primary
Boost Performance and Scale
Predictability and low latency at any size.
What’s your role and industry?
We've inferred your role based on your scenario. Modify or confirm and select your industry.
Select your industry
Financial services
Government
Healthcare
Education
Telecommunications
Automotive
Hyperscaler
Electronic design automation
Retail
Service provider
Transportation
Which team are you on?
Technical leadership team
Defines the strategy and the decision making process
Infrastructure and Ops team
Manages IT infrastructure operations and the technical evaluations
Business leadership team
Responsible for achieving business outcomes
Security team
Owns the policies for security, incident management, and recovery
Application team
Owns the business applications and application SLAs
Describe your ideal environment
Tell us about your infrastructure and workload needs. We chose a few based on your scenario.
Select your preferred deployment
Hosted
Dedicated off-prem
On-prem
Your data center + edge
Public cloud
Public cloud only
Hybrid
Mix of on-prem and cloud
Select the workloads you need
Databases
Oracle, SQL Server, SAP HANA, open-source

Key benefits:

  • Instant, space-efficient snapshots

  • Near-zero-RPO protection and rapid restore

  • Consistent, low-latency performance

 

AI/ML and analytics
Training, inference, data lakes, HPC

Key benefits:

  • Predictable throughput for faster training and ingest

  • One data layer for pipelines from ingest to serve

  • Optimized GPU utilization and scale
Data protection and recovery
Backups, disaster recovery, and ransomware-safe restore

Key benefits:

  • Immutable snapshots and isolated recovery points

  • Clean, rapid restore with SafeMode™

  • Detection and policy-driven response

 

Containers and Kubernetes
Kubernetes, containers, microservices

Key benefits:

  • Reliable, persistent volumes for stateful apps

  • Fast, space-efficient clones for CI/CD

  • Multi-cloud portability and consistent ops
Cloud
AWS, Azure

Key benefits:

  • Consistent data services across clouds

  • Simple mobility for apps and datasets

  • Flexible, pay-as-you-use economics

 

Virtualization
VMs, vSphere, VCF, vSAN replacement

Key benefits:

  • Higher VM density with predictable latency

  • Non-disruptive, always-on upgrades

  • Fast ransomware recovery with SafeMode™

 

Data storage
Block, file, and object

Key benefits:

  • Consolidate workloads on one platform

  • Unified services, policy, and governance

  • Eliminate silos and redundant copies

 

What other vendors are you considering or using?
Thinking...
Your personalized, guided path
Get started with resources based on your selections.
My Updates
No updates at this time.