00:06
Good morning, good afternoon, good evening. Welcome to our webinar today, our TechTalk, Virtualization Reimagined Inside the Everpure Journey. Thank you so much for joining, today for what's gonna be a really nice discussion about what we're doing, around virtualization.
00:22
Before we get started, I wanna introduce my two colleagues who are joining me today. Greg, I'll start with you. You wanna give a quick introduction on, on yourself and your background? Sure. Thank you, Andy. My name is Greg McNutt, a Technical Director here at Pure.
00:35
Been here about nine years, working on all kinds of things related to infrastructure, and excited to talk today. Thank you. And, Sagar, how about yourself? Hey. Thank you, Andy. my name is Sagar Srinivasa. I've been in Pure for almost three years now.
00:51
Same as Andy, working on infrastructure, automation, and Kubernetes. So excited to be here. Excellent. And, I'm your host today, Andy Gower. I am, the Director of, Product and, and Solutions Marketing on the Portworx side, Portworx by Everpure.
01:09
And I'm really excited about today's conversation because it's really all about virtualization and, virtualization being reimagined in this modern world. I think we all know about the changing landscape in virtualization, and we're looking for what are companies doing? How are companies responding to that virtualization shift, and what are they doing
01:27
as they look at alternatives? And today's discussion is really about what has Everpure, what has our own team done, in reaction to some of the, the trends and some of the things happening, around virtualization as a whole. Four parts to today's webinar and, and it's really gonna be more of a discussion than, than a slide presentation.
01:43
But, we're gonna talk about, you know, rethinking virtualization at its highest level. We're gonna dive into and understand what were the reasons, both technical and economic, behind our revirtualization or our rethink of virtualization. We're gonna talk a little bit then about the architecture and design principles. How do we build a unified platform for running VMs and containers?
02:02
What are the design principles? What are the North Stars for VMs and containers? How do we manage that? Then we're gonna get into the migration journey itself. This is an area that as we talk to as I talk to customers, as the team has been out there talking virtualization, people are
02:15
really interested in. It's great to talk about a new design. It's great to talk about a new How do you actually migrate those VMs? How do you build a repeatable migration set of tooling and frameworks that allow you to go
02:26
through the journey? We're gonna dive into a couple of things, that Greg and Sagar have, have really built to help, help drive that forward. And then we're gonna end with oper-- outcomes and next steps, right? What's been the early operational impact of what we've done? What are the lessons we've learned, and, and where does the platform go, from here?
02:42
Now, throughout this, I have a couple colleagues who are on the chat. Feel free to ask questions. They can answer questions throughout, and then we'll have time at the end for a Q&A. So feel free to ask us those questions and, hold them for the end if you're interested, and ask them live. Otherwise, we'll, we'll tackle them in the
02:56
chat throughout the conversation today. Okay. With that, let's jump into that first topic, which is really about rethinking virtualization. Why do we do it? How do we do it? What are the factors behind it?
03:08
And, and Greg, I'm gonna start with you on this one. Can you, you know, walk through some of the key drivers between rethinking our virtualization strategy? What, what came to you? What were the things that were coming to you, that really pressed you to explore a new virtualization strategy?
03:22
That's good, good question. The tri-- the, the catalyst was se-- several things, but really the thing renewal costs with VMware basically pushed this thing over the line. Prior to this, we've had a big estate of a lot of different systems, a lot of different hypervisor systems, and we've been a long time about consolidating that
03:41
was, this was essentially the trigger. So, you know, once we, once we went down that path, then it's like, okay, where do we go, right? Where do we go for consolidating these things? How do we manage our complexity? How do we get over the line?
03:53
And, you know, our pick was something that was probably something we could control, and the idea with KubeVirt and Kubernetes is it's something that we'll talk about in detail, but it really has a lot of the checkboxes we want for simpli-simplified, we keep control, some kind of future costs that we can manage. And, and it's also kind of the state-of-the-art, so we, you know, we, we
04:13
basically-- that's our journey. And Sagar, I know, you know, we talked a little bit about cost and licensing pressure. Talk a little bit about how this been driven by platform monetization, some of the things that we're doing around VMs, containers, and looking forward to AI. How did that play a role in the strategy overall?
04:32
Yeah. So we, we've been having talks about how we can bring both VMs and containers into one picture and platform admins and as well as, infrastructure owners, like how do-- can see a single pane of glass, to manage both of them, monitor both of them, have real alerting around them. So that was another driving factor for us, when we were looking at costs and, you know,
04:58
how we can be, not be vendor locked in and, you know, go through that strategy. Yeah. Yeah. I think that last one is particularly important too. You know, you've, you've gone through this, this challenge once with the renewal costs and trying to figure out how to get to the next step.
05:15
I think avoiding that vendor lock-in long term, prioritizing open platforms, ecosystems, I think was particularly important for our strategy. We'll get a little bit into that design principles. But Greg, do you wanna comment a little bit more on that, that lock-in piece and how that played a role in developing the strategy?
05:31
This is really kind of a business cycle thing. In the beginning, when our company was very small, we were building most of the stuff on, you know, manually managed KVM platforms. And we've, you know, went through OpenStack. We've gone certainly through VMware and, and the sprawl gets there.
05:49
You have more and more platforms to manage. There are more and more to keep up to date. And so we said, "Look, we wanna go into this." You know, VMware is the gold standard. There's no denying that. But with costs and, and control and kind of our future plans, this wound up being our
06:04
right story. So we said, "Look, let's go with something that is probably likely to remain open source for a long time, and there's a around it. Let's get on board." And that's, that's, that's why Kubernetes. So we, we talked a bit about, you know, the drivers behind the virtualization strategy, and you hinted at a little bit of what it looked like or where
06:22
we're starting. Sagar, I'll, I'll start with you on this one. Can you talk a little bit about what the virtualization footprint looks like today or, or coming into this- Mm-hmm this effort? Yeah. Definitely, yeah. So currently we have workloads, a majority is in VMware. We have around close to 10,000 of VMs, spread
06:44
across, multiple data centers, all under VMware with different versions of ESXi. And also multi-tenancy, so we have different BUs and different teams. We do also have some VM footprints in OpenStack, for dedicated workloads, and their R&D and testing.
07:06
So that's kind of where we look at it from a, you know, a 10,000-foot distance like, okay, this is our environment and footprint, so how do we make sure we, you know, unify them and kind of bring them together, in a single pane of glass? And I'm gonna, I'm gonna come back to you o- on this one too. You know, you mentioned the VM footprint and, and OpenStack.
07:30
You also have a Kubernetes footprint there in pockets as well, right? So starting with- Yes a little bit of containerization that had occurred across- Yes your organization in different pockets. Yes, that is true. So we do have workloads which are truly containerized, and, we use STaaS products, as well as they are hosted on cloud, be it
07:48
AWS, Azure or on-prem. But the on-prem majority of them have been in single silos, right? Some running under VMware as a VM, Kubernetes hypervisors, some running, maybe, in OpenShift world. So this was like, a, a reach out to see if how, how can we what are the, what are the gains
08:10
that we're gonna get when we try to bring them together, or do we need to So there were a lot of thinking and, I would say, whiteboarding that happened, to identify those workloads and m- make them part of this journey. Y- you talked about bringing them together. I wanna Greg, I'll, I'll pass this one to you.
08:29
Take a step back and talk about the operational complexity that led you to want to bring them together. What, what does it look like today with all those separate platforms or VMs, containers, and different infrastructure stacks? Yeah, as most, operations teams recognize, there's, there's sprawl. It's lots of different implementations of roughly the same thing,
08:49
and we're, we're the same. You know, it, it's like, you know, projects to do upgrades, they're not leveraged. You wind up having four or five disciplines that you've got to go through there. So from the, just the lifecycle these systems, it was definitely a do some consolidation.
09:04
Now, consolidation is sometimes we call it a forklift move, sometimes we call it a replatforming, and, and these were all part of the analysis about where we're at. Again, in this case, we said, "Look, let's redouble, let's focus, let's go platform that's gonna address, you know, all of our current needs and, and head that way and get everybody behind that." That focus and that leverage that we're now getting is really,
09:27
really paying off. And everybody, everybody says that we all go through these periodic consolidations, and this for us is one of the larger ones, definitely, and it's, and it's already paying off. And I will talk about outcomes, shortly here, but before we get there, you know, I want to talk about, and we've hinted at a few of these already, but some of the
09:48
the new virtualization platform. And, and Greg, I'll start with you on this. Talk about, you know, a couple of these key important ones around unification, around, you know, you mentioned it, reducing the, the platform sprawl. You know, talk about some of those key things you wanted to get out of a new
10:01
solve if you were gonna make a shift away, from what you knew for quite a long time with VMware. Right. So the There's a, there's a joke saying it's all about the networking, and that, holding true here. But what has happened, the way we approach workloads and deployment is far more configurable than before.
10:21
Before it would be heavyweight static deployments for, for a lotta, a lot of applications. And, and we're more and more going to these are more ephemeral, they're coming and going a lot. And when you put together a application built of a bunch of parts, you have to network them together.
10:35
That's just pretty standard, and having them under one control plane is a huge So the ability to manage networking, you know, hardware level networking too, between boxes, whether they're virtual machines or containers, and consolidating workloads from everything under our control using open source products is the, is the key drivers. Like this is the f- the future.
10:56
This is the approach for these hybrid systems which are going to be in our So yes, one control plane means one policy engine, one networking platform, you know, one imaging platform. Observability winds up being unified. So a lot of things come together, and that- that's your kinda number one and two there on this list.
11:13
And, and Sagar, I'll go to you for a couple of the o- other ones. Talk a little bit about, you know, the developer self-service and the platform for engineering and how-- why those were such important requirements for the new virtualization platform. Sure. So currently, if I'm a, a end user and I need
11:30
to provision, a workload in Kubernetes or, say, any platform, be it VMware, I have to go through multiple systems. There's multiple tickets I need to open. There's self-service, yes, but I see the orchestration going to different, APIs, and then they're all working together.
11:49
They are combined and intertwined. And there's also some, I would say, automation benefits that we-- that, that, that was amiss, right, before going in. And then we kind of looked at, how we can make it unified that way. We have one system and truly networking, be it storage or compute, they all are,
12:12
you know, well defined under this, big, I would say, fence. So that way we kind of-Make sure that the automation is API-driven, not too much, talking to multiple platforms and multiple different endpoints. Just one Kubernetes API, and then you get the benefits of faster provisioning and then, the truly benefit from the Kubernetes, I would say, ecosystem.
12:41
So that kind of paved the way for most of our, outstanding, automation as well, right? So if be it application onboarding, be it, multi-tenancy access to multiple teams and users, now it's all integrated with, Okta and OIDC, so they can also, kind of they just choose which VMs and what kind of workload. Does it have to be VM?
13:07
Does it have to be a container? And what are their images where they would live? And just kind of truly, I would say, utilize and benefit from this, one ecosystem rather than, you know, orchestrating between multiple platforms. Yeah. And I'm, I'm gonna come back to this point, in
13:26
a couple of slides, but I th- I, I think this is a, a key element of kind of the outcome side of things. And I like the term that, that we used when we were talking before this around vertical app integration. This notion of right now the apps are kind of disparate, and the ticketing systems don't talk necessarily to each other with the networking, with the, the
13:43
various other requests. By going down this kind of route that we're gonna talk to in, in the next couple slides, you now have that vertical app integration where everything for a platform team, for an application owner is integrated in one stack that they can, address as you were talking about just now. So I think that's a very powerful benefit.
14:01
We'll come back to that, a-after we talk about the stack itself. And with that, I'm gonna use that as, transition to get into what is the actual architecture and the design principles, we put together in this new environment. And I'm gonna start with, with these, with these design principles, and we've touched on these in a little bit.
14:19
But Greg, I'll start with you here. Talk a little bit about, you know, these key design principles, the openness, the ability to scale, the, this GitOps operation we've been talking about and, and that hybrid cloud flexibility. Why were these the kind of four north stars, of any architecture that,
14:32
that were, was put together? Yeah. That's, you know, it's always the desire to get to some place where you maximize control, and this was a lot of these are ways of saying that we have control. You know, yes, using open source means in, in emergencies we can make direct
14:49
changes as, if necessary, and we do those occasionally, and we work with the, you know, we work with the ecosystem to give and, and take on both directions. The-- I think, you know, things like scale are fine with other, other platforms and, you know, code-driven operations. These are all, these are all things that we've done in pockets, like,
15:10
like was mentioned earlier. I, I think this unification, this sort of, this, this strategy of pressure towards consistency is really a story about leverage. It's like, can we have a few experts that understand the entire platform? And it's not, it's not an OpenStack specialist or a, you know, a KVM
15:27
Kubernetes or something. They're all, they're all essentially one. So, so yes, we're building the hybrid apps, which we talked about earlier, and we wanna go with open, open source, so that we have some level of control. And this is a, this is a industrial-grade platform that we've chosen, so it'll work at scales, you know, that, that we certainly need.
15:46
So yeah, these are all, these are all just the s- the center. Exactly. So let's, let's talk about that platform itself. And Sagar, I'll, I'll go to you. Can you walk us through the architecture and, and some of both the, the benefits, but also some of the challenges you encountered along the way as you were putting together the
16:02
architecture and starting to, you know, put it into, test and production versus putting it down, on the page to look at? Definitely, yeah. So starting off our challenges was scaling, right? So we need to make sure the architecture is scalable.
16:19
At any point, there should not be, a point of, you know, aggregation or so we wanted to go with somewhere where we have, infrastructure as code, definitely. But if you think about it, like today, everything on cloud, like it's, it's easy to manage and, as, as, as compared you have Terraform, you have CloudFormation for AWS and other kind of tools which kinda, well, and then within a button you have the
16:47
a-architecture up. When it comes to bare metal Kubernetes, it's a different game. So we do want to get the, c- we, we need, we need to address the complexity behind it. How do you provision? How do you keep the desired state? So we kind of chose the, Argo CD as our defining factor for,
17:10
all of the infrastructure, based tooling, that's gonna go and live inside our Kubernetes cluster. Whereas, we chose GitHub Actions, as our pipelines to manage our bare metal, nodes as well as FlashArray, right? So we, we developed a pipeline which actually gets inputs like, hey, what kind of cluster are you building? Is it four node, 10 node?
17:36
And we do have an internal API, where see and, you know, pick and choose what and, of nodes that we need. Kinda like cloud, but, but still, operating under, you know, data center, complexities. Once we had that maturity, we found out that there could be requests
17:58
Like, there could be like a big cluster of ten nodes or maybe a hundred nodes or 200 nodes. How do we scale this? How do we make sure, we can provision them without any bottlenecks and multiple teams involved, right? So that was the challenge that we dove into. And the result was pretty much-Um, like, to, to come up with automated pipelines.
18:22
And we use Ansible, and we use, something called, Kubespray. So Kubespray is an open source, Ansible, heavy, automation that is available workflow to manage your, Kubernetes environment. And it need not be bare metals, but it also supports bare metals as well. So there's a big, open source community which supports it and maintains it.
18:44
So, that kind of helps us in upgrades, adding nodes to our clusters, or doing all the, day two operations on the clusters. Going from there, we kind of now wanted to see how we're gonna gain the same, I would say, operational functionalities that comes with VMware, right? Within VMware, like you, you-- if, if you are a VMware admin, you would have to go into
19:12
VMware vProvision, or if you have a self-service which is tied to VMware, there's like a life cycle where you request VMs. So we wanted to s- we wanted to bridge that gap when we're going with KubeVirt, right? So, we have some-- we have a VM lifecycle UI, which is called as KIF. Internally, we call it KIF.
19:31
So it already supports VMware and OpenStack today. So we were like, why don't you use the same, UI? Why don't you use the same platform, and see how it can integrate with KubeVirt? And, obviously it's a cloud-native environment, so it is API-driven. So we had all the, I would say, architecture in front of us, which was easily able to
19:51
integrate with our existing life cycle. You talk over SDKs, we talk over APIs. And, the team, the, the, the team that works, and maintains KIF was very, very helpful and accommodative to come up with this kind of solutioning. So there was a lot of back and forth, we identified what kind of, I would say day
20:11
two operations a VMware admin does, right? Like turn off, turn on VMs, take snapshots, add users, maybe increase the configuration, increase CPU, RAM, or, you know, retire a VM. So all those, actions were identified, and then we kind of gave them to our, UI so that it can process and it can send those over to KubeVirt over APIs.
20:36
And if you look at the, ActiveCluster, and I want to talk about a little bit about, what's inside the cluster, right? We talked about how we integrated, how we, you know, get in and request things. But inside the cluster, the major role played by, I would say the two, two components are the networking and storage.
20:59
So we have Portworx, which is basically, backing, backed by our FlashArray, product, which the CSI for this cluster. And we use, Cilium for the CNI, which is actually, talks to our switches and things or BGP to, expose our network, to the entire infrastructure. So these two were, are maintained by Argo, CD.
21:25
So whenever we need to upgrade them, they are upgraded in a cloud-native fashion because we are, we ha- Argo is always monitoring, over a GitHub repository. So a PR to change my Cilium version would just trigger the upgrades inside the cluster. Same way for Portworx as well. And KubeVirt itself.
21:45
KubeVirt itself has its own C- CRDs and version controlling that we do through Argo. So now it kind of came to picture like, we have the nuts and bolts in place, and this is how our upgrade cycle is gonna look like. This is how our software upgrades will look like. And then we kind of de- delved into how do we do maintenances using Kubespray.
22:05
Yeah, that's pretty much how everything ties together. So, so I wanna kind of double-click on, two areas, and I'm, I'm gonna start on the networking side. I know- Mm-hmm this is, you know, a- and guys have mentioned it a couple of times, networking is one of the, if not the biggest challenge as you go through this.
22:23
Talk a little bit more about both Cilium, but also some of the other things you have to explore in the, in the broader community to, uh- Mm-hmm make, make this architecture work and really make it scale as you started to build out, build out this architecture into an actual, deployment. Mm-hmm. Definitely. Yeah. So when we started out, obviously in
22:43
Kubernetes world, if you're familiar, you have containers living in there, throwing into in VMs, right? So, there's, there's a different architecture when it comes to pod versus VM, right? So a pod usually lives, in be- behind a load balancer of sort, which is exposed to the outside world as a service, right?
23:02
So how do we expose a VM? How do we basically make sure, we access the VM from outside world? So that challenge was pretty, common, and we knew that, okay, if we gonna do an ingress or some kind of, load balancer way, then to define ports and if we have to define, what's gonna get in, how are we gonna get out of the systems, right?
23:23
And that also came with challenge from our end users like, "Hey, I wanna be VM from its, you know, FQDN, whatever the DNS name is every time. I don't want the VM to lose connectivity, whenever, whenever it is migrating between nodes." usually that, that's what happens, even with Kubernetes, right? Where you have your pod scheduling to different nodes, during outages or
23:46
self-healing, I would say, operations. So, we, we started with something called as, pod subnet. We c- we, we use BGP with Cilium, and, advertise, the pod subnet, from our top rack switches to outside world. So that way, we have a IP, which is defined to your VM, and we-- all we needed to do was,
24:12
register that IP, as a, in DNS. That way, whenever the VM owner wants to talk to the VM, he-- it's, it's, it's already up to date. It's already there, and they can just, you know, use their FQDN, and then they can get into the VM. They can-- no ports attached.
24:28
You can do whatever you want. You wanna do SSH, you wanna ping, you wanna run a HTTPS service, w- all kinds of stuff, right?And, it worked, theoretically. It worked when we practically, you know, did our testing. And then we found out, okay, let's try migrations, right?
24:46
We obviously with Portworx, we, we get storage migration, we get, same vMotion that we have in VMware. But we found out that when we migrate, the IPs change. And when the IPs change, we need a method to be able to update our DNS records, right? So that's a challenge now.
25:05
Now the IP is gone, I cannot reach the VM. So then we developed, an operator which will talk to the DNS at any time, and it's a, it's a Kubernetes-native operator, so it, it looks at all the events, and any time there's a vMotion or some migration, not vMotion, but some migration that happens within, KubeVirt, it kind of immediately talks to DNS or API and then updates everything.
25:28
Everything was instantaneous. But then it was, DNS, TTL and caching, right? That was, again, another challenge that came up. So, so these kind of challenges did come up, and then we discovered, maybe make our IP static.
25:44
That way, when we are migrating, you know, we make sure the IP stays stable. We then hit a challenge with Cilium, that Cilium's principle says this can only be one IP for a pod, right? So even it's migrating, when it's going to another node, it's gonna stick to that IP. So, then we started evaluating different, plugins and different offerings that is
26:06
available from open source, like KubeVirt and different TNIs. Do we need to go there? Should we leave Cilium's, great functionalities and features? So we had some, re- whiteboarding and, we c- we decided, did some prototyping, and we decided to use multiple NICs.
26:25
Keep what Cilium is using for its, primary interface and kind of start using, secondary interfaces for stable IPs. And that kind of clicked. That kind of, also made sure the IPs remain stable. They're registered with a stable IP all the time on the DNS.
26:44
And we also get the, you know, benefits of, migration using, you know, Cilium as our CNI. And you, you mentioned this briefly, and this can be for, for you or Greg, you know, Portworx Enterprise and kind of the role plays and things like storage vMotion or vMotion or, or live migration type of capabilities. Can, can you talk a little bit more about, you know, kind of the role it plays in providing
27:08
some of that storage and data management, those workflows that you really need that you're used to in VMware that you now are kind of translating to a Kubernetes world? And this can be either of you wanna just comment on that briefly. Sagar, talk about Well, we might have another slide about it, but there's, reality of moving, storage around, and we're starting to use that now with this system.
27:28
We didn't really have that ex- except for in, in VMware. I wanna-- I j- I-- Just one quick thing, but I'll give it back to Sagar here. I can't emphasize enough that the work that Sagar, and in particular, there's a guy on our team, Joon Park is the guy, that have pioneered some of the routing stuff doing with Cilium and the challenges that we found.
27:47
These are all issues where most technology, including a lot of this, is not as dynamic as you would think. As we build more and more ephemeral workloads that come and go a lot or move around for failover or whatever, we find, you know, gaps and creaks and groans in the systems. And so the team's out in front of that.
28:03
And so I have to celebrate that we've done a good job here about tunneling through these kind of problems towards a more, dynamic environment. So it's, it's quite impressive and, you know, betting on hardware-based networking is gonna pay off for the, you know, especially through. But Sagar, explain maybe now or later what the deal is with, with, with Huawei
28:21
doing the equivalence or the moral equivalent of, of vMotion. Yeah. So like, like I mentioned before, when we do migrations, right? So that's the notion within Kubernetes, the pod, the IPs change, and then during migration, there's some functionalities or feature sets and requirements saying, "Yes, you need to You can-- My piece would change,"
28:45
but, you have, a load balancer, right? So you have a STaaS- a VIP that you can always access through behind. But that's true for containers. That's true for workloads which live in containers today. But how would you manage VMs?
28:59
So that was the challenge and, during vMotioning, in vCenter, and VMware is pretty, pretty good about this, right? The-- It's pretty fast. You have your dedicated vMotion network you can assign. KubeVirt kind of gives us the same feature set as well.
29:13
Like, you can do the same thing. You can assign a dedicated, migration network, between hypervisors, between Kubernetes nodes. You can, you can also, Then also there, there's also, IP address management. There's an IPAM, that comes within Cilium, which is homegrown within enables you to, you know, dedicate IPs to workloads, through automated fashions, and
29:37
that way we can seamlessly migrate, VMs without losing any packets, like how we do in, VMware today. So those kind of, those are the some of the tests that we had to run through, with our, end users and, and end users were also pretty, you know, good about this. They were giving feedbacks, "Hey, today I was just doing an associate session and I observed,
30:01
like, you know, my session dropped. What happened?" And we kind of found out, "Oh, you were migrated to a different node." oh, maybe because we haven't implemented stable IP yet. So those were kind of the things that I think we should look out as, Also to our customers, I would say is, have some early feedbacks.
30:17
That way you're not, you know, driving on a, a tunnel with no visions, right? Like what can happen, outside the world, like how would we treat something like these? Like there, there could be scenarios that would come up in front of you, so you have to re-architect and you have to pivot. So that kind of gave us more and more challenges, but we were able to accomplish and,
30:39
you know, sail through most of them, pretty, pretty, pretty, I would say, um-Uh, in a, in a fast manner. And I, I wanna move on to, to the migration component in a second. And so I'll let you save your voice, and I'll, I'll take, I'll take this slide to talk to you cause I'll, I'll be coming back to you in a second.
30:57
But I think, you know, what, what the really came to be from the architecture it matched up with the design principles, right? By, by building it on Kubernetes and KubeVirt, it provided that open ecosystem-based platform. You've been able to, you know, have the scale you need.
31:14
With KubeVirt, you can both run large workloads greater than terabyte, but also smaller VMs and containerized apps, including the ephemeral ones. You know, with things like, you know, the automation of Kubernetes, but also adding on things like, Portworx Enterprise, which automates a lot of the storage and data management tasks. You talk about things like, you know, vMotion
31:33
and, and backup and DR and some of the other components of, what Portworx Enterprise provides. You've been able to really dive into the, software-driven operations. And then, of course, that flexibility of both, you know, Portworx to manage data anywhere, but also Kubernetes to run anywhere, that's on-prem or in the public cloud.
31:49
You know, that architecture really hit home with those four design principles that, that you're after. So, okay, I-I've let you breathe for a second, Sagar. I'm gonna come back to you now 'cause I wanna talk about migration. And, and really, I think this is where a lot of people listening, especially, you know,
32:05
this is a tech talk, we wanna get a little bit in the weeds. Talk to me about this process. I know it's not always straightforward, sometimes it can be quite painful. Talk about, you know, the, the discovery and analysis, then the workload classification, and then, you know, how that turned into some validation and, and the
32:20
actual migration itself. But let's start with that, that first area, that discovery analysis, how you, started going through that VMware footprint. Sure. So yeah, we had the, we had the goal in front of us, right? So move off of VMware.
32:34
So how do we get there? And, and workloads are pretty, they could be of different sizes a-and varieties, right? They could be some test, test workload, they could be some production workload. Within them, we could have some different kind of, you know, colors, of workloads, which are pretty CPU intensive, some disk intensive.
32:55
So how do we, you know, categorize them? Definitely, we need to make sure that, we reduce the downtime, right? We don't wanna make sure, whenever we are migrating, how are we gonna choose between, hot migration, warm migration, cold migration? Is, is this okay to, you know, do this, with some downtime or some window so that we, our
33:17
end users are not affected? All of those played a role in categorizing, and, you know, making sure what do we choose and how do we, pick a guinea pig, right? So we started, talking to our end users, and we found out, like, most of the footprint in VMware was pretty much, I would say, workstations and developer VMs.
33:40
So these VMs are, kind of the jump hosts or the go-to, virtual machines that, developers use on a day-to-day basis, right? So this gets them into the data center, and then they can talk to other, workloads within the ecosystem and data center. So that was pretty big footprint, almost, forty percent of our entire, workload.
34:02
And we looked at if we gonna go with this particular classification or this particular workload, we're gonna save a lot of cores faster, right? And then we also kind of thought, are all these the same? Like, are they all of same operating systems? Do they have same kind of, requi-- specs so that, you know, we can, draft our target
34:25
Kubernetes-KubeVirt cluster? Like, what ki- what kind of nodes and what kind of, storage, backend storage, and networking we need for that kind of infrastructure. And then we decided, okay, since these are not, like, very critical, they are developers, but not for the company, like, not for the production workload.
34:45
So we'll start with small, and we-- maybe we'll start with, you know, some sort of, you know, pilot users, and we can just see how we can, you know, migrate them, right, to this workload. So then came, um-- So with that, with the workload class-classifications finished, we then was-- wanted to start with, okay, are the VMs and these are the, you know,
35:06
source, clusters which are spread across, two data centers. We have, Utah and here in California. So we were looking at, now we need to do sizing, and well, now we need to talk about, you know, data center capacity, power, cooling. Do we have them all in one rack or separated across different racks?
35:27
And we need to talk about, think in mind, high availability as well, right? So if one rack goes down, make sure that the VMs can, can and powered on, on the separate rack. So those kind of high availability and, some kind of, I would say, hardware-related constraints were one of the, talks that to go through and then decide before diving in.
35:52
And if I think about, actual migration, right? So now we have the, hardware, we have, the, the landing zone, we have the source, everything is aligned. We have the networks, we can reach each other. So how do we migrate them, right? How do we-- Do we-- Can we do, hot
36:12
apparently, there are different techniques, and we'll go through this in later slides, like how we can migrate them. But, today, as of now, there's no, like, a hot migration available across platforms. Um-There's, there's some leeways where you can do kind of like a wall migration. You can start some snapshots and start them, and then during the cutover, you can
36:34
just do a final snapshot, and move them. But there's still, like, a small, downtime expected. And then it also depends on what kind of workload are we doing. So if this is a developer VM, maybe we can do this when all the, end users are maybe we can come up with some time zones so we actually discovered and coordinated with
36:55
our, team support team and the developer team and our current administrative team in India, in Prague, as well as here in US, to make sure that when we do these migrations, make sure the developers are offline, or we don't do their-- do this during their work hours. So yeah, good kind of scheduling, and that kind of paved our way into the waves that we would use for our migration, right?
37:18
So this is wave one, wave two, wave three. We'll hit, EMEA time zone. We'll hit, AMR time zone. And that's how we came up with the planning. Yeah. And I, there's a couple of key things here,
37:32
and, and Greg, I wanna get your, your, you know, kind of thoughts on this too. But it's very clear that this is a journey. This is not something that happens overnight. This is not something that, you know, happens even within just a month or two. This is a, this is a journey, and being intentional about that journey is key.
37:46
You know, you talked a little bit about starting with the tier four developer VMs. Big footprint, so significant cost savings by decommissioning those cores and, and putting them on the new platform, but also gave you an opportunity to prove out and validate the system before moving to those more heavy production workloads. So it, it really is a journey, and Greg, I wanna get your thoughts on kind of that
38:07
migration journey as a whole. Yeah. I-- Some of the talks you have in these kind of things are do you move the high-risk workloads first or the low risk? You know, what do you learn now vs later? You know, us choosing these Actually, these dev VMs are m-much more than a
38:22
They're, they're large platforms. They run application software and simulator environments, and as a result, we're building large images. There's a lot of traffic in and out of these machines, and they're very diverse. We're a development company, so everybody's doing experiments on their dev VMs.
38:36
And so while they're all, you know, x eighty-six, networked, there's lots of different services running on these things. And so, so yes, they're relatively simple to move, but the reality is about all their services and connectivity have been, have been big work. So this is a good, a good first wave.
38:54
In that for, for other kind of production workloads, let's say there's some common services that we run on the inside. They'll come, they'll come later, but-- and those are decisions are not necessarily as much of, you know, a forklift move. Some of them are re-platformed.
39:07
We'll take some of the apps, and we'll move them into Kubernetes. They'll be container-native. And some will be just moved over, some will be maybe moved to Amazon or some from Amazon moved on-prem. I mean, so we always have motion like this anyway, and, and this system is allowing us to, to, you know, address each particular workload
39:24
as, a-as needed. So we will have You know, as the system gets bigger, we're gonna, you know, find if there's limits and the creaks and groans or we didn't understand things and, and it's a, it's a ramp, it's a journey. I-- This kind of journey is, you know, at least a couple of years, while everything is moving.
39:42
So, so, so far so good. It's been You know, we found real issues, and boy, with open source, that's where it paid off because it's like we could figure out what the root cause is right away. So, so we're, we're in the journey right now. Sagar, question about, you know, for those tier fours, are we That, that's
39:57
short-term project, isn't it, from start to end? We'll be done with this thing pretty quick, right? Yes. So I think we're pretty much done with seventy percent of this, tier four environments, now we're focusing on moving on to tier tier two workloads.
40:12
And what we call them tier one, tier two are the high, visible, highly, efficient and, like, I would say, high-priority VMs production workloads, which kind of, k- like, like, for example, DNS, right? And I would say Artifactory or your, build tools and how, with where they're hosted. So those are the pretty big ones that we're going next with.
40:40
So like, like Greg said, when we move along, when we plan for these things, something new that we learn every time. Like, okay, so this is a new requirement, so how do we pivot? How do we make sure our, target infrastructure supports it? So there could be, some new, I would say, feature set that, new versions of Kubernetes
41:00
and new versions of KubeVirt come out with, and we would like to test them in our dev and staging so that we can make sure, yes, is a perfect solution that we needed for this kind of workload in production. Yeah. Yeah. I, I'm gonna move on in a second to the, some of the tooling that you use to help with this journey, but before I do, I just-- I wanna,
41:20
you know, Greg, back on, on your point, this has been a really impactful, opportunity to go classify in that discovery those workloads that, like you mentioned, to not only say, Here are the VMs and let's tier them for moving," but also maybe we retire some of them. Maybe we, you know, decommission them but build them, net new. Maybe we move to, like you said, you know, a STaaS version of the, of the product as opposed
41:43
to just, you know, maintaining it ourselves on-prem. So there's a lot of value even in just that discovery and analysis that goes beyond the migration itself. Yeah. I, you know, I've been doing this for a while, and every one of them does their and track and have an inventory of what we have.
41:58
And, and you miss a few things, and we're no different here. And, one of the things that this team is developing here is this muscle memory for these so-called replatforming tasks. And so that means it's a team that's got better and better awareness about what the whole estate looks like.
42:15
We're building and tearing down crazily fast on all, all kinds of things. And so having those tools and that knowledge or that expertise is really, really helping here. So Sagar's teams, you know, they, they dig into the, "Oh gosh, you didn't know about this," right? And, and, and it's, it is no different for us.
42:30
We're just rapidly responding.And, and, and not only, you know, are they, are they digging in, I know Sagar, you and your team have, you know, explored a few tools that have made the process easier. And one of those, i- is around Xcopy and some of the things that, that allow more rapid migration.
42:49
Talk a little bit about kind of where you started with the migration over the network- Mm-hmm and then the role that Xcopy in facilitating a faster migration, as you- Mm-hmm moved through the journey. Sure, yeah. So I wanna reiterate on what I said earlier that, you know, there's no, there's no hot migration currently available between platforms.
43:08
So we kind of, go back to thinking, how do we make it faster if not We know there's a downtime, so how do we make sure that, you know, it is, feasible, the application can take it, maybe we can schedule it, maybe a few minutes is fine. And if it is, you know, load balanced between different workloads, can we do one VM at a time, and then the, the second one after one's done?
43:29
Things like that, right? During So when So that journey also started with, not going with the advanced, what you see here in the screen, right? Ma-make sure using Xcopy feature set. So we started off with something where we were not doing storage-assisted migrations. We, we did the traditional, you know, way of migrating the disks from VMware to
43:56
KubeVirt, over the network using the host networks. So from VMware, it goes to the host network, and in KubeVirt, Kubernetes, the top of the rack switches. And for that to be fast enough, you need to We, we kind of thought about how do we make sure our uplinks are faster, make sure our switches are faster?
44:14
Do we How, how do we come up with, you know, a way so that if a 1 TiB worth of VM needs to be f- lifted and put into KubeVirt, how much time it's gonna take? So if it's gonna t- if it's gonna going through same data hole, same rack, same data center, it's not a big problem. It's We're gonna get a good speed and bandwidth there.
44:33
But what if, we're going across data centers? What if we're going across, different regions, right? So, and, being a bigger VM, definitely we're gonna have like, really, really, bigger downtime, and with bigger downtime, maybe it's not affordable for the end user. So that kind of, you know, made us to look for other alternatives, how we can make
44:56
it faster. So definitely, snapshotting and then, you know, using those snapshots and then, you time, when we schedule a downtime, we those snapshots and recover VMs on the target side. That was one option. And then we, we found out, you know, within Pure, like, there's a dedicated team which worked with the forklift, tool by Red Hat to,
45:20
use the storage-assisted, migration feature. It's called Xcopy. So with this, what we can do is we can have the even one TiB disks, which would take close to, depending on your network which was taking almost, five, more than five hours if we are in separate, regions, to down to, 40 minutes, right?
45:49
So we were like, "How is that possible? How are we doing this? So what's the technology?" So basically, Xcopy, with Xcopy, when you use, FlashArray as your source as well as FlashArray as your destination, you kind of create, replications between them. So your volumes, which is VMware volume, gets replicated to your, Kubernetes
46:11
FlashArray. And within that, once it's replicated, within forklift, you can use the Xcopy feature, and it will actually, do a storage-assisted cloning, which is much, much faster because the data is already there, right? So you're seeding it, it's already there in the destination FlashArray, which KubeVirt, where, where your KubeVirt VMs can, Pods can utilize them and convert them into PVCs.
46:36
All you need to do is, you know, wait for the exact timing, and then you, specify, "Okay, I'm ready to cut over." and then Forklift will actually create a hot clone, and as soon And the clone depends on the size of your actual data. And even if it is one terabyte of worth of data, how much is the used data? If it is 500, 600, depends on, depending on that, the speed will decide
46:58
It, it decides the speed. But we saw, which were taking hours, was down into minutes. And if it's, smaller, even faster. If it is not many used data, even faster, right? So this kind of, was a breaking deal for us, and we immediately, started using
47:17
Forklift with Xcopy storage-assisted migrations. But there were some, I would say, prerequisites as you, as, as I mentioned before. You need to have replications going on, so obviously you need to have a replication network between these FlashArrays.
47:34
So there was some time spent for the architecture. And also there are different feature sets that you, you wanna go for, right? So if you're using revaults within VMware, you can do, what you call is an active DR, like a synchronous replication with faster uplinks available. That is instantaneous.
47:52
That also really uses your, time to migrate and come up on the KubeVirt side, right? So there are different options we went through, and then finally we stuck, okay, this is perfect. This makes sense. Even with larger VMs, critical VMs, we can have them migrated within few minutes.
48:10
So our downtime window drastically reduced from hours to minutes. And I, I want to, I wanna make sure I leave a couple minutes at the end, for questions. So I, I, I'll give you kind of 30 seconds just to talk about this. But you alluded to this already once. I know you also, in addition to using forklift and, and Xcopy, you took advantage of an
48:30
internal tool that you built to help really, streamline and address, the migration. Just talk, talk for a minute about this tool and kind of the role it's So this is a tool, that we built out, just a UI to give everybody, even our the management basically, where we are, how many cores have we saved, how many VMs have we migrated.
48:51
Kind of reporting tool which started built out, from scratch. It's pretty much built, in React and FastAPI. It's a three-tier application. It's surprisingly living in our same Kubernetes cluster on-prem. So this reporting tool kind of tells you, like, how many have been migrated, what how do you
49:08
wanna plan. You can create batches here, and things like that. But then, I took it over in Ash and said, "Why not use it for migration as well?" You We can It's Kubernetes, and KubeVirt is API-driven, so why don't just click a button to migrate a bunch of, 50 VMs and see the magic happen?
49:25
So we were able to do that with this tool. So if folks are interested, we can also open source it, and everybody can utilize it. Yeah. Yeah, and I, I think it's a, a great example of kinda some of the innovation that comes along the way with, with this type of journey. And, and you know, I, I'm gonna jump into outcomes and next steps so that we have a
49:45
couple minutes for questions, and just talk about kinda where we are in the. So, like we said, it is a journey. It obviously is not, anywhere near complete. But even with that being said, we have moved about 20,000 cores and, you know, about 300 to, to, you know, maybe even bumping up a little more, migrated per day.
50:04
So moving at a rapid pace and continuing to expand. I think a key thing here too is th- only a handful of folks working on this project, so it may seem like a big task, but with some of the automation and some of the other things that you've been able to get to, now you're in a place where it's not necessarily taking, you know, a, a team of 50, a team of 100 to get done.
50:24
You've got eight engineers who are, you know, focused on, on driving this project. And importantly, you know, we talked about this, and, and Greg, I do wanna get your thoughts on this, the governance piece. The Not only the developer of self-service, but also being able to automate some of the governance that comes along, with that to ensure that those app teams, those platform
50:42
teams now have, adherence with some of the key governance requirements within the organization. I think that's become very important as an outcome. Can you 1touch on that just for a minute? Yeah. As our company matures, a lot more attention to details like, you know, business continuity, a continuous security over time,
51:02
capacity planning. A lot of, a lot of higher level functions which we, as a s- a relatively young company, have always dragged along a little bit. These are, these are being added as we go. So while we're under the hood on this thing, we're also putting in strong access controls and delegated authorities and, and backup storage and all these.
51:19
And, and we had them before, but we at the level that we have them now is getting much, much closer to a very mature enterprise. And so that's, that's, that's something we have to do, and we've, we've been able to do it not as separate projects, but as part of this. And, and I think, you know, there's some real operational impacts that have been felt
51:36
already, and, and we touched on a couple of those. But the governance, the self-service, the ability for developers to provision directly, and also faster provisioning. So things that used to take, you know, quite a lot of time, whether it's because they have to go through a ticketing system or have to go through, a, a longer provisioning process now
51:51
can take minutes. And, and that development team really can, you know, embrace self-service and, and, really, you know, focus on what they're doing, which is building and, and developing. Just to touch on a couple of the, the other outcomes. So, you know, in addition to that th- that unified platform has helped
52:08
some of that platform sprawl, and I think the number we talked about, before something like 90% of apps will eventually land on the platform. So the, you know, about 10% you've gotta maintain. They may be in some other places. But that platform sprawl is, is really reduced as you expand the platform to higher
52:24
tier workloads, to hybrid applications. And importantly, you know, as we touch on both from the, the Kubernetes side but also the Portworx side, that intelligent day two operations is, is critical in helping to, you know, reduce the drag on, on your team, on the, you know, infra teams, on the internal teams, so developers can really focus on, on building, a- and, and doing what they do best.
52:46
Before I 1touch on, on questions, I'll leave the five minutes for questions, I do quickly wanna just get each of you, I'll give you maybe 30 seconds each. Talk about some of the lessons learned and, and we've touched on some of these bit, but, you know, give me maybe your top lesson learned through the journey, something people should walk away with a- as they look at this journey for themselves.
53:03
And, and Greg, I'll start with you. Yeah, sure. Thank, thank you. Well, it's all about the networking. It always is on this kind of computing. And so, y- you know, you think you know it until you actually load it all up.
53:15
So we've, we've basically Y- you heard some of the journeys on data migration. Those were issues, but also just network configuration and, and going into this world of more and more dynamically configured workloads, not just statically configured. Address that and, and be part of it and, and build for it and that's, that's what we're doing. Yeah.
53:32
I want- Sagar? Yeah, I wanna inter the same thing. I think automation is, like, pretty critical, for not just starting your journey, but also to sustain and make sure your day two operations are handed out smoothly, especially when you have multiple teams involved, right? So you have, you have a separate support team who's looking over your cluster or be it, a
53:56
new upgrade that you wanna do. If it is automated in the right fashion, you would expect less hurdles, and I would say, you know, a faster efficiency in your, day-to-day life cycle. So with that, I do wanna leave a minute for, for questions or five minutes for questions. I'm gonna stop sharing for a second, take a look at the, at the chat, but
54:21
for, for any additional, questions. And, you know, I'll start, I'll start off with one, for, the two of you while, while I wait and see if there's questions. Talk a little bit about, you know, again, and, and I, I wanna go back to the day two operations to, to Portworx and some of the role it plays.
54:39
Talk a little bit about the importance of automating day two operations and, and how that really has helped you scale the platform and also provide real benefit back to, the development team. And, and Sagar, I'll start with you on this one.Yeah, sure. I would say, first of all, like whenever What-- This is a great example I can talk
54:59
about, is, you get these alerts about volumes and your, disks filling up, right? So within VMware, there are also, when it really integrate with, outside, networking tools or, monitoring tools, you get to see, okay, my VM is about ninety percent full, so I need to raise a ticket maybe with some team, and somebody will take a look at automate it, maybe, someone will get it faster, right?
55:22
But, the, the, with Portworx Enterprise, with the feature of autopilot, you can specify the threshold, and it will automatically, you know, upgrade it for you. If you, if you, if it's a critical production workload, it can increase by five percent, ten percent. All those can be done and tweaked within Portworx Enterprise, which was pretty good
55:42
feature for us. You know, saves a lot of engineering works, and they do, support teams works as well. And, you know, the, the other question too I-I'll bring in is around, you know, we talked, about this a little bit in the workload classification. As you're going through and looking at what app to migrate or refactor or re-architect,
56:04
how are you making that decision? How do you make the decision of, hey, this application, let's go ahead and re-architect it and, and containerize it and, and, not even worry about, about transferring or, or even retire it. And Greg, talk about that maybe a little bit, the mindset of how you decide those different buckets to put the workloads in.
56:23
Well, it, you know, it, there It depends on the team that's available, who the product is for, where it's in its life cycle. So we, we have a shared effort with the teams that own an application. Say, "Look, is this time for you to move to Kubernetes? Is it not?" You know, is the stuff got third-party libraries that are really kind of
56:42
stuck, or is it something that's open source? Do you wanna combine that feature and deprecate it with some other feature? So these discussions for each one of these things are, are, are happening. And so we wind up with, you know, okay, this app is just forklift, this app is moving to Kubernetes, this app's being combined.
56:58
You know, and, and that's, that is We're getting good at that because there's a lot of those discussions are pitching over and over again. I mean, an important part inside here is that there's, we'll say table stakes in all cases is all these workloads shall be software defined. So some of the very old ones have been handcrafted.
57:13
People, you know, root login as a machine and then twiddle the machine We're not doing that anymore. This is far, far more software defined, which is gonna help us with, you know, future migrations or workload identification. So we say, "Look, we got to do this software defined.
57:27
We have some old workloads. We're not quite sure what they do. What are we gonna do here?" And so this is the story has been much more about, at scale dynamic configurations and, and that's, and that's everybody's on board with that. And the, the last question I want to ask, you know, maybe thirty seconds before we have to wrap up here. You know, w-what does the future look like?
57:49
We've touched on it a little bit, but what i- what is the future? And, and Sagar, I'll talk about it with you. You know, five years from now, where do you hope, Everpure is in their journey, and, and what does that, what does it look like, within Everpure in five years? Well, definitely we would be looking at, investments of, on KubeVirt heavily.
58:09
We would be looking at, m-migrating pretty much all of our heavy workloads into KubeVirt, assessing each one of based on their criticality and, the visibility across the platforms, right? And I think we would be at a point where we are close to ten percent of what our current VMware footprint is, and maybe the same thing for OpenStack. And we would, might not-- we would also be
58:36
utilizing Portworx's, coming off new features and Storage Vault. That would be, I would say like, a new, storage journey that we would, you know, g-go through together. Yeah. And, and Greg, quickly to you, any- to add to that as we wrap up here?
58:59
Yeah. Hey, you know, the world is evolving, so other CPU architectures, definitely GPU scheduling is coming with this thing, solutions between cloud and on-prem. These are all, these are all in our journey. Well, with that, I know we're at the, the top of the hour, so I do wanna just point out two
59:16
things as we wrap up here. First of all, thank you very much for spending time with us today. If you're in Europe, a-as I am at the moment, I am I'm here getting ready for KubeCon next WEKA, in Amsterdam, twenty-third through twenty-sixth. Please stop by. We'll be at Booth four-fifty.
59:31
Learn more about our journey, learn more what we're doing, broadly around the, the virtualization space. We have a specific VM on Kubernetes Day, event happening next Monday. So if you're, if you're in Amsterdam, please reach out. Please, reach out to us if you want, to join us for that event.
59:47
And we hope to see you there. We also, next week, if you're in the, US, we'll be at RSAC, talking about security, talking about, what Everpure is doing around, you know, eliminating non-disruptive outages, accelerating threat detection. So stop by Booth two four four nine, to, to learn more there. And last but not least, we've got Pure//Accelerate coming up, June
01:00:08
sixteenth to eighteenth. We'll talk about this and many other topics there. Our registration for that just opened, so feel free to go ahead and register at the QR code. We hope to see you there, to dive more deeply in some of these discussions. So with that, thank you again very, very much for joining today.
01:00:24
Please join our community. We'll continue having this discussion there. We'll continue talking about, all things around KubeVirt and Kubernetes and modern virtualization and, the trends that are going on. And thank you again very much for joining us, and we hope to see you again, soon, for another Tech Talk. With that, thank you very much and have a
01:00:42
great day. Thank you, Andy. Thank you.