Truth in IT
    • Sign In
    • Register
        • Videos
        • Channels
        • Pages
        • Galleries
        • News
        • Events
        • All
Truth in IT Truth in IT
  • Data Management ▼
    • Converged Infrastructure
    • DevOps
    • Networking
    • Storage
    • Virtualization
  • Cybersecurity ▼
    • Application Security
    • Backup & Recovery
    • Data Security
    • Identity & Access Management (IAM)
    • Zero Trust
    • Compliance & GRC
    • Endpoint Security
  • Cloud ▼
    • Hybrid Cloud
    • Private Cloud
    • Public Cloud
  • Webinar Library
  • TiPs
  • DRAW

Verge.io: GPU Virtualization Without Specialists: VergeOS + NVIDIA

VergeIO
07/30/2026
0 (0%)
Share
  • Comments
  • Download
  • Transcript
Report Like Favorite
  • Share/Embed
  • Email
Link
Embed

Transcript


management. But we have some special folks here with us today to help guide that discussion. Of course, I am Aaron Richman, I am a field evangelist here at Verge IO. And we have Jimmy Rutella with us, my co-host today from NVIDIA. He'll tell you more about himself. We have Paul Hodges, Verge IO's field CTO. And we have George Crump, who is our chief marketing officer and a longtime industry analyst as well. In fact, guys, can we just do some introductions maybe in that order? Start with you, Jimmy. Yeah, sure. Hey, everybody. Jimmy Rutella here. So I'm a senior solutions architect with NVIDIA. And I focus specifically on our VGPU product. It's been great working with the Verge OS team on this integration and really excited to show you guys what we have today. Awesome. Thanks, Jimmy. Paul? Hi, everyone. Paul Hodges. I'm the field CTO with Verge IO. I've actually been with the company since day one. So going on about 16 years, I've kind of made my way around the product from development to engineering and, you know, wear a lot of hats with the company. So excited to show you guys what we've come up with here. And I'm George Crump, as Aaron said, chief marketing officer prior to joining Verge. I was an industry analyst, ran my own analyst firm called Storage Switzerland for about 14 years. I also want to add that we refer to Paul as P4. Feel free to do the same. And also, I was going to say he was being a little humble. He is a key component of this particular, I don't know what we call it, capability, let's call it. And that's why I'm really glad to have him on the line. So P4, thanks. I know you're busy developing the next great thing. So I appreciate you taking the time and working with us today. Absolutely. Thanks, guys. You know, what are we going to talk about in general? Well, we're going to talk about some things that might help you change the way you view the concept of GPU management. I've got five sort of bullet points to give you a sense of where we're going. We're going to talk about the problem first, what gets in the way. We're going to talk about a very exciting brand new NVIDIA vGPU20 that has just recently been announced. And Jimmy, of course, is the subject matter expert with regards to that. We're going to talk to you a little bit about the Verge OS architecture, which some of you may already be familiar with. Many of you will not be until we talk about it here. And then also P4 is going to do a live demo for us of this vGPU20 supported capabilities that we bring to the market. And we'll have some time at the end for scenarios, questions, and answers. We do want to ask you to keep this interactive, right? You may have thoughts or questions along the way. Feel free to use the chat. Let us know. Some of you already have. I see folks from all over the world. That's pretty awesome. The UK, South America, California. Let us know where you're at. And when you have questions, just throw them right in there. Don't wait. And we'll try to either answer them via chat or perhaps live, et cetera. There'll also be some meeting links throughout the webinar. Feel free if you want to schedule some time to discuss this or other aspects of your infrastructure and what we bring more privately. We love doing that. So please take advantage of that. And we're going to have some poll questions, some polling questions. I'll mention them. When you see them, please give us a sense of just answer them if you can. And so that'll help guide the discussion a little bit as we move forward. So thank you for that. And all right, speaking of polls, why don't we jump right into the first one to help us kind of get going on the right foot there. Question one, right? We're looking to see what kind of hypervisor you're currently running. Are you using GPUs? And if so, tell us about the use case there. Then we're going to move on and we're going to start talking about the problem, which you guys are probably just as familiar with, many of you are, as we are, because you're living it. All right. So the problem that we want to talk about, right, what is the problem? Why every IT team is stuck and what it's costing them. And trust us, it is costing. And so we're going to dive into that here a bit. Let me, there we go. So many teams, many teams, right, they're being asked to embark on GPU projects. But George, maybe you want to kind of frame the problem a little bit here for the folks that have joined us. Yeah, I think, you know, obviously there's a hardware cost to a GPU. But what we hear, I think most customers know that there's going to be a hardware cost to the GPU and are comfortable with that cost. What I think takes a lot of customers off guard is the, is all the other stuff, the driver management, setting up MIGs, which we'll define in a little bit, but essentially the ability to divide, if you will, hard divide a GPU up, the tooling is a little fragmented. And I don't really have it on the list here, but the, and then doing all that in a virtual environment underneath the hypervisor or with a hypervisor adds a layer of complexity that I think a lot of people struggle with. And so I, the result is really, there's two results that I typically see. One is people give up, maybe there's three, they give up, they don't deploy GPUs as widely as they might, because they feel like each time they deploy a GPU, it's kind of like starting from scratch, it's another exercise, or as the slide says, they hire a specialist to do that. But Jimmy, I'd like to ask you, I mean, you're probably more in the trenches on this stuff than I am, is the operational overhead one of the bigger challenges to this? I think it weighs into it. I think the biggest thing is sometimes people overlook or are not aware of the vGPU functionality and they look at it as traditionally pass-through was the only option and dedicated resources to a high-end data center GPU is often overkill for a lot of organizations and what they need to do. So, awareness of the technologies and then the technology itself is a bit complicated. And you guys do a great job of simplifying that. Yeah. And just going back to our poll question real quick, I misquestioned what the results were on question two, but question three is interesting. We asked kind of how you either are or would use it, and it's interesting, it's split pretty evenly between VDI, AI, data, well, VDI and AI are probably the two leaders and then data science and visualization right behind it, right? So it's a pretty even spread. So we'll talk about some of the different ways you might leverage this as we go forward. Go ahead, Eric. Absolutely. And I noticed earlier in the earlier polling question, about half of you, which does not surprise me based on the title and theme of today's webinar, are either using GPUs today or testing, piloting, et cetera. So you've come to the right place for sure. All right. So let's advance here. So we have some stats just to kind of give you a high level overview of what we're seeing and hearing that there is a real expertise gap and that that is an expensive problem and a hindrance to many of probably of your organizations and certainly your peers, right? 73% lack of in-house GPU experience. Time to production is sizable and the cost, right, of a specialist, you're a specialist for a reason and it comes at a premium. I think we have another, yep, we've got another poll up there regarding GPU expertise. So feel free to let us know what your thoughts are there. We talk about hardware, we talk about in-house expertise, et cetera. And I think the results are kind of what we would see, right, is its expertise in hardware. I think hardware is interesting because a hardware, the cost of the hardware can be offset if you can use it to its fullest extent. And I think really, you know, Jimmy, I'll go to you again, but I think that's one of the reasons you guys have put vGPU and MIG out there so that people can fully utilize the full capabilities of the hardware. Is that a fair guess on my part? Yeah, 100%. I mean, it's all about that management of the investment accordingly and the flexibility that comes with it. You know, being able to allocate whatever resources you need as needed and being able to adapt those and change them as needed too becomes even more important, especially in an environment today where, to your poll question earlier, you know, it was pretty evenly split amongst what people are using those GPUs for. And many people don't realize they can do all of those things at the same time and oftentimes are. It's not just one use case anymore. Right. Yeah. Right. Awesome. Thank you. We're in a part of our discussion where we really want to hear from Jimmy first and foremost about this brand new, this exciting NVIDIA vGPU 20, you know, what's new and why does it matter that we are a supported platform and that we are working engaged together in this endeavor? So, Jimmy, why don't I turn it over to you at this point? Yeah, sure. So before I jump into the new features in vGPU 20, for those of you that aren't familiar with what vGPU is at all, I want to give you just a super high level overview of what it is. So vGPU is a software that we sell that gives you the ability to take a data center GPU and fractionalize it into smaller pieces and share that through your hypervisor to your individual VMs. You can do this for a wide variety of reasons, you know, mentioned some of them earlier, basically, you know, visualization use cases, AI use cases, even just for general knowledge workers, giving you the ability to offload that streaming of pixels to the GPU instead of your CPU, giving you a more smooth end user experience for things like, you know, video conferencing, streaming video, multi high res monitor support, things like that. So that's indeed, you know, that's sort of the core of what the tool does. It's a very mature product, you know, vGPU 20, it's the 20th release of it. We add, because it's a software based platform, we have the ability to add features and new things to it over time. Definitely one of the biggest announcements with vGPU 20 was the support for Verge OS. This basically gives us the ability for Verge and NVIDIA to work together, should issues arise or, you know, give them the ability to add new features quickly and more easily. I can tell you from my personal experience, the first time that I saw Verge OS, I was blown away by the GPU integration already baked into it without our support officially. So yeah, that was probably the biggest eye opener for me. So that was very cool. A couple of key things, though, that were added with vGPU 20, in addition to Verge. So if you're not familiar with our RTX Pro 6000 Blackwell Server Edition GPU, I know it's a very long name, but it's a very cool card. So it is the first GPU that we've offered that includes what we refer to as universal MIG and vGPU support. So if you're familiar with some of our A100 or A30 cards that are compute only cards previously, those featured MIG, MIG stands for multi-instance GPU. And what that is, is the ability to take one big GPU and at the hardware level, fractionalize it into smaller partitions. Those partitions are then presented to the hypervisor as if they were individual GPUs. So this gives you true multi-tenancy at the hardware level. And then on top of each of those MIG partitions, you can then further subdivide through vGPU on top of it. MIG eliminates some of that noisy neighbor situation that you would potentially get, especially in more dense deployments. For example, with the RTX Pro 6000 specifically, if you were to use just vGPU, just software fractionalization on it, you would only get a maximum of 32 users that could share it, 32 VMs that could share it. But by leveraging MIG plus vGPU together, you can actually get up to 48 VMs that would share that. So lots of benefits to MIG and vGPU. A couple other features that were released with it with respect to vGPU on KVM specifically. We have MIG and time slicing, so MIG plus vGPU on top of it. We did also add Wayland display server protocol support for any of your, if you're running Linux desktops. This basically replaces the older Xorg protocol that was used on those Linux desktops, basically giving you a more secure, smoother, better modern graphics support on a Linux OS, essentially bringing it up to par with a lot of the Windows OSs. This is also a, our most recent production branch was released just last month at our GPU technology conference. Protocol support until March of 2027. We do have two different branches of support. We have what we refer to as a production branch. I like to think of it as like a new feature branch. So new things are added to that first. And then following that, we'll have a long-term support branch that'll come with this support and additional features later on. Yeah. We had a question come in the chat that either you or P4, I'm sure could answer, but can you explain specifically how MIG helps with like noisy neighbors to avoid one shared user from crashing into another? Yeah, sure. So with vGPU by itself, what that entails is you have a hard allocation of frame buffer or video memory that gets assigned to an individual VM. The GPU compute that's on that GPU is then time-sliced. So what that means is, let's say you have, you know, 12 users sharing one GPU. Each of them has a specific hardware allocated memory or frame buffer associated with it that cannot bleed into others. But they can potentially take over the entire GPU from a compute perspective, which is what's typically referred to as a noisy neighbor. Basically creating wait times for others that might need to use that. With MIG, by hardware allocating the resources, you're actually taking fractions of both GPU compute and frame buffer and passing that to the hypervisor as if it was a maximum resource that that partition can potentially leverage. So if you have 12 users that are sharing that one GPU, the max that they would only be able to take over from a GPU compute perspective would be a quarter of that GPU. So it would be less likely you would create a noisy neighbor type situation. Perfect. Cool. Thanks, Jimmy. Yep. And then, yeah, so from a Verge perspective, you know, the drivers that they deploy, vGPU is a dual driver system. So there is a driver that gets deployed on the host itself. That basically allows the hypervisor to see the GPU, allocate the resources and features that go along with it. And then you have a second driver that gets installed on the guest VM itself. And I think I think Paul will show you some of the details of that later on in the demo. So. All right. And so now, Jimmy, we've got another slide for you just to talk a little bit to us about this specific, the RTX. Yeah. Yeah. So basically, yeah, just a little bit about how they how they essentially work together and how we work together, you know, as partners. So I mentioned earlier the RTX Pro 6000 first GPU with MIG plus vGPU support. That GPU features up to four MIG slices on each GPU, so you can divide it into up to four pieces and then further subdivide that with vGPU on top of it. It is it is a larger GPU. It is a full size double with GPU. It is certified for the majority of your OEMs, regardless of of who you work with from a hardware perspective. It does have a high power consumption up to 600 watts. However, that power is actually variable, which is another unique feature of that GPU. So it can actually be dialed down all the way down to 450 watts. Potentially, this gives you a little bit more flexibility for specific OEM servers. Typically the OEM will be the one that will recommend what power settings to use and things like that, and they'll certify at a certain configuration. Also keep in mind that vGPU is a licensed product. It does require licenses to go along with it. There are basically three vGPU license types, depending on what you're looking to do with the GPU. The sort of base license is our virtual apps or virtual applications license. This is for like app streaming, things like that. Virtual PC, which is your sort of general knowledge worker type license type, gives you a maximum of three gigabytes of frame buffer. It's really suited to sort of your knowledge workers that just need pixel acceleration or multi-monitor support. And then there's our virtual workstation license, which is the superset license of all of them. It basically gives you full capabilities of the GPU. That includes full CUDA support up to maximum amounts of frame buffer, depending on whatever your GPU is. Gives you the ability to do high-end graphics workloads or AI workloads with that same license. One of the benefits of the RTX Pro 6000 with MIG and vGPU support is because you're presenting those four MIG slices to the hypervisor as if they were individual GPUs, each of those can actually host different profile types for different use cases. So you may have a situation where you have one MIG slice that's dedicated to an entire one user that's doing AI workloads. You then have maybe the other slices dedicated to high-end visualization type workloads or CAD workloads, something like that. And then another set of users that might be using it for just general knowledge worker type use cases. All of those can coexist on the same GPU, which makes this a very, very unique offering. Yeah. Go ahead. No. I was just agreeing with you. Sorry. Okay. No, it's okay. So, you know, and a little bit about what this partnership means. So I mentioned earlier, I was thoroughly impressed with what the Verge team had already integrated into their product from a GPU perspective. The official support between NVIDIA and Verge together just makes it even better. Like I said earlier, this gives us the ability for both of our engineering teams to work on any issues that might come up. One call can get both vendors working on an issue together if something were to arise. In addition to that, like I said earlier, this gives the Verge team the ability to integrate new features as they come about faster than they otherwise would have before because they'll be made aware of things earlier and integrated into that development process earlier on. Yeah, so, and this is ready today. I think that's a thing you can't stress enough. With vGPU 20's release March, you know, last month, that added support for all of this. It's already up and running today and you can use it today. Awesome. Thanks. Yeah. Yeah. And a question came in on the chat that I just want to kind of take, I'll just take verbally because, Jimmy, you've got to, you've got to be stored Switzerland a little bit on this one. So I appreciate that. So the question came in, is using us with your guys' products easier than VCF, VMware Cloud Foundation? I mean, we're going to say yes, I think, and I truly would believe that P4 could talk to it probably a little bit more, but you'll see it here in a second. And the best thing to do is let us get it set up for you, take you through the product. You know, that's, Sarah's put the meeting link in there a couple of times. Click on that. Let's get you set up and take you through the process. It's, you're the only one that can really tell the difference, but I can tell you right now that we have customers constantly tell us how much easier it is. So short answer is yes. So anyways, go ahead, Aaron. Yeah. Well, and to your point, George, and to the question, thank you for the questions. Keep them coming. You know, when we're talking about VergeOS and GPU architecture, right, we've already heard a little bit about three modes, but what we're getting ready to see is it's all one interface and zero specialists require the power behind what we are presenting here today. So with that in mind, maybe Jimmy, I think, let's see, would it be Jimmy or maybe would it be Paul that we would like to have us through the three modes? Yeah. Let's P4, please, if you will, take us through these three modes, the one interface and the no CLI. What does that all mean? Yeah. So in the Verge operating system, we try to make this very easy to consume so that you're not going to 20 different places, having to log into a CLI to pre-configure all these tweak kernel parameters. So we've kind of simplified, you know, how do you take a device, how do you break it up? And this doesn't just apply to GPUs, right? You've got GPU, GPU passthrough, vGPUs, MIG, that's one subset, but we've even taken it further as far as like network cards, you know, you can break those up into SROV devices, very similar concept and actually same technology that NVIDIA is leveraging for some of this. So we basically took that concept of, I have a physical piece of hardware, I want to give it to a virtual machine. We put together a common interface to basically make that happen. And, you know, traditionally, as Jimmy had mentioned, you know, when you want to pass a virtual GPU into a guest, you're pretty much just doing GPU passthrough. So you're very limited on how much you can do. And, you know, with the novel idea of virtual GPU that they released, it really opens up a lot of possibilities to be able to buy, you know, maybe it's a little bit more expensive of a piece of hardware, but now you can leverage that. Especially for like a managed service provider or, you know, development team, you can get more bang for your buck out of that product because you can issue so many different workloads to them. Once I first found out about MIG and vGPU, I mean, my eyes kind of opened, like, you know, one of the questions that came up is the noisy neighbor thing. It's like now as the administrator, you can make that decision. You can say, no, these people, I can, you know, they can go and do whatever they want. And it's not real important of workload. But if you've got something that latency is very sensitive, you know, MIG really allows you to put them in their own bucket and leverage that physical hardware separation so that they, you know, those workloads can get the IOPS and, you know, the CUDA cores that they need. Also, I'll piggyback on top of that from what I mentioned earlier. So while we do offer with just plain vGPU what we call heterogeneous profiles, that's the mixing of different profile types on one GPU. We do offer that, but I'll be honest, in a large enterprise organization, it can be quite difficult to manage to do it right. By leveraging MIG with vGPU, it makes that so much easier because each MIG slice by default can host a different vGPU profile on top of it. So for situations like I mentioned, where you might have different use cases in your organization, you might have, you know, AI use cases, you might have general knowledge worker use cases, you might have graphics work use cases. Those can all coexist on an individual GPU within a cluster and easily be mixed. Awesome. Thank you, Jimmy. And Paul, just to be clear, if not for the integration that we're talking about with one interface, how in the world does somebody set up, manage MIG in particular? What's different here? You know, you're basically dropping to a console. You are determining which compute instances, which GPU instances are available. You have to go and issue them, bring them up, tear them down, and then also create all these profiles. It's a lot of tooling that you have to do that we've kind of just taken and wrapped up into a single form that says, this is actually what I want. And we're doing that lifting behind the scenes to basically start breaking things up and ensuring that the right things get put in the right buckets. Awesome. Thank you. We've advanced the slide for our viewers, Paul, to talk about this one upload automatic deployment to every GPU-enabled VM. Maybe you can talk about this for us for a bit. Yeah. So maybe I'll take it from the concept of, you know, hey, I have VergeOS installed. I have a GPU plugged into my environment. What do I do next, right? So the first step is obtain the actual, the valid license, or to obtain the driver directly from NVIDIA's portal. Once you have that portal, you basically just need to upload that one file, right? That one file gets uploaded to VergeOS. When you upload it, you're basically going to say, now I have this graphics card. I want to turn this into a virtual GPU or MIGV GPU. During that step, you pick the driver, and we'll give you a very friendly list of it. Hey, we found version 20.0. Pick this one from the list. And then another cool feature that we actually do as well is not only are you saying, I want to leverage for vGPU, you also can have our operating system build a driver ISO for all of your guests that will automatically be attached. One thing that's not mentioned here as well, if you have your client configuration token, which is what NVIDIA uses to actually license the vGPU, you can upload that same file through the same interface you did with the drivers and actually bundle that in with the ISO so that when you go to boot your, you know, you plug this vGPU into your guest, everything you need to get it licensed, installed is all already there at your disposal. So there's no, okay, now I need to go log into the portal and put this here. And, you know, there's all these other steps that we're just allowing you to bypass and simplify into one process. And that attaching the ISO, it's an option, it should be on by default, but we'll handle all that for you so that when you install Windows and it's booted, you've got everything you need right there. And that pretty much covers one through five. I took three through five and kind of summarized them, right, but. Sure, sure. But point and click, right? Not CLI. No CLI needed. Very different. Jimmy, anything here that you wanted to add? If not, I don't think so. There is a question in the chat. It doesn't look like I was able to answer it when I tried to post, but somebody asked about support for the RTX Pro 4500 Blackwell Server Edition that we just announced at GTC last month. So that GPU is supported with vGPU 20. I don't believe it has been officially tested yet as those GPUs are not really out in the wild for the most part. But because it's part of the same driver set and package, I don't see any reason why it wouldn't be working here. Awesome. Thank you. And I'll second what Jimmy said there. All right. So let's see, Paul, just to maybe continue on a little bit here. This concept that GPUs are first class resources in VergeOS, what does that mean? How does that work in the real world? Yeah. So basically, when you buy one of these GPUs, you obviously have to have a use case for it, right? You're not just buying these just to have it. Hey, I need to build inference or I need AI running. But if you were to do something like this in Proxmox, you're going to be jumping to the CLI. Yeah, they do have some niceties in there, but getting everything working, they didn't do any of the heavy lifting for you. There's not a lot of automation that we're going to provide that we tried to simplify for you so that it's more click button as opposed to let's go read a manual. Let's go read NVIDIA's documentation, Proxmox's documentation. How do I add this? It's really just a simple single form. So we've kind of brought that in as like, I guess you would call it like that first class resource, right? You have the use case for it. We're going to treat it like our baby and we're going to make it as easy to use as possible. Now, moving on to snapshots and replication, since the way that we handle these, we call them pass through devices, right? Because even though it's a VGPU, we're still passing, even if it's software, not emulated, but software generated, it still is a device that we're passing through. All of that, that configuration, everything gets encapsulated with our snapshot engine. So if you're doing things like replication and DR planning using VergeOS to VergeOS, that entire config is going to go along for the ride. So assuming if you have a DR event, you've got those same GPUs, you're going to be able to move VMA to site B and power it on there as well. So GPU utilization, we're going to monitor what's in use. There's things that we move very fast in development over here at VergeOS and more and more is being added, but we're going to track the inventory, the lifecycle, where you have resources available. It's all kind of pooled together in one logical group so that if you need to do, hey, I need to live migrate this virtual machine from one node to the other, we're going to figure all that out for you. You don't have to go and figure it out manually. You're just basically going to say, hey, I need to reboot this physical host. We're going to monitor all that, figure out where the resources are available and let you know if they are and give you options, right? And yeah, I mean, you know, single management plane, just like VergeOS, there's only one endpoint that you connect up to. It's API first, so you don't have to go to multiple places. We're not stitching different things together. There's one single API, so one user interface will accomplish it. Or if you're using something like Terraform, et cetera, for automation, that all goes to the same endpoint and it's the same integration. A question came in that I can't even make up an answer for, so I'm going to give it to you. Our use case is to have multiple teams do prototype work, then freeze the VMs. Can another team member fire up his VMs and use the GPU and do the same thing while those VMs are frozen? When you say frozen, do you mean hibernated or powered off? What? I guess I would need a little clarification on that. Pick one, because I don't have the detail. Yeah, I mean, if once a resource is, yeah, if it's powered off or paused. Yeah, if it's powered off, then, you know, that's the resource has been given up. You know, we can only just like think of it like you have a single GPU and GPU passthrough. A vGPU is a resource. When a virtual machine is using it, it's locked. Nobody else can touch it until you're down. So if we were to, you know, live migrate that somewhere else, all the VRAM is going to go along for the ride. It's going to move to another node, but you still can't consume that resource. But yeah, they can be taken on demand. You know, if a developer powers up his workload, runs his things, he powers it off, that resource is gone. We take our resources and put them into a pool. It's just a pool, and someone takes one out. When they're done, they put it back in the pool. We're not thick allocating things. It's all just you basically come up with a logical group, and then there's so many things in that group, and someone just takes out what they need, and then they put it back when they're done. Hopefully that answers the question. Yeah, and then there's another question came in about us talking a lot about not needing a CLI, but what about using CLI? You were almost going to go there anyway, so talking about automation through Terraform or something like that, still automatable? Yeah, I mean, when you talk about CLI, I'm talking about the Verge OS. You know, you want to go to the physical host and log into the terminal. You're not having to do any of that in Verge OS. Now, if we're talking CLI running something like Terraform, you're generally doing that on your desktop, and there's going to be API keys that are connecting up to our API. All that automation can still happen. So, you know, there is CLI access and things that can be done with the CLI for automating, but from our user interface perspective, there's no need to jump in and manually do anything other than upload some files. That makes sense. Hey, Aaron, we've talked a lot about the deployment, I think, throughout. Let's kind of skip that. I think people want to see this thing happen. So let's jump right to it. I think you're right. But there is something I want to share as we get ready to jump right in, and that is a really keen observation from Jason Yeager, one of our team members, our SVP of engineering. You guys can probably all see this. I'm just going to read it quickly. This is the piece that surprises IT directors every time they expect a command line, right? We were just talking about that. When they see MIG configuration done entirely through a point and click interface, the conversation shifts from can we afford to do this to when do we start? These are not just musings of an SVP of engineering. These are things that we're seeing and hearing from you and your peers. OK, we do have some resources that should be in the resource section at the bottom of the webinar screen. And you can go ahead and grab any of these that you like. You know, take a look there. But what we're all waiting for now, a live demo. So why don't I stop sharing my screen before you can fire up what you've got going on and really let everybody see what they're actually here to see. All right, so let me. I've stopped sharing. OK, so I'm just going to touch this real quick. And to be honest, there really isn't a lot to show here because we do a lot of the automation. So this should be relatively quick. But I'm going to start by, you know, in those steps one through five. The first one is you need to obtain the drivers, right? So in this lab environment that I have here, I've already uploaded version 20.0. So I download this from NVIDIA directly. I uploaded it right from here. You know, you can basically upload it from your PC or you can actually just point the URL and we'll actually behind the scenes go and download it directly if you happen to have that URL. Once this is in a system, now what I need to do is I'm going to go and look at my nodes. Node 3 is the server that has the RTX 6000 Pro. From here, I'm going to go and look at my PCI devices. And let's go ahead and filter on the display controller. Now, since this card is so new, we don't have a nice friendly description for it yet. But I do know that this is the actual RTX Pro. So all I do is I click this. I say make resource. Now what it's going to do is I'm going to be creating what we call a resource group. The resource group is basically a pool of pass-through devices. They can be things like, you know, a host GPU for AI and inference. A VGPU, a direct pass-through, SRIOV. We also do support USB as well. But obviously in this, we're going to go and look at VGPU. Now, this is basically it, right? This is the form that we're interested in. And right down here, I get to pick the driver that I want. So as you can see, there's a lot of versions on here. But version 20 is the one that we're interested in here because that's what actually has official support. Um, so now I've already applied this. So I kind of skipped a little bit here. But generally, this list, if this is a vanilla system, we're not going to know what VGPU profiles this card supports yet. I've already run this test before. So these are already generated. If this list is empty, it's as simple as go through the steps I'll show you in a minute to apply the drivers, come back to this page. We're going to basically query the card and NVIDIA drivers ourselves, figure out what's available and present it in this list to you. So in this list, I mean, there's a lot of options here, right? So the main one that we're kind of showing off here is MIG. So looking at this, I know this card is 96 gigs of VRAM. MIG is going to let me break this thing up any way I see fit. So we've got 24 gig or there's we have 48 or you can actually just make a single MIG compute instance that's 96 gig. For this test, what I'm going to do is I'm going to pick the 12Q, 24 gig. Now what this will auto fill in here, based on the profile that I picked, I can create up to this number of MIG instances. So I can create up to four. What this enables us to do is kind of that heterogeneous topic that Jimmy was talking about. I can say, you know what, I want to take this MIG card. I want to break it up into a single 24 gig. But then also I want to fill it up with, I only want to create one of them, but I also want to fill it up with this particular VGPU instance. If I leave this at zero, it's going to create as many as it possibly can. Since I picked a 12Q, that means 12 gig of VRAM. This is a 24 gig MIG instance, so I can create two of them. I could limit this to only create one, but I'm just going to leave it at the default. So basically I'm going to create one MIG instance. It's going to create two VGPU profiles. The next fun part is, okay, I'm going to create this resource group right here. Let's just give this a name. 6000 Pro SE 12Q. So now what I want to do is I want to click make guest drivers. What this is going to do, it's going to take this NVIDIA driver bundle. It's going to turn it into an ISO so that we can automatically plug it into the guest upon creation or plugging in a device. A newer thing, did someone say something? Well, I said, wow, sorry. Another cool new feature is if you upload your client config tokens, these can be grabbed from the NVIDIA portal. You know, you'll have a license server. You generate your client config token. If you upload it to the files, we'll show it in this list. And I can basically just hit this button. I know this is the most recent one that I generated today. I'll hit submit. So now it's going to generate that ISO. If I go back and look at node three, it's going to say that I need to reload the drivers. So I'm just going to follow the steps here, put this node in maintenance mode. It's going to evacuate all workloads. That's what maintenance mode does. It gets everything off of the node. Once it's done, it's in maintenance mode. I click reload drivers. And then if I go down here, I'll actually just go ahead and pull this out of maintenance mode while that's installing. And if we go down here, it's waiting for the drivers to finish reloading before processing maintenance mode because we had to be in maintenance mode. And as soon as this is done, drivers will be installed. Reload's complete. That's one of my favorite features that you guys offer is that single, that one window that you just showed that did basically three different steps that would normally be done elsewhere. You configured vGPU. You configured MIG. And you basically bundled the drivers and the token files for the license configuration file all in that one window. That's like three or four different steps you would normally have to do. Plus, you guys, I believe, are the only ones that actually offer that MIG configuration without dropping to the command line and doing that elsewhere. So it's a very, very cool feature. It was a lot of fun to work. Thank you for what we've got to break there. A question came in. Does the VMs that are going to use the GPU have to be running on the same host as the GPU? Yep. Okay. And then you can't over-provision VRAM, is that correct? Yeah, it's a hard allocation. So you can think of your VRAM as your finite resource. So when you're sizing for a deployment, depending on whatever your workloads are, that's going to be the determining factor of what GPU you need, what size profile you need, or what size MIG slices you would need to run that workload, because that is your hard allocation of resources. Gotcha. The GPU compute itself, though, that is and can be time-sliced. So that can be not over-provisioned per se, but can be utilized by multiple people that are sharing that resource. Gotcha. So after I reloaded the drivers, you'll notice that I have two node resources here. They're both at 12Q profile. So this is us basically taking that resource group, which is also kind of like a pool, we add the resources to them. So if you picked a different profile, like a 2Q, we're going to create more of them. You will see them in this list here, and none of them are currently in use. Now what I'm going to do is I'm going to go over to a virtual machine right here. And basically, I'm going to add a new device. So when I say add device, it automatically selected vGPU, because that's the only one that I have available. It picked the 12Q resource group. Optionally, I could give it a name if I want, or we'll just auto-generate one. And then the real friendly feature here is attach the guest drivers. So this is going to take the ISO that we generated and automatically plug it into this guest. So in fact, I'll show you right here what it's doing. So you'll notice I have three drives here, an EFI, CD-ROM, and just a plain disk. If I go and add the device, I click attach drivers, I hit submit. We'll automatically insert the guest drivers right here, so that when I power this thing on, what we're going to see is this ISO will be available. It'll have the token in there. It'll have the drivers to go ahead and actually do the install. So let's console in here. All right. So if I go down here, there's an NVIDIA drivers CD-ROM that gets created here. Here's that client config token. So when I'm ready to register this, you basically just copy it into a directory. And if I click into here, I can go into guest drivers. I've got the Windows ones right here. And all I need to do at this point is install this, which this will take a little bit. We don't really need to show this off, but basically install this. It's going to get the drivers already. Then we need to copy that config token in, and then we're done. We're off and running. Cool. And that is about it. Most of you guys can think of something else that maybe you want to see from the panel here. I think that's part of the power of what we're talking about here is that that is it. Yeah, right. Yeah, yeah. I suspect this would have taken someone else an awful lot longer period of time than it just took you. Awesome. Jimmy, a fair question came in that I certainly can't answer that maybe you can, because I'm sure you probably get this kind of question all the time from customers. The NVIDIA DGX Spark GB10 supports VGB natively via NVIDIA AI Enterprise. What are the benefits of VGB from Verge versus this one? And then you don't need to pick a side, but just when would you advise a customer one way or the other? Well, I mean, so Spark is a very different use case than what we're talking about here. I also, I don't believe that that does support VGPU. So I don't know that that's a correct statement. Okay, fair enough. And yeah, I think if there are any final words, Jimmy P4, before we wrap up. No, thanks for all the work you guys put in supporting this. And it's a great solution and a great partnership. So thank you. Yeah, we really appreciate it as well. Thank you so much for spending time with us today. And we hope to see you again soon. We do a webinar every week or two. So keep your eyes peeled and we'll see you again soon. Take care. Transcribed by https://otter.ai

TL;DR

  • VergeOS eliminates the GPU expertise barrier by providing point-and-click management for NVIDIA vGPU 20, including MIG configuration, driver deployment, and licensing — all without CLI access or specialized skills.
  • The official NVIDIA partnership enables joint engineering support, faster feature integration, and support for the latest RTX PRO 6000 Blackwell Server Edition hardware with 96GB VRAM.
  • A single driver upload automatically generates guest ISOs with bundled licensing tokens and deploys to all GPU-enabled VMs, reducing deployment time from days to minutes.
  • MIG slicing allows heterogeneous workloads (VDI, AI, data science, general computing) to coexist on the same GPU with hard VRAM allocation and time-sliced compute resources.
  • The unified interface treats GPUs as first-class infrastructure, managing passthrough, vGPU, and MIG modes through the same workflow used for network cards and other resources.
  • Organizations can fully utilize expensive GPU hardware across multiple use cases without hiring specialists or maintaining fragmented tooling.

The GPU Management Expertise Gap

The webinar opens by addressing a critical challenge facing IT organizations: 73% of teams lack in-house GPU expertise, yet are increasingly being asked to support GPU workloads for VDI, AI inference, data science, and visualization. Traditional GPU deployment requires specialized knowledge of driver management, MIG (Multi-Instance GPU) configuration, and hypervisor integration — skills that come at a premium. Organizations face a difficult choice: hire expensive specialists, limit GPU deployment scope, or struggle with operational complexity. The session frames this as not just a technical problem but a business constraint that prevents organizations from fully leveraging their GPU hardware investments. VergeIO and NVIDIA position their partnership as a solution that eliminates the expertise barrier through unified infrastructure management.

NVIDIA vGPU 20 and the VergeOS Partnership

Jimmy Rotella from NVIDIA introduces vGPU 20, the latest release that enables official support for VergeOS integration. The partnership provides joint engineering support, faster feature integration, and access to new capabilities as they're developed. A key hardware announcement is support for the RTX PRO 6000 Blackwell Server Edition, which combines with vGPU 20 to enable flexible GPU resource allocation. The technology allows multiple use cases — VDI, AI workloads, data science, and general knowledge workers — to coexist on the same GPU hardware. This heterogeneous profile capability, especially when combined with MIG slicing, addresses the historical challenge of GPU resource management where organizations either over-provisioned dedicated GPUs or struggled with complex time-slicing configurations. The partnership makes these advanced features accessible without requiring command-line expertise.

VergeOS Architecture: Three Modes, One Interface

Paul Hodges, VergeOS Field CTO, demonstrates the platform's unified approach to GPU management across three modes: GPU passthrough, vGPU, and MIG-enabled vGPU. The architecture treats GPUs as first-class infrastructure resources, similar to how the platform handles network cards with SR-IOV. The key differentiator is a single point-and-click interface that eliminates the need for CLI commands, kernel parameter tweaking, or multi-tool workflows. Driver management is simplified to a single upload that automatically generates guest driver ISOs, bundles NVIDIA client configuration tokens for licensing, and deploys to all GPU-enabled VMs. MIG configuration — traditionally requiring console access and manual instance creation — is handled through dropdown menus that automatically calculate available slices based on VRAM allocation. The system queries GPU capabilities, presents available profiles, and manages resource pools without requiring administrators to understand the underlying complexity.

Live Demonstration: End-to-End Workflow

The demonstration walks through the complete deployment process on an RTX PRO 6000 SE with 96GB VRAM. Hodges shows how uploading a single NVIDIA driver file triggers automatic profile discovery, presents MIG slicing options (24GB, 48GB, or 96GB instances), and creates resource groups. The interface allows administrators to select a 12Q profile with 24GB VRAM, which automatically calculates that four MIG instances are possible. The system generates guest driver ISOs that include both NVIDIA drivers and licensing tokens, then automatically attaches them to VMs when GPU devices are added. The entire process — from driver upload to VM with functioning vGPU — takes minutes and requires no command-line interaction. Maintenance mode automation handles driver reloads, workload evacuation, and node management. The demonstration emphasizes that VRAM is a hard allocation (no over-provisioning), but GPU compute can be time-sliced across multiple users sharing the resource.

Chapters

0:00 - Introduction & Panelists
2:00 - The GPU Management Problem
6:00 - Poll Results & Expertise Gap
9:00 - NVIDIA vGPU 20 Overview
18:00 - RTX PRO 6000 Blackwell & Partnership
21:00 - VergeOS GPU Architecture
28:00 - Driver Management Workflow
30:00 - GPU as First-Class Infrastructure
33:00 - Deployment Scenarios
40:00 - Live Demo: Complete Workflow
53:00 - Q&A Session

Key Quotes

6:55 "... 73% lack of in-house GPU experience. Time to production is sizable and the cost, right, of a specialist, you're a specialist for a reason and it comes at a premium."
9:53 "I think the biggest thing is sometimes people overlook or are not aware of the vGPU functionality and they look at it as traditionally pass-through was the only option and dedicated resources to a high-end data center GPU is often overkill for a lot of organizations and what they need to do."
22:07 "All of those can coexist on the same GPU, which makes this a very, very unique offering."
23:32 "This is ready today. I think that's a thing you can't stress enough. With vGPU 20's release March, you know, last month, that added support for all of this. It's already up and running today and you can use it today."
28:25 "By leveraging MIG with vGPU, it makes that so much easier because each MIG slice by default can host a different vGPU profile on top of it."
46:39 "That's one of my favorite features that you guys offer is that single, that one window that you just showed that did basically three different steps that would normally be done elsewhere. You configured vGPU. You configured MIG. And you basically bundled the drivers and the token files for the license configuration file all in that one window."

FAQ

Do VMs using GPUs need to run on the same physical host as the GPU hardware?

Yes, VMs must run on the same physical host where the GPU is installed. GPU resources cannot be shared across hosts in the cluster, though VergeOS manages resource pools and allocation across nodes through its unified interface.

Can you over-provision GPU VRAM like you can with CPU or memory?

No, VRAM is a hard allocation and cannot be over-provisioned. When sizing GPU deployments, VRAM is the finite resource that determines what GPU hardware, profile size, or MIG slices you need. However, GPU compute itself can be time-sliced and shared across multiple users.

How does VergeOS compare to VMware Cloud Foundation for GPU management?

VergeOS provides a significantly simplified workflow with point-and-click MIG configuration, automatic driver ISO generation with bundled licensing tokens, and unified GPU resource management — all without CLI requirements. The platform is designed specifically to eliminate the complexity and specialist requirements of traditional hypervisor GPU implementations.


Categories:
  • » Data Management » Virtualization
  • » Webinar Library » Verge.io
  • » Data Protection » Backup & Recovery
  • » Cybersecurity » Cloud Security
  • » Data Protection
Channels:
News:
Events:
Tags:
  • Cloud Security
  • Data Protection
  • AI & Machine Learning
  • Technical Deep Dive
  • Demo
  • Webinar
  • GPU Virtualization
  • NVIDIA vGPU 20
  • MIG
  • Multi-Instance GPU
  • Infrastructure Management
  • VDI
  • Virtual Desktop Infrastructure
  • AI Inference
  • Data Science Workloads
Show more Show less

Browse videos

  • Related
  • Featured
  • By date
  • Most viewed
  • Top rated
  •  

              Video's comments: Verge.io: GPU Virtualization Without Specialists: VergeOS + NVIDIA

              Industry Events (Sponsor Hosted)

              • Aug
                03

                Discover DLP Memories: The ever-evolving triage agent enhancing efficiency each shift.

                08/03/202611:00 AM ET
                • Aug
                  06

                  Safeguarding Sensitive Data in the Era of Public AI Platforms

                  08/06/202604:00 AM ET
                  • Aug
                    06

                    Same Tactics, Enhanced Speed: AI Agents’ Impact on Identity Attacks

                    08/06/202602:00 PM ET
                    More events

                    Upcoming Webinar Calendar

                    • 08/03/2026
                      11:00 AM
                      08/03/2026
                      Discover DLP Memories: The ever-evolving triage agent enhancing efficiency each shift.
                      https://www.truthinit.com/index.php/channel/2062/discover-dlp-memories-the-ever-evolving-triage-agent-enhancing-efficiency-each-shift/
                    • 08/06/2026
                      04:00 AM
                      08/06/2026
                      Safeguarding Sensitive Data in the Era of Public AI Platforms
                      https://www.truthinit.com/index.php/channel/2058/safeguarding-sensitive-data-in-the-era-of-public-ai-platforms/
                    • 08/06/2026
                      02:00 PM
                      08/06/2026
                      Same Tactics, Enhanced Speed: AI Agents’ Impact on Identity Attacks
                      https://www.truthinit.com/index.php/channel/2064/same-tactics-enhanced-speed-ai-agents-impact-on-identity-attacks/
                    • 08/07/2026
                      11:30 AM
                      08/07/2026
                      Refreshing Beverage Ideas Paired with Essential Cybersecurity Insights
                      https://www.truthinit.com/index.php/channel/2063/refreshing-beverage-ideas-paired-with-essential-cybersecurity-insights/
                    • 08/13/2026
                      12:00 PM
                      08/13/2026
                      Harnessing AI for Secure Innovation in the Enterprise with Netskope & Omada
                      https://www.truthinit.com/index.php/channel/2065/harnessing-ai-for-secure-innovation-in-the-enterprise-with-netskope-omada/
                    • 08/19/2026
                      12:00 PM
                      08/19/2026
                      Becoming Agent Ready: Insights and Strategies with Cyera
                      https://www.truthinit.com/index.php/channel/2036/becoming-agent-ready-insights-and-strategies-with-cyera/
                    • 09/02/2026
                      12:00 PM
                      09/02/2026
                      Unified Data Security in Action: Uncover, Analyze, and Resolve Threats
                      https://www.truthinit.com/index.php/channel/2045/unified-data-security-in-action-uncover-analyze-and-resolve-threats/
                    • 09/30/2026
                      04:00 AM
                      09/30/2026
                      AI Command Center: Optimizing Visibility and Control in Your Operations
                      https://www.truthinit.com/index.php/channel/2024/ai-command-center-optimizing-visibility-and-control-in-your-operations/
                    Truth in IT
                    • Sponsor
                    • About Us
                    • Terms of Service
                    • Privacy Policy
                    • Contact Us
                    • Preference Management
                    Desktop version
                    Standard version