Truth in IT
    • Sign In
    • Register
        • Videos
        • Channels
        • Pages
        • Galleries
        • News
        • Events
        • All
Truth in IT Truth in IT
  • Data Management ▼
    • Converged Infrastructure
    • DevOps
    • Networking
    • Storage
    • Virtualization
  • Cybersecurity ▼
    • Application Security
    • Backup & Recovery
    • Data Security
    • Identity & Access Management (IAM)
    • Zero Trust
    • Compliance & GRC
    • Endpoint Security
  • Cloud ▼
    • Hybrid Cloud
    • Private Cloud
    • Public Cloud
  • Webinar Library
  • TiPs
  • DRAW

Scale Computing: Edge AI Deployment with GPU Offload on SC//HyperCore

Scale Computing
06/21/2026
0 (0%)
Share
  • Comments
  • Download
  • Transcript
Report Like Favorite
  • Share/Embed
  • Email
Link
Embed

Transcript


Hypercore nodes using GPU offload. So this particular cluster is equipped with NVIDIA GPUs. We're using the virtual GPU provider to, if we wanted to be able to partition and carve up the GPU into multiple profiles, you could even have multiple workloads on each node accessing these. This tag is all I need to do to specify the GPU resources that I want to be assigned to this particular virtual machine. In this case, we're all up and running and provisioned. And on the console, I just have a tool called nvtop set up and running so that we can see the GPU utilization. This is a SSH connection into the terminal, and what I want to show here is that everything is running here in containers. There's different ways to deploy it, but the key thing here that actually has everything, the web user interface, the model, is deployed in this open web UI container. It's kind of all in one, but I have some other containers that are running here. And then let's go ahead and we'll just do a demo of some of the capabilities. So let's paste in a query here. Actually, not the one. Let's write a white paper about deploying AI inferencing at the edge using scale computing hypercore. And we can see that even as I start to type and then I submit the chat, the GPU utilization is taking advantage of the NVIDIA L4 card and the profile that's available to it. Really little CPU utilization across here. It took 12 seconds to kind of think through here, and it's satisfying my request. All types of different stacks could be deployed in virtual machines. They could be Windows, Linux, containerized. In this case, we're using Docker. Other container runtimes like Podman, Kubernetes, all could be utilized to run various software stacks that provide large language models, computer vision, and so forth. We can help provide automation to be able to deploy the applications initially, deploy updates to the models across an entire edge fleet, and keep your applications fresh and new and also give you the ability to rapidly deploy new applications that you might need on that same consolidated and resilient infrastructure. Thank you.

TL;DR

  • SC//HyperCore enables edge AI inferencing using NVIDIA GPU offload with virtual GPU profiles that allow multiple workloads to share GPU resources efficiently across nodes.
  • The platform supports containerized LLM deployment using Docker, Podman, or Kubernetes, with real-time GPU monitoring showing minimal CPU utilization during AI workload processing.
  • Scale Computing provides automated fleet management for deploying and updating AI models across distributed edge locations, enabling rapid deployment of new applications on consolidated infrastructure.

Summary

This technical demonstration showcases how Scale Computing's SC//HyperCore platform enables edge AI inferencing workloads through GPU offload capabilities. The walkthrough illustrates the deployment of containerized large language models using NVIDIA GPUs with virtual GPU profiles, allowing multiple workloads to share GPU resources across nodes. The demonstration features a live example of an LLM generating a white paper about edge AI deployment, with real-time GPU utilization monitoring via NVTop showing efficient resource allocation with minimal CPU overhead. The platform supports various deployment methods including Docker containers, Podman, and Kubernetes, providing flexibility for different AI applications such as computer vision and natural language processing. Scale Computing emphasizes the ability to deploy and update AI models across distributed edge locations through automated fleet management, enabling organizations to maintain current applications and rapidly deploy new AI workloads on consolidated, resilient infrastructure designed specifically for edge computing environments.

Chapters

0:00 - Introduction to Edge AI on HyperCore
0:13 - GPU Configuration and Virtual Profiles
0:46 - Container Deployment Architecture
1:11 - Live LLM Demonstration

Key Quotes

0:17 "We're using the virtual GPU provider to, if we wanted to be able to partition and carve up the GPU into multiple profiles, you could even have multiple workloads on each node accessing these."
1:41 "Really little CPU utilization across here. It took 12 seconds to kind of think through here, and it's satisfying my request."
2:10 "We can help provide automation to be able to deploy the applications initially, deploy updates to the models across an entire edge fleet, and keep your applications fresh and new."

FAQ

What GPU hardware does SC//HyperCore support for edge AI workloads?

SC//HyperCore supports NVIDIA GPUs including the L4 card demonstrated in the video. The platform uses virtual GPU providers to partition GPU resources into multiple profiles, allowing multiple workloads to access GPU resources on each node.

What container runtimes are supported for deploying AI models on SC//HyperCore?

SC//HyperCore supports multiple container runtimes including Docker, Podman, and Kubernetes. The demonstration uses Docker with the Open Web UI container, which includes the web interface and model in an all-in-one deployment, but organizations can choose the runtime that best fits their requirements.


Categories:
  • » Data Protection
Channels:
News:
Events:
Tags:
  • AI & Machine Learning
  • Edge Computing
  • Demo
  • Technical Deep Dive
  • Infrastructure Management
  • Edge AI deployment
  • GPU offload
  • Virtual GPU profiles
  • Containerized LLMs
  • NVIDIA GPU integration
  • Edge computing infrastructure
  • AI inferencing
  • Fleet management automation
Show more Show less

Browse videos

  • Related
  • Featured
  • By date
  • Most viewed
  • Top rated
  •  

              Video's comments: Scale Computing: Edge AI Deployment with GPU Offload on SC//HyperCore

              Industry Events (Sponsor Hosted)

              • Aug
                06

                Safeguarding Sensitive Data in the Age of AI Platforms

                08/06/202604:00 AM ET
                • Aug
                  06

                  AI Agents Transforming Identity Attack Tactics and Speed

                  08/06/202602:00 PM ET
                  • Aug
                    13

                    Harnessing AI for Secure Innovation in the Enterprise with Netskope & Omada

                    08/13/202612:00 PM ET
                    More events

                    Upcoming Webinar Calendar

                    • 08/06/2026
                      04:00 AM
                      08/06/2026
                      Safeguarding Sensitive Data in the Age of AI Platforms
                      https://www.truthinit.com/index.php/channel/2058/safeguarding-sensitive-data-in-the-age-of-ai-platforms/
                    • 08/06/2026
                      02:00 PM
                      08/06/2026
                      AI Agents Transforming Identity Attack Tactics and Speed
                      https://www.truthinit.com/index.php/channel/2064/ai-agents-transforming-identity-attack-tactics-and-speed/
                    • 08/13/2026
                      12:00 PM
                      08/13/2026
                      Harnessing AI for Secure Innovation in the Enterprise with Netskope & Omada
                      https://www.truthinit.com/index.php/channel/2065/harnessing-ai-for-secure-innovation-in-the-enterprise-with-netskope-omada/
                    • 08/19/2026
                      12:00 PM
                      08/19/2026
                      Becoming Agent Ready with Cyera: Essential Strategies and Insights
                      https://www.truthinit.com/index.php/channel/2036/becoming-agent-ready-with-cyera-essential-strategies-and-insights/
                    • 09/02/2026
                      12:00 PM
                      09/02/2026
                      Unified Data Security in Action: Uncover, Analyze, and Resolve Threats
                      https://www.truthinit.com/index.php/channel/2045/unified-data-security-in-action-uncover-analyze-and-resolve-threats/
                    • 09/30/2026
                      04:00 AM
                      09/30/2026
                      AI Command Center: Optimizing Visibility and Control in Your Operations
                      https://www.truthinit.com/index.php/channel/2024/ai-command-center-optimizing-visibility-and-control-in-your-operations/
                    • 11/19/2026
                      01:00 PM
                      11/19/2026
                      360View: Govern, Secure & Recover Your Microsoft 365 Environment
                      https://www.truthinit.com/index.php/channel/2076/360view-govern-secure-recover-your-microsoft-365-environment/
                    Truth in IT
                    • Sponsor
                    • About Us
                    • Terms of Service
                    • Privacy Policy
                    • Contact Us
                    • Preference Management
                    Desktop version
                    Standard version