Truth in IT
    • Sign In
    • Register
        • Videos
        • Channels
        • Pages
        • Galleries
        • News
        • Events
        • All
Truth in IT Truth in IT
  • Data Management ▼
    • Converged Infrastructure
    • DevOps
    • Networking
    • Storage
    • Virtualization
  • Cybersecurity ▼
    • Application Security
    • Backup & Recovery
    • Data Security
    • Identity & Access Management (IAM)
    • Zero Trust
    • Compliance & GRC
    • Endpoint Security
  • Cloud ▼
    • Hybrid Cloud
    • Private Cloud
    • Public Cloud
  • Webinar Library
  • TiPs
  • DRAW

Scale Computing: Edge AI Deployment with GPU Offload on SC//HyperCore

Scale Computing
06/21/2026
0 (0%)
Share
  • Comments
  • Download
  • Transcript
Report Like Favorite
  • Share/Embed
  • Email
Link
Embed

Transcript


Hypercore nodes using GPU offload. So this particular cluster is equipped with NVIDIA GPUs. We're using the virtual GPU provider to, if we wanted to be able to partition and carve up the GPU into multiple profiles, you could even have multiple workloads on each node accessing these. This tag is all I need to do to specify the GPU resources that I want to be assigned to this particular virtual machine. In this case, we're all up and running and provisioned. And on the console, I just have a tool called nvtop set up and running so that we can see the GPU utilization. This is a SSH connection into the terminal, and what I want to show here is that everything is running here in containers. There's different ways to deploy it, but the key thing here that actually has everything, the web user interface, the model, is deployed in this open web UI container. It's kind of all in one, but I have some other containers that are running here. And then let's go ahead and we'll just do a demo of some of the capabilities. So let's paste in a query here. Actually, not the one. Let's write a white paper about deploying AI inferencing at the edge using scale computing hypercore. And we can see that even as I start to type and then I submit the chat, the GPU utilization is taking advantage of the NVIDIA L4 card and the profile that's available to it. Really little CPU utilization across here. It took 12 seconds to kind of think through here, and it's satisfying my request. All types of different stacks could be deployed in virtual machines. They could be Windows, Linux, containerized. In this case, we're using Docker. Other container runtimes like Podman, Kubernetes, all could be utilized to run various software stacks that provide large language models, computer vision, and so forth. We can help provide automation to be able to deploy the applications initially, deploy updates to the models across an entire edge fleet, and keep your applications fresh and new and also give you the ability to rapidly deploy new applications that you might need on that same consolidated and resilient infrastructure. Thank you.

TL;DR

  • SC//HyperCore enables edge AI inferencing using NVIDIA GPU offload with virtual GPU profiles that allow multiple workloads to share GPU resources efficiently across nodes.
  • The platform supports containerized LLM deployment using Docker, Podman, or Kubernetes, with real-time GPU monitoring showing minimal CPU utilization during AI workload processing.
  • Scale Computing provides automated fleet management for deploying and updating AI models across distributed edge locations, enabling rapid deployment of new applications on consolidated infrastructure.

Summary

This technical demonstration showcases how Scale Computing's SC//HyperCore platform enables edge AI inferencing workloads through GPU offload capabilities. The walkthrough illustrates the deployment of containerized large language models using NVIDIA GPUs with virtual GPU profiles, allowing multiple workloads to share GPU resources across nodes. The demonstration features a live example of an LLM generating a white paper about edge AI deployment, with real-time GPU utilization monitoring via NVTop showing efficient resource allocation with minimal CPU overhead. The platform supports various deployment methods including Docker containers, Podman, and Kubernetes, providing flexibility for different AI applications such as computer vision and natural language processing. Scale Computing emphasizes the ability to deploy and update AI models across distributed edge locations through automated fleet management, enabling organizations to maintain current applications and rapidly deploy new AI workloads on consolidated, resilient infrastructure designed specifically for edge computing environments.

Chapters

0:00 - Introduction to Edge AI on HyperCore
0:13 - GPU Configuration and Virtual Profiles
0:46 - Container Deployment Architecture
1:11 - Live LLM Demonstration

Key Quotes

0:17 "We're using the virtual GPU provider to, if we wanted to be able to partition and carve up the GPU into multiple profiles, you could even have multiple workloads on each node accessing these."
1:41 "Really little CPU utilization across here. It took 12 seconds to kind of think through here, and it's satisfying my request."
2:10 "We can help provide automation to be able to deploy the applications initially, deploy updates to the models across an entire edge fleet, and keep your applications fresh and new."

FAQ

What GPU hardware does SC//HyperCore support for edge AI workloads?

SC//HyperCore supports NVIDIA GPUs including the L4 card demonstrated in the video. The platform uses virtual GPU providers to partition GPU resources into multiple profiles, allowing multiple workloads to access GPU resources on each node.

What container runtimes are supported for deploying AI models on SC//HyperCore?

SC//HyperCore supports multiple container runtimes including Docker, Podman, and Kubernetes. The demonstration uses Docker with the Open Web UI container, which includes the web interface and model in an all-in-one deployment, but organizations can choose the runtime that best fits their requirements.


Categories:
  • » Data Protection
Channels:
News:
Events:
Tags:
  • AI & Machine Learning
  • Edge Computing
  • Demo
  • Technical Deep Dive
  • Infrastructure Management
  • Edge AI deployment
  • GPU offload
  • Virtual GPU profiles
  • Containerized LLMs
  • NVIDIA GPU integration
  • Edge computing infrastructure
  • AI inferencing
  • Fleet management automation
Show more Show less

Browse videos

  • Related
  • Featured
  • By date
  • Most viewed
  • Top rated
  •  

              Video's comments: Scale Computing: Edge AI Deployment with GPU Offload on SC//HyperCore

              XStreaminars (watch here)

              • Jul
                28

                Illumio + Netskope: Zero Trust in the Age of AI Autonomy

                07/28/202601:00 PM ET
                • Jul
                  29

                  Ask Your Cloud Anything: Unlocking Governance Silos in your Environments

                  07/29/202601:00 PM ET
                  More events

                  Industry Events (watch there)

                  • Jul
                    14

                    Crafting a Championship-Caliber Security Team for Lasting Defense

                    07/14/202601:00 PM ET
                    • Jul
                      14

                      Understanding the Crucial Role of Context in Safeguarding AI-Accessible Data

                      07/14/202602:00 PM ET
                      • Jul
                        22

                        Insights from Attackers During the FIFA World Cup: A HUMAN Dialogue

                        07/22/202601:00 PM ET
                        More events

                        Upcoming Webinar Calendar

                        • 07/14/2026
                          01:00 PM
                          07/14/2026
                          Crafting a Championship-Caliber Security Team for Lasting Defense
                          https://www.truthinit.com/index.php/channel/2025/crafting-a-championship-caliber-security-team-for-lasting-defense/
                        • 07/14/2026
                          02:00 PM
                          07/14/2026
                          Understanding the Crucial Role of Context in Safeguarding AI-Accessible Data
                          https://www.truthinit.com/index.php/channel/2037/understanding-the-crucial-role-of-context-in-safeguarding-ai-accessible-data/
                        • 07/21/2026
                          04:00 AM
                          07/21/2026
                          Strategies for Managing AI Governance and Securing App-to-LLM API Traffic
                          https://www.truthinit.com/index.php/channel/1967/strategies-for-managing-ai-governance-and-securing-app-to-llm-api-traffic/
                        • 07/22/2026
                          06:30 AM
                          07/22/2026
                          Insights and Strategies in Data Protection and Privacy Management
                          https://www.truthinit.com/index.php/channel/2000/insights-and-strategies-in-data-protection-and-privacy-management/
                        • 07/22/2026
                          01:00 PM
                          07/22/2026
                          Insights from Attackers During the FIFA World Cup: A HUMAN Dialogue
                          https://www.truthinit.com/index.php/channel/2029/insights-from-attackers-during-the-fifa-world-cup-a-human-dialogue/
                        • 07/28/2026
                          01:00 PM
                          07/28/2026
                          Illumio + Netskope: Zero Trust in the Age of AI Autonomy
                          https://www.truthinit.com/index.php/channel/2031/illumio-netskope-zero-trust-in-the-age-of-ai-autonomy/
                        • 07/29/2026
                          04:00 AM
                          07/29/2026
                          Real-Time Strategies for Safeguarding Against Prompt Injections
                          https://www.truthinit.com/index.php/channel/1968/real-time-strategies-for-safeguarding-against-prompt-injections/
                        • 07/29/2026
                          01:00 PM
                          07/29/2026
                          Ask Your Cloud Anything: Unlocking Governance Silos in your Environments
                          https://www.truthinit.com/index.php/channel/2048/ask-your-cloud-anything-unlocking-governance-silos-in-your-environments/
                        • 08/19/2026
                          12:00 PM
                          08/19/2026
                          Becoming Agent Ready: Insights from Cyera's Expertise
                          https://www.truthinit.com/index.php/channel/2036/becoming-agent-ready-insights-from-cyeras-expertise/
                        • 09/02/2026
                          12:00 PM
                          09/02/2026
                          Unified Data Security in Action: Uncover, Analyze, and Resolve Threats
                          https://www.truthinit.com/index.php/channel/2045/unified-data-security-in-action-uncover-analyze-and-resolve-threats/
                        • 09/30/2026
                          04:00 AM
                          09/30/2026
                          AI Command Center: Optimizing Visibility and Control in Your Operations
                          https://www.truthinit.com/index.php/channel/2024/ai-command-center-optimizing-visibility-and-control-in-your-operations/
                        Truth in IT
                        • Sponsor
                        • About Us
                        • Terms of Service
                        • Privacy Policy
                        • Contact Us
                        • Preference Management
                        Desktop version
                        Standard version