Truth in IT
    • Sign In
    • Register
        • Videos
        • Channels
        • Pages
        • Galleries
        • News
        • Events
        • All
Truth in IT Truth in IT
  • Data Management ▼
    • Converged Infrastructure
    • DevOps
    • Networking
    • Storage
    • Virtualization
  • Cybersecurity ▼
    • Application Security
    • Backup & Recovery
    • Data Security
    • Identity & Access Management (IAM)
    • Zero Trust
    • Compliance & GRC
    • Endpoint Security
  • Cloud ▼
    • Hybrid Cloud
    • Private Cloud
    • Public Cloud
  • Webinar Library
  • TiPs
  • DRAW

OneLLM: Native AI Inference in OpenNebula Sunstone

Open Nebula
06/13/2026
0 (0%)
Share
  • Comments
  • Download
  • Transcript
Report Like Favorite
  • Share/Embed
  • Email
Link
Embed

Transcript


In this screencast, we will show the preview of the upcoming 1LLM feature that brings the AI Inference directly into the platform. Running AI Inference workloads on-premises might easily become challenging. Both the administrators and users often struggle with the lack of unification across the process, finding themselves managing GPU servers, model weights, and inference software outside of the graphical UI. In the upcoming release of Open Nebula, we are addressing these challenges by adding an AI Inference section directly to Sunstone, a single place to define hardware profiles, curate AI models, and deploy production-ready models with just a few clicks. 1LLM is a new section inside Sunstone that brings AI Inference into the platform. No external tools, no separate infrastructure to manage. There are two perspectives in this demonstration. First, the admin, who sets the things up, then the tenant, who puts them to work. As an admin, you start by defining instance types. Each one specifies the GPU, the VRAM, the compute tier, and the model it can handle. Small for lightweight use, medium for most workloads, large for the biggest models. Every type is fully specified. H100, 10 gigs of VRAM, the exact model size range it supports. Define it once, reuse it everywhere. The second thing an admin manages is the model catalog, the library of AI models available inside the data center. Models are downloaded, versioned, and access controlled. Tenants only see what's ready. The admin decides what's available. From here, this is what a tenant experiences. They pick a ready model, pick their instance type, and deploy. OpenAbility handles the rest. It provisions the virtual machine, loads the weights, and starts the Inference engine. No manual steps, no SSH, no scripts. When it's live, the tenant gets an OpenAI-compatible API endpoint. Any app already leveraging the OpenAI SDK connects with zero code changes. And they can test it right here, inside Sunstone. A live conversation with the model. No external tooling needed. Admin sets it up, tenant puts it to work. One platform, from the bare metal, to a virtual machine. And they can test it right here, inside Sunstone. A live conversation with the model. Admin sets it up, tenant puts it to work. One platform, from the bare metal, to a live AI endpoint. And this concludes this feature preview demonstration. Thank you for watching, and see you in the next screencast.

TL;DR

  • OneLLM brings native AI inference capabilities into OpenNebula's Sunstone GUI, eliminating the need for external tools or separate infrastructure management for on-premises AI workloads.
  • Administrators define reusable hardware profiles specifying GPU types, VRAM, and supported model sizes, then curate a versioned model catalog with granular access controls for tenant consumption.
  • Tenants deploy production-ready inference endpoints by selecting pre-configured models and instance types, with OpenNebula automatically provisioning VMs, loading weights, and exposing OpenAI-compatible APIs for zero-code integration.

Summary

This demonstration previews OneLLM, an upcoming OpenNebula feature that integrates AI inference capabilities directly into the Sunstone GUI. The feature addresses common challenges organizations face when running on-premises AI workloads by eliminating the need to manage GPU servers, model weights, and inference software outside the platform. OneLLM provides a unified interface where administrators can define hardware profiles with specific GPU and VRAM configurations, curate AI model catalogs with version control and access management, and enable tenants to deploy production-ready inference endpoints with minimal configuration. The system provisions virtual machines, loads model weights, and exposes OpenAI-compatible API endpoints that work with existing SDK integrations, allowing organizations to run AI inference workloads alongside their existing cloud and edge infrastructure without external tooling or manual intervention.

Chapters

0:00 - Introduction to OneLLM
0:52 - Admin Perspective: Instance Types
1:35 - Admin Perspective: Model Catalog
1:49 - Tenant Workflow: Deployment

Key Quotes

0:16 "Both the administrators and users often struggle with the lack of unification across the process, finding themselves managing GPU servers, model weights, and inference software outside of the graphical UI."
0:33 "We are addressing these challenges by adding an AI Inference section directly to Sunstone, a single place to define hardware profiles, curate AI models, and deploy production-ready models with just a few clicks."
2:14 "Any app already leveraging the OpenAI SDK connects with zero code changes."

FAQ

What problem does OneLLM solve for organizations running AI workloads on-premises?

OneLLM addresses the lack of unification in managing on-premises AI inference by bringing GPU server management, model weights, and inference software directly into OpenNebula's Sunstone GUI. This eliminates the need to manage these components separately outside the platform, providing a single interface for defining hardware profiles, curating AI models, and deploying production-ready endpoints.

How does OneLLM handle compatibility with existing AI applications?

OneLLM exposes OpenAI-compatible API endpoints for deployed models, allowing any application already using the OpenAI SDK to connect with zero code changes. This ensures seamless integration with existing AI workflows and tooling without requiring custom development or API adaptation.


Categories:
  • » Cybersecurity » Cloud Security
  • » Data Protection
Channels:
News:
Events:
Tags:
  • AI & Machine Learning
  • Cloud Security
  • Technical Deep Dive
  • Demo
  • Getting Started
  • AI inference
  • GPU resource management
  • LLM deployment
  • on-premises AI infrastructure
  • model catalog management
  • OpenAI API compatibility
  • cloud management platform
  • multi-tenancy
Show more Show less

Browse videos

  • Related
  • Featured
  • By date
  • Most viewed
  • Top rated
  •  

              Video's comments: OneLLM: Native AI Inference in OpenNebula Sunstone

              XStreaminars (watch here)

              • Jul
                28

                Illumio + Netskope: Zero Trust in the Age of AI Autonomy

                07/28/202601:00 PM ET
                • Jul
                  29

                  Ask Your Cloud Anything: Unlocking Governance Silos in your Environments

                  07/29/202601:00 PM ET
                  More events

                  Industry Events (watch there)

                  • Jul
                    14

                    Crafting a Championship-Caliber Security Team for Lasting Defense

                    07/14/202601:00 PM ET
                    • Jul
                      14

                      Understanding the Crucial Role of Context in Safeguarding AI-Accessible Data

                      07/14/202602:00 PM ET
                      • Jul
                        22

                        Insights from Attackers During the FIFA World Cup: A HUMAN Dialogue

                        07/22/202601:00 PM ET
                        More events

                        Upcoming Webinar Calendar

                        • 07/14/2026
                          01:00 PM
                          07/14/2026
                          Crafting a Championship-Caliber Security Team for Lasting Defense
                          https://www.truthinit.com/index.php/channel/2025/crafting-a-championship-caliber-security-team-for-lasting-defense/
                        • 07/14/2026
                          02:00 PM
                          07/14/2026
                          Understanding the Crucial Role of Context in Safeguarding AI-Accessible Data
                          https://www.truthinit.com/index.php/channel/2037/understanding-the-crucial-role-of-context-in-safeguarding-ai-accessible-data/
                        • 07/21/2026
                          04:00 AM
                          07/21/2026
                          Strategies for Managing AI Governance and Securing App-to-LLM API Traffic
                          https://www.truthinit.com/index.php/channel/1967/strategies-for-managing-ai-governance-and-securing-app-to-llm-api-traffic/
                        • 07/22/2026
                          06:30 AM
                          07/22/2026
                          Insights and Strategies in Data Protection and Privacy Management
                          https://www.truthinit.com/index.php/channel/2000/insights-and-strategies-in-data-protection-and-privacy-management/
                        • 07/22/2026
                          01:00 PM
                          07/22/2026
                          Insights from Attackers During the FIFA World Cup: A HUMAN Dialogue
                          https://www.truthinit.com/index.php/channel/2029/insights-from-attackers-during-the-fifa-world-cup-a-human-dialogue/
                        • 07/28/2026
                          01:00 PM
                          07/28/2026
                          Illumio + Netskope: Zero Trust in the Age of AI Autonomy
                          https://www.truthinit.com/index.php/channel/2031/illumio-netskope-zero-trust-in-the-age-of-ai-autonomy/
                        • 07/29/2026
                          04:00 AM
                          07/29/2026
                          Real-Time Strategies for Safeguarding Against Prompt Injections
                          https://www.truthinit.com/index.php/channel/1968/real-time-strategies-for-safeguarding-against-prompt-injections/
                        • 07/29/2026
                          01:00 PM
                          07/29/2026
                          Ask Your Cloud Anything: Unlocking Governance Silos in your Environments
                          https://www.truthinit.com/index.php/channel/2048/ask-your-cloud-anything-unlocking-governance-silos-in-your-environments/
                        • 08/19/2026
                          12:00 PM
                          08/19/2026
                          Becoming Agent Ready: Insights from Cyera's Expertise
                          https://www.truthinit.com/index.php/channel/2036/becoming-agent-ready-insights-from-cyeras-expertise/
                        • 09/02/2026
                          12:00 PM
                          09/02/2026
                          Unified Data Security in Action: Uncover, Analyze, and Resolve Threats
                          https://www.truthinit.com/index.php/channel/2045/unified-data-security-in-action-uncover-analyze-and-resolve-threats/
                        • 09/30/2026
                          04:00 AM
                          09/30/2026
                          AI Command Center: Optimizing Visibility and Control in Your Operations
                          https://www.truthinit.com/index.php/channel/2024/ai-command-center-optimizing-visibility-and-control-in-your-operations/
                        Truth in IT
                        • Sponsor
                        • About Us
                        • Terms of Service
                        • Privacy Policy
                        • Contact Us
                        • Preference Management
                        Desktop version
                        Standard version