Truth in IT
    • Sign In
    • Register
        • Videos
        • Channels
        • Pages
        • Galleries
        • News
        • Events
        • All
Truth in IT Truth in IT
  • Data Management ▼
    • Converged Infrastructure
    • DevOps
    • Networking
    • Storage
    • Virtualization
  • Cybersecurity ▼
    • Application Security
    • Backup & Recovery
    • Data Security
    • Identity & Access Management (IAM)
    • Zero Trust
    • Compliance & GRC
    • Endpoint Security
  • Cloud ▼
    • Hybrid Cloud
    • Private Cloud
    • Public Cloud
  • Webinar Library
  • TiPs
  • DRAW

VM High Availability Setup & Testing in OpenNebula

Open Nebula
10/02/2026
0 (0%)
Share
  • Comments
  • Download
  • Transcript
Report Like Favorite
  • Share/Embed
  • Email
Link
Embed

Transcript


In this video we will show how to set up Virtual Machine High Availability and verify that a virtual machine will be migrated to another host when the initial host is powered down. Virtual Machine High Availability is a powerful feature with a lot of fine-tuning. However, it requires shared storage to be in place. Without shared storage, virtual machine can be respawned on a different host, but will lose all the data. This may be suitable for stateless machines where the data is not a concern. However, true Virtual Machine High Availability is only available in an environment with shared storage. This functionality relies on the Hooks subsystem. The Hooks subsystem is powerful automation and integration functionality available in OpenEBOLA. This functionality allows the execution of any script based on a virtual machine, host, or image state change, or an API call. To learn more about the Hooks system, please visit the official OpenEBOLA documentation. Fencing is important to prevent split-brain issues. It is crucial to ensure that the host is fenced. You must edit the fence-host.shell script and supply it with the correct fencing command. Virtual Machine High Availability also relies on the built-in monitoring of hosts and virtual machines. You may need to adjust the monitoring settings as well to achieve faster failover. The environment for this demonstration consists of one host with OpenEBOLA frontend software, two KVM hosts managed by OpenEBOLA, each host has a virtual machine running. One virtual machine is a Flask-based web application, and the second one is the database that stores data for this app. Shared storage that exposes system and image data stores and is mounted to each node, including the frontend, using the NFS protocol. To enable Virtual Machine High Availability, we must configure the hook first. You can find the hook template in the official OpenEBOLA documentation. There are additional flags that can be passed to the script, each of them is described in the documentation as well. To create a hook, we're going to create a file and insert the hook template contents. The default hook settings are very cautious, which may lead to longer virtual machine migration times. In the scope of this demonstration, we're going to make them much more aggressive. Additionally, there is no need to fence the host because we're going to power it down ourselves, hence the no-fencing flag has been added. Please note that the no-fencing flag works in a demo environment but should never be used in the production environment. The hook will execute the host error script if any host changes its state to error. Save the file and then use the one-hook create command line tool and point to the file template. Now the hook has been created successfully. Using the one-hook list command, we can find all hooks configured in this environment. Using the one-hook show command, we print the details of a specific hook and its execution log. In your production environment, please don't forget to configure the fencing mechanism by editing the following file. Once again, you shouldn't run virtual machine high availability without fencing in a production environment. Now it's time to test and verify whether it works as expected. We have two virtual machines running on two different hosts. These two represent a typical application with a database, a flask-based web application is serving the data while the database stores it. In this test scenario, the web application will remain running while the host with the database virtual machine is brought down. Let's navigate to the web page and confirm that the database isn't empty and contains the data. The data is there. In our test scenario, we are supposed to lose connectivity with the database for a brief period, but once it's back online, the data must not be lost. Now let's open a VNC connection to the application virtual machine and log in using the configured credentials. Once logged in, let's start pinging the second virtual machine. We are going to use the ping command to capture the moment when the virtual machine goes down along with the host and when it successfully started on another host. Switch back to the command line interface, this time to the KVM node that will be powered down. Now let's execute the shutdown command. Now let's observe the state change of a host using the OpenAbility built-in command line tool one-host-top. Let's switch to the VNC session and see whether the virtual machine is still responding to sent ICMP packets. As we can see, the virtual machine stopped responding. Back to the command line, we can confirm that the host state switched to error. This will trigger the execution of the high availability hook. Now let's look back at the ICMP packets. We can see that the virtual machine starts to respond to ICMP requests once again. The virtual machine went down after the 44th request and started responding on the 66th, meaning that only 21 ICMP packets went unresponded. Switching to the application, we can see that the database is not yet running as the operating system and database software are in the process of starting up. By trying one more time, we can confirm that the database software is back online and the data has been preserved. By switching to the SunStone web UI, we can see that both virtual machines are now running on the same host. This is the result of the virtual machine high availability hook execution. And this concludes our screencast. We have shown how to configure the virtual machine HA capability of OpenNebula and verify that the functionality works as expected. Thank you for watching and we will see you in the next screencast.

TL;DR

  • VM High Availability in OpenNebula requires shared storage — without it, VMs can be respawned elsewhere but will lose all persistent data on failover.
  • HA is implemented via the Hooks subsystem, which executes a host-error script automatically when a host transitions to an error state.
  • Fencing must be configured in production environments to prevent split-brain issues; the no-fencing flag demonstrated here is strictly for lab use.
  • A live failover test showed only 21 missed ICMP packets before the database VM resumed on a surviving host, with all application data preserved.

Summary

This screencast walks through configuring and validating Virtual Machine High Availability (VM HA) in OpenNebula, demonstrating how a running VM is automatically migrated to a healthy host when its original KVM host is powered down. The demonstration environment consists of an OpenNebula frontend, two KVM hypervisor nodes, and shared NFS-backed storage — a prerequisite for true VM HA, since without shared storage a VM can only be respawned on another host at the cost of losing all disk data. The HA mechanism is built on OpenNebula's Hooks subsystem, which triggers automation scripts in response to host or VM state changes. A hook template is created from the official documentation, configured with aggressive failover timing for demo purposes, and registered using the one-hook CLI tool. The no-fencing flag is used in this demo but is explicitly called out as unsuitable for production — real deployments must configure the fence-host.sh script to prevent split-brain scenarios. The live test shuts down one KVM host running a Flask web application's database VM, captures the outage window via ICMP ping, and confirms that the VM restarts on the surviving host with all data intact. Only 21 ICMP packets went unanswered during the failover window, and the Sunstone web UI confirms both VMs are subsequently running on the same host.

Chapters

0:00 - Introduction & Prerequisites
0:38 - Hooks Subsystem Overview
1:18 - Demo Environment Setup
1:48 - Configuring the HA Hook
3:33 - Live Failover Test
6:18 - Results & Wrap-Up

Key Quotes

0:21 "It requires shared storage to be in place. Without shared storage, virtual machine can be respawned on a different host, but will lose all the data."
0:38 "This functionality relies on the Hooks subsystem. The Hooks subsystem is powerful automation and integration functionality available in OpenNebula."
2:27 "Please note that the no-fencing flag works in a demo environment but should never be used in the production environment."
5:44 "The virtual machine went down after the 44th request and started responding on the 66th, meaning that only 21 ICMP packets went unresponded."
6:04 "By trying one more time, we can confirm that the database software is back online and the data has been preserved."

FAQ

Is shared storage mandatory for VM High Availability in OpenNebula?

Yes. True VM HA requires shared storage mounted to all nodes, including the frontend, so the VM's disk image is accessible from any host. Without it, OpenNebula can respawn the VM on another host, but all data will be lost — suitable only for stateless workloads.

What is the no-fencing flag and when should it be used?

The no-fencing flag disables the host fencing step in the HA hook, which is acceptable in a controlled demo where the host is manually shut down. It must never be used in production, where fencing is required to guarantee the original host is truly offline before the VM is started elsewhere, preventing split-brain data corruption.


Categories:
  • » Data Protection » Backup & Recovery
  • » Data Protection
Channels:
News:
Events:
Tags:
  • Cloud Infrastructure
  • Data Protection
  • Demo
  • How-To
  • Technical Deep Dive
  • Backup & Recovery
  • Virtual Machine High Availability
  • OpenNebula Hooks Subsystem
  • KVM Hypervisor Management
  • Shared Storage with NFS
  • Host Fencing and Split-Brain Prevention
  • VM Migration
Show more Show less

Browse videos

  • Related
  • Featured
  • By date
  • Most viewed
  • Top rated
  •  

              Video's comments: VM High Availability Setup & Testing in OpenNebula

              Industry Events (Sponsor Hosted)

              • Oct
                13

                Ensuring Compliance Through Audit Evidence: From CJIS to FERPA

                10/13/202601:00 PM ET
                • Oct
                  15

                  Risk in Real Time Demo Series: Virtual Patching: Protection at the Speed of Exploitation

                  10/15/202611:00 AM ET
                  • Oct
                    20

                    Harnessing Data Governance for AI with Cyera and Snowflake

                    10/20/202611:00 AM ET
                    More events

                    Upcoming Webinar Calendar

                    • 10/13/2026
                      01:00 PM
                      10/13/2026
                      Ensuring Compliance Through Audit Evidence: From CJIS to FERPA
                      https://www.truthinit.com/index.php/channel/2159/ensuring-compliance-through-audit-evidence-from-cjis-to-ferpa/
                    • 10/15/2026
                      11:00 AM
                      10/15/2026
                      Risk in Real Time Demo Series: Virtual Patching: Protection at the Speed of Exploitation
                      https://www.truthinit.com/index.php/channel/1372/risk-in-real-time-demo-series-the-autonomous-era-orchestrating-a-resilient-enterprise/
                    • 10/20/2026
                      11:00 AM
                      10/20/2026
                      Harnessing Data Governance for AI with Cyera and Snowflake
                      https://www.truthinit.com/index.php/channel/2137/harnessing-data-governance-for-ai-with-cyera-and-snowflake/
                    • 10/27/2026
                      01:00 PM
                      10/27/2026
                      The HUMAN Experience: Real-Time Insights into Page Intelligence
                      https://www.truthinit.com/index.php/channel/2139/the-human-experience-real-time-insights-into-page-intelligence/
                    • 11/04/2026
                      11:00 AM
                      11/04/2026
                      Leveraging CISA’s Zero Trust Maturity Model for an AI-Driven Landscape
                      https://www.truthinit.com/index.php/channel/2149/leveraging-cisas-zero-trust-maturity-model-for-an-ai-driven-landscape/
                    • 11/05/2026
                      01:00 PM
                      11/05/2026
                      HUMAN Dialogue: Redefining Authentic Trust in the Agentic Internet
                      https://www.truthinit.com/index.php/channel/2160/human-dialogue-redefining-authentic-trust-in-the-agentic-internet/
                    • 11/19/2026
                      01:00 PM
                      11/19/2026
                      360View: Govern, Secure & Recover Your Microsoft 365 Environment
                      https://www.truthinit.com/index.php/channel/2076/360view-govern-secure-recover-your-microsoft-365-environment/
                    Truth in IT
                    • Sponsor
                    • About Us
                    • Terms of Service
                    • Privacy Policy
                    • Contact Us
                    • Preference Management
                    Desktop version
                    Standard version