Conversation
|
@blueorangutan package |
|
@DaanHoogland a [SL] Jenkins job has been kicked to build packages. It will be bundled with KVM, XenServer and VMware SystemVM templates. I'll keep you posted as I make progress. |
Codecov Report✅ All modified and coverable lines are covered by tests.
Additional details and impacted files@@ Coverage Diff @@
## 4.22 #13738 +/- ##
=============================================
- Coverage 17.69% 3.69% -14.01%
=============================================
Files 5925 449 -5476
Lines 533534 38176 -495358
Branches 65273 7072 -58201
=============================================
- Hits 94421 1409 -93012
+ Misses 428434 36580 -391854
+ Partials 10679 187 -10492
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
Packaging result [SF]: ✖️ el8 ✖️ el9 ✖️ debian ✖️ suse15. SL-JID 18704 |
|
@poddm can you look at the build errors, please? |
Description
This PR fixes HA-enabled KVM instances getting stuck in the
Runningstate after anout-of-band stop (e.g. the QEMU process is OOM-killed or crashes on the host).
Actual behaviour (before):
When the power-state sync detected that a running, HA-enabled KVM instance was no longer
present on the host, the management server called
HighAvailabilityManager.scheduleRestart(vm, true)fromhandlePowerOffReportWithNoPendingJobsOnVM(). That path invokesadvanceStop(), whichsubmits a VM work job and blocks waiting for it to complete. Because this handler runs
on the
AgentManager-Handlerthread — a context that does not dispatch VM work jobs — thecall blocks indefinitely. The instance was left in
Running, and the follow-up HAinvestigation (host still up, domain already gone) concluded "no need to restart".
Additionally, a libvirt
VIR_DOMAIN_CRASHEDstate was not mapped to a power state, so acrashed domain was reported as
PowerUnknownand only detected later via the "missing VM"threshold, delaying recovery.
Expected behaviour (after):
An HA-enabled KVM instance that stops out-of-band is transitioned to
Stoppedand an HArestart is scheduled promptly, without blocking the AgentManager thread.
Changes:
Management server —
engine/orchestration/.../VirtualMachineManagerImpl.javaIn the HA out-of-band-stop branch of
handlePowerOffReportWithNoPendingJobsOnVM(),replace the blocking
scheduleRestart(vm, true)with a non-blocking path:Stopped(FollowAgentPowerOffReport),op_ha_workwithStep.Scheduledfor the HAworker to pick up asynchronously.
A guard ensures a valid
host_idis present before enqueuing HA work (avoids violatingthe FK constraint), and existing pending HA work is not duplicated.
KVM agent —
plugins/hypervisors/kvm/.../LibvirtComputingResource.javaMap
DomainState.VIR_DOMAIN_CRASHEDtoPowerState.PowerOffso a crashed domain isreported as off immediately instead of being omitted from the report (which delayed
detection via the missing-VM threshold).
The power-state report already confirms the instance is off, so re-investigation is
unnecessary in this path.
Types of changes
Feature/Enhancement Scale or Bug Severity
Feature/Enhancement Scale
Bug Severity
Screenshots (if appropriate):
N/A
How Has This Been Tested?
RunningtoStopped, a row is created inop_ha_workwithStep.Scheduled, and the HA worker restarts the instance.AgentManager-Handlerthread no longer blocks (no stuck power-off handler),and that
management-server.logno longer shows the instance stuck inRunningafter theinvestigation.
VIR_DOMAIN_CRASHEDdomain is now reported asPowerOffand triggers the samerecovery path without waiting for the missing-VM threshold.
How did you try to break this feature and the system with this change?
instance is still recovered rather than being deemed "alive" by the investigator.
host_id/last_host_idis unavailable: HA work is not enqueued(guarded) to avoid violating the
op_ha_workFK constraint.instance.
the branch is guarded by
vm.getState() == State.Runningand the hypervisor-type checks.