New features and enhancements

This release adds improvements related to the following components and concepts.

  • Storage-Based Remediation (SBR) Operator is now generally available in Workload Availability for Red Hat OpenShift.
    • The SBR Operator is an OLM Operator for managing STONITH Block Device (SBD) configurations and remediations for high-availability clustering. The SBR Operator provides automated node remediation when nodes become unresponsive by leveraging shared storage for fencing operations. With the Node Health Check (NHC) Operator which detects that a node is offline, the SBR Operator steps in to power-cycle, reset, or forcefully disconnect the node.
  • Support updated for the kube-rbac-proxy:
    • The usage of kube-rbac-proxy containers is replaced with code-based metrics, endpoint authentication, and authorization. This update applies to the Fence Agents Remediation (FAR) Operator, Self Node Remediation (SNR) Operator, Machine Deletion Remediation (MDR) Operator, and the Node Maintenance Operator (NMO).
  • Update to the metrics configuration for the Node Health Check (NHC) Operator:
    • NHC metrics are now collected using the Platform Prometheus. The User Workload Monitoring is no longer used. This change eliminates manual token management, and adopts secure, in-process mTLS for OpenShift Container Platform (OCP).
  • Workload Availability for Red Hat OpenShift Operators deployment updated:
    • Best practices for Operator deployment are updated to include lifecycle hooks and startup probes. This update applies to the Fence Agents Remediation (FAR) Operator, Self Node Remediation (SNR) Operator, and the Node Health Check (NHC) Operator.
  • A further consideration when using the Self Node Remediation (SNR) Operator / safeTimeToAssumeNodeRebootedSeconds parameter:
    • As the SNR operator relies on a hardware watchdog timer to reboot unhealthy nodes, if a kernel panic occurs, kdump begins capturing the vmcore. However, if this capture process takes longer than the watchdog timeout, the hardware timer forces a reset before kdump finishes, leaving the vmcore incomplete. Increasing the watchdog timeout duration, also requires a corresponding increase in safeTimeToAssumeNodeRebootedSeconds.