Skip to content

Set up volume/hardware monitoring

This page explains how to configure volume and hardware monitoring in config/volumes.yaml, including check frequency, test type, and alert thresholds.

You set up volume/hardware monitoring in config/volumes.yaml. For each volume, you can specify whether or not it is to be checked, the subdirectory where the check is performed, the check and reporting frequency, the test type, and the impact of a failed test. Publishes from the hardware monitor to MonVols.yaml happen either at your configured frequency or immediately if a serious condition is detected.

The rules governing publish events are extensible but come with a set of defaults.

Rule Description
critical transition Publish if a volume enters or leaves a critical state.
monotonicity Publish if a particular check on a volume continuously gets slower.
sustained warning Publish if a particular check stays at the warning level.

You can define your volume checks at the node level or with any other set of overlays that you want. For example, each node can be set up differently with different values. Edits to this table are dynamic and take effect immediately without the need to restart a process. The name column has to be unique.

You can look up your configured monitoring tests in config/volumes.yaml. Sample entries might be as follows (the zfact column, or Z-factor, is described in more detail below):

Name subdir testType zfact impact Comment
MyVolume1 ABC R 10 50 If the Z-factor score is 10 or greater, a warning message may be generated. If the test fails, the user-assigned impact will be 50 and a fail-over may be triggered.
MyVolume2 ABC W 5 70 This test is similar to the previous test. The only differences are that it is a write test rather than a read test and the Z-factor is lower (that is, more sensitive) and the impact higher because a faulty write is considered more serious than a faulty read for this subdirectory.
MyVolume3 DEF WRV 5 85 A critical subdirectory with a low (that is, very sensitive) Z-factor and a very high impact value on a write/read/verify test.
MyVolume4 GHI WR 20 25 A non-critical subdirectory with a high Z-factor value and a low impact value on a write/read test.
- - WRV 5 50 Default values applicable to all volumes on a node if you do not create any records in volumes.yaml and if you set isEnabled to true.
volumes:
  values:
    <default>:
      isEnabled: false
      chkFreq: 10
      repFreq: 60
      testType: wrv
      wlen: 8
      tmo: 3
      intv: 200
      tries: 3
      zfact: 5
      impact: 50
    DATA:
      isEnabled: true
      subdir: "${MISC}/voltest" # See miscDataDir in systemParams.yaml
    RT:
      isEnabled: true
      subdir: "${EMSLOGDIR}/voltest" # See emsLogDir in systemParams.yaml
      impact: 100
Column Definition Example
name The name of the volume test. This name must be unique across all nodes.
isEnabled Indicates whether or not the test is enabled false
subdir Relative subdirectory where read/write/verify checks are performed
chkFreq Frequency of checks in milliseconds 10
repFreq Reporting frequency in seconds 60
testType R (Read) W (Write) WR (Read/Write) WRV (Read/Write/Verify) WRV
wlen Number of bytes to write 8
rlen Number of bytes to read. If not specified, this value defaults to wlen
tmo Timeout interval in seconds 3
intv Retry interval in milliseconds 200
tries Number of times to attempt the test before considering it to have failed 3
zfact Z-score standard deviation factor for calculating the dynamic warning threshold. KX Sensors works out what "normal" is based on various tests carried out continuously and uses that value to compute a dynamic trigger threshold that adjusts with system and hardware performance. A high value (for example, 20) indicates a very insensitive warning threshold, while a low value (for example, 5) indicates a highly sensitive warning threshold with very little deviation from the mean being tolerated. 5
impact This is a user-defined value that indicates how serious you consider a failed test to be. Valid values are 0 to 100. The higher the value, the more serious you consider the failure to be. For example, if you assign an impact of 80 to a failed test on subdirectory ABC, you are indicating that the loss of ABC would have a serious impact on the overall operation of the system. If, on the other hand, you assign an impact of 30 to the same test on the same subdirectory, you are indicating that ABC is not a critical resource and the impact of its loss would, in isolation, be low. 50

In config/systemParams.yaml, you define the critical threshold for assessing overall failure impact and the maximum number of samples used for analysis. If the total test scores for all monitored volumes on the node exceed this threshold, KX Sensors marks a node as failed and triggers an AUTOFOVPROC type fail-over.

Parameter Description Example
haVolImpactThr The threshold that will trigger feeds to fail-over from a node when the monitored volumes on that node are deemed to be in a critical state. For example: WRV test on volume 1 failed (impact = 30); R test on volume 2 failed (impact = 20); WRV test on volume 3 failed (impact = 35). Total impact = 85, which exceeds the specified critical threshold of 75. 75
haVolImpactQueryThr The threshold that will change GW routing preferences when the monitored volumes on that node are deemed to be in a critical state. Should be lower than your haVolImpactThr value. 50
monHwWindow The maximum number of sample records held for analysis at any given time 1000

Next steps