Set up volume/hardware monitoring¶
This page explains how to configure volume and hardware monitoring in config/volumes.yaml, including check frequency, test type, and alert thresholds.
You set up volume/hardware monitoring in config/volumes.yaml. For each volume, you can specify whether or not it is to be checked, the subdirectory where the check is performed, the check and reporting frequency, the test type, and the impact of a failed test. Publishes from the hardware monitor to MonVols.yaml happen either at your configured frequency or immediately if a serious condition is detected.
The rules governing publish events are extensible but come with a set of defaults.
| Rule | Description |
|---|---|
| critical transition | Publish if a volume enters or leaves a critical state. |
| monotonicity | Publish if a particular check on a volume continuously gets slower. |
| sustained warning | Publish if a particular check stays at the warning level. |
You can define your volume checks at the node level or with any other set of overlays that you want. For example, each node can be set up differently with different values. Edits to this table are dynamic and take effect immediately without the need to restart a process. The name column has to be unique.
You can look up your configured monitoring tests in config/volumes.yaml. Sample entries might be as follows (the zfact column, or Z-factor, is described in more detail below):
| Name | subdir | testType | zfact | impact | Comment |
|---|---|---|---|---|---|
| MyVolume1 | ABC | R | 10 | 50 | If the Z-factor score is 10 or greater, a warning message may be generated. If the test fails, the user-assigned impact will be 50 and a fail-over may be triggered. |
| MyVolume2 | ABC | W | 5 | 70 | This test is similar to the previous test. The only differences are that it is a write test rather than a read test and the Z-factor is lower (that is, more sensitive) and the impact higher because a faulty write is considered more serious than a faulty read for this subdirectory. |
| MyVolume3 | DEF | WRV | 5 | 85 | A critical subdirectory with a low (that is, very sensitive) Z-factor and a very high impact value on a write/read/verify test. |
| MyVolume4 | GHI | WR | 20 | 25 | A non-critical subdirectory with a high Z-factor value and a low impact value on a write/read test. |
| - | - | WRV | 5 | 50 | Default values applicable to all volumes on a node if you do not create any records in volumes.yaml and if you set isEnabled to true. |
volumes:
values:
<default>:
isEnabled: false
chkFreq: 10
repFreq: 60
testType: wrv
wlen: 8
tmo: 3
intv: 200
tries: 3
zfact: 5
impact: 50
DATA:
isEnabled: true
subdir: "${MISC}/voltest" # See miscDataDir in systemParams.yaml
RT:
isEnabled: true
subdir: "${EMSLOGDIR}/voltest" # See emsLogDir in systemParams.yaml
impact: 100
| Column | Definition | Example |
|---|---|---|
| name | The name of the volume test. This name must be unique across all nodes. | |
| isEnabled | Indicates whether or not the test is enabled | false |
| subdir | Relative subdirectory where read/write/verify checks are performed | |
| chkFreq | Frequency of checks in milliseconds | 10 |
| repFreq | Reporting frequency in seconds | 60 |
| testType | R (Read) W (Write) WR (Read/Write) WRV (Read/Write/Verify) | WRV |
| wlen | Number of bytes to write | 8 |
| rlen | Number of bytes to read. If not specified, this value defaults to wlen |
|
| tmo | Timeout interval in seconds | 3 |
| intv | Retry interval in milliseconds | 200 |
| tries | Number of times to attempt the test before considering it to have failed | 3 |
| zfact | Z-score standard deviation factor for calculating the dynamic warning threshold. KX Sensors works out what "normal" is based on various tests carried out continuously and uses that value to compute a dynamic trigger threshold that adjusts with system and hardware performance. A high value (for example, 20) indicates a very insensitive warning threshold, while a low value (for example, 5) indicates a highly sensitive warning threshold with very little deviation from the mean being tolerated. | 5 |
| impact | This is a user-defined value that indicates how serious you consider a failed test to be. Valid values are 0 to 100. The higher the value, the more serious you consider the failure to be. For example, if you assign an impact of 80 to a failed test on subdirectory ABC, you are indicating that the loss of ABC would have a serious impact on the overall operation of the system. If, on the other hand, you assign an impact of 30 to the same test on the same subdirectory, you are indicating that ABC is not a critical resource and the impact of its loss would, in isolation, be low. | 50 |
In config/systemParams.yaml, you define the critical threshold for assessing overall failure impact and the maximum number of samples used for analysis. If the total test scores for all monitored volumes on the node exceed this threshold, KX Sensors marks a node as failed and triggers an AUTOFOVPROC type fail-over.
| Parameter | Description | Example |
|---|---|---|
| haVolImpactThr | The threshold that will trigger feeds to fail-over from a node when the monitored volumes on that node are deemed to be in a critical state. For example: WRV test on volume 1 failed (impact = 30); R test on volume 2 failed (impact = 20); WRV test on volume 3 failed (impact = 35). Total impact = 85, which exceeds the specified critical threshold of 75. | 75 |
| haVolImpactQueryThr | The threshold that will change GW routing preferences when the monitored volumes on that node are deemed to be in a critical state. Should be lower than your haVolImpactThr value. |
50 |
| monHwWindow | The maximum number of sample records held for analysis at any given time | 1000 |