Set up a watch¶
This page describes how to set up statistical and log watches in KX Sensors, using config/watch.yaml and config/logWatch.yaml to monitor threshold violations and log events.
There are two types of watches in KX Sensors:
- Statistical watches are triggered when a threshold violation occurs for a value in a monitoring table that is being watched.
- Log watches are triggered when certain logging levels within a logging facility are reported by a process that is being watched.
Set up a statistical watch¶
-
You set up statistical watches in
config/watch.yaml, using thewarnThr,critThrandisHighattributes to specify your threshold violation conditions:watch: values: name: watch1 process: dbw description: DBW memory watch table: column: memUsed warnthr: 150000 criThr: 2000000 duration: 50 suppressBusy: rptCount: rptInterval:
If you are monitoring a state or condition (for example, looking for a process state of "inactive", "error" or "no response"):
- Leave
durationblank for an instantaneous alert. -
Enter the numeric value of the states in the
warnThrandcritThrproperties. For example, fromconfig/enum/procState.yaml, "inactive" is 10 and "no response" is 12.watch: values: name: procState1 process: node: description: General process state watch table: column: procState warnthr: 10 criThr: 12 isHigh: true duration: suppressBusy: rptCount: rptInterval:
| Parameter | Description |
|---|---|
| name | The name of your statistical watch. |
| process | The process name to which the watch applies, or null for all processes. |
| node | The specific node to be monitored, or null for all nodes. |
| description | Optional description of the watch. |
| isEnabled | Whether or not the watch entry is enabled. The default for this value is true. |
| alertName | The name of the alert that is to be invoked when the watch triggers. See Set up a watch-based alert. |
| table | The name of the monitoring table to which the watch applies. |
| column | The column from the monitoring table that you want to watch. |
| warnThr | The warning threshold value. If you are setting a watch on a process state or any other numeric enumeration, use the numeric value here. |
| criThr | The critical threshold value. If you are setting a watch on a process state or any other numeric enumeration, use the numeric value here. |
| isHigh | If true, the warnThr and critThr values are maximums and a warning is triggered if the monitoring table values reach or exceed these values. If false, the warnThr and critThr values are minimums and a warning is triggered if the monitoring table values fall below these values. The default for this value is true. |
| duration | Duration in seconds of condition before the watch is triggered. If you enter zero or leave this column blank, KX Sensors captures a single instantaneous reading for that moment in time. If you enter a non-zero value, KX Sensors triggers the watch only if the value is continuously out-of-bounds of at least one threshold for the specified duration. |
| suppressBusy | If true, suspend the watch if the underlying process is busy, initializing, recovering or upgrading. |
| rptCount | Maximum number of times watch triggers an alert within the interval specified by rptInterval (0 indicates no limit). |
| rptInterval | Interval in seconds for limiting the number of alerts specified by rptCount (0 indicates no limit). |
Note
Instantaneous readings are not recommended for monitoring conditions such as memory usage, because they might report on isolated spikes and dips that are not significant.
Set up a log watch¶
You set up log watches in config/logWatch.yaml to report on the log level recorded by a given process. Dedicated log files are produced for each process in KX Sensors; subject to the configured per-facility log levels for the process, they contain a complete record of every significant action the process took and every situation that it encountered.
Log watches let you monitor log levels such as WARN or ERROR in your log files, or monitor for specific strings in the log message text. For each record in config/logWatch.yaml, you define the severity of the log message and, optionally, the alert string.
logWatch:
values:
name: monLogWatch
process:
description: Log watch for MON errors (all processes)
level: error
facility: mon
rptCount:
rptInterval:
alertStr:
| Parameter | Description |
|---|---|
| name | The name of your log watch entry. |
| process | The process name to which the watch applies, or null for all processes. |
| node | The node to which the watch applies, or null for all nodes. |
| description | Optional description of the watch. |
| isEnabled | Whether or not the log watch is enabled. The default for this value is true. |
| alertName | See Set up a watch-based alert. |
| level | The log level or severity of the error as defined in config/enum/logLevel.yaml. This can be one of warn, error or fatal and must be specified. |
| facility | Log facility to which the watch applies, or null for all facilities. |
| rptCount | Maximum number of times watch triggers an alert within the interval specified by rptInterval (0 indicates no limit). |
| rptInterval | Interval in seconds for limiting the number of alerts specified by rptCount (0 indicates no limit). |
| alertStr | If you enter a regular expression or string to match (for example, not found), an alert is generated only if the expression occurs within the log message. See Set up a watch-based alert. |