Skip to content

Operator-initiated actions

This page explains how to perform manual failover, failback, and forced synchronization using HA maintenance console commands, and how to troubleshoot failover events.

KX Sensors supports a number of operator-initiated actions such as manual failover, manual failback and forced synchronization. These commands are all optional; their purpose is to allow system administrators to make their own determination regarding the health of a node or feed and initiate manual failovers and failbacks on demand.

Perform a manual failover

You can move a feed to another node by means of the failover command if the target node is in at least as healthy a state as the current primary node. The feed will remain on that node as long as that node remains in a healthy state. If you have activated auto-failback and if the current node enters an unhealthy state, the feed will attempt to shift back to the previous node in your node preference sequence.

The node must be in the feed's failover list, must not be the current primary for the feed and must be less preferred than the current primary in the node preference sequence.

If the current acting primary is the least preferred node on your node preference list, you cannot perform a manual failover. Your only option is to perform a manual failback.

Syntax `failover;`<feed>;`<node>
Optional N/A
Arming required Yes
Confirmation required No
q).maint.ha[`sysstate;`feed1]
2025.09.04 14:45:03.124 DEBUG [mymaint] DISC Fetching key /config/haState
feed node|  feedName sndNode ts                          cc mode        event
---------|  -----------------------------------------------------------------
100  A    |  feed1     A       2025.09.04D14:23:43.879628643 2  PRIMARY     HACTL
100  B    |  feed1     B       2025.09.04D14:23:43.888926777 2  SECONDARY   HACTL
q).maint.arm`ha
Now armed: `ha
q).maint.ha[`failover;`feed1;`B]
2025.09.04 14:45:41.162 DEBUG [mymaint] DISC Fetching key /config/haState
Feed 100 (feed1): primary=A nodes=`A`B failover=`A`B
2025.09.04 14:45:41.165 DEBUG [mymaint] DISC Fetching key /config/nodeState
Issuing failover of feed 100 from node A to B
q)2025.09.04 14:45:41.183 DEBUG [mymaint] DISC Received update for /config/haState
2025.09.04 14:45:41.183 INFO  [mymaint] HA New registry HA state:
2025.09.04 14:45:41.183 INFO  [mymaint] HA     feed node|  sndNode ts                          cc mode  event
2025.09.04 14:45:41.183 INFO  [mymaint] HA     ----------|  ------------------------------------------------
2025.09.04 14:45:41.183 INFO  [mymaint] HA     200  A     |  A        2025.09.04D14:23:42.376031010 2  3     1
2025.09.04 14:45:41.183 INFO  [mymaint] HA     200  B     |  B        2025.09.04D14:23:42.755643715 2  4     1
2025.09.04 14:45:41.183 INFO  [mymaint] HA     100  B     |  B        2025.09.04D14:45:41.172759624 3  12    9
2025.09.04 14:45:41.183 INFO  [mymaint] HA     100  A     |  B        2025.09.04D14:45:41.172759624 3  11    9
2025.09.04 14:45:41.183 INFO  [mymaint] HA     0    B     |  B        2025.09.04D14:23:44.850117027 2  4     1
2025.09.04 14:45:41.183 INFO  [mymaint] HA     0    A     |  A        2025.09.04D14:23:44.814655581 2  3     1
2025.09.04 14:45:41.201 DEBUG [mymaint] DISC Received update for /config/haState
2025.09.04 14:45:41.201 INFO  [mymaint] HA New registry HA state:
2025.09.04 14:45:41.201 INFO  [mymaint] HA     feed node|  sndNode ts                          cc mode  event
2025.09.04 14:45:41.201 INFO  [mymaint] HA     ----------|  ------------------------------------------------
2025.09.04 14:45:41.201 INFO  [mymaint] HA     200  A     |  A        2025.09.04D14:23:42.376031010 2  3     1
2025.09.04 14:45:41.201 INFO  [mymaint] HA     200  B     |  B        2025.09.04D14:23:42.755643715 2  4     1
2025.09.04 14:45:41.201 INFO  [mymaint] HA     100  B     |  B        2025.09.04D14:45:41.189042911 4  3     9
2025.09.04 14:45:41.201 INFO  [mymaint] HA     100  A     |  B        2025.09.04D14:45:41.189042911 4  4     9
2025.09.04 14:45:41.201 INFO  [mymaint] HA     0    B     |  B        2025.09.04D14:23:44.850117027 2  4     1
2025.09.04 14:45:41.201 INFO  [mymaint] HA     0    A     |  A        2025.09.04D14:23:44.814655581 2  3     1

q).maint.ha[`sysstate;`feed1]
2025.09.04 14:45:52.373 DEBUG [mymaint] DISC Fetching key /config/haState
feed node|  feedName sndNode ts                          cc mode        event
---------|  -----------------------------------------------------------------
100  A    |  feed1     B       2025.09.04D14:45:41.189042911 4  SECONDARY   MANFOVFEED
100  B    |  feed1     B       2025.09.04D14:45:41.189042911 4  PRIMARY     MANFOVFEED

Perform a manual failback

If auto-failback is not activated and you wish to manually fail back to the previous node, you must use the failback command. The node must be in the feed's failover list, must not be the current primary for the feed, and must be more preferred than the current primary.

Before entering the failback command, you should:

  1. Identify the root cause of the problem.
  2. Return the node to good health.
  3. Restart the failed components to ensure the databases are synchronized and functioning normally.
  4. Implement safeguards to prevent the root cause from recurring.
  5. Issue the manual failback command to return the system to its initial operating mode.
Syntax `failback;`<feed>;`<node>
Optional N/A
Arming required Yes
Confirmation required No
q).maint.arm`ha
Now armed: `ha
q).maint.ha[`failback;`feed1;`A]
2025.09.04 14:48:39.809 DEBUG [mymaint] DISC Fetching key /config/haState
Feed 100 (feed1): primary=B nodes=`A`B failover=`A`B
2025.09.04 14:48:39.812 DEBUG [mymaint] DISC Fetching key /config/nodeState
Issuing failback of feed 100 from node B to A
2025.09.04 14:48:39.832 DEBUG [mymaint] DISC Received update for /config/haState
q)2025.09.04 14:48:39.832 INFO  [mymaint] HA New registry HA state:
2025.09.04 14:48:39.832 INFO  [mymaint] HA     feed node|  sndNode ts                          cc mode event
2025.09.04 14:48:39.832 INFO  [mymaint] HA     ----------|  ------------------------------------------------
2025.09.04 14:48:39.832 INFO  [mymaint] HA     200  A     |  A        2025.09.04D14:23:42.376031010 2  3     1
2025.09.04 14:48:39.832 INFO  [mymaint] HA     200  B     |  B        2025.09.04D14:23:42.755643715 2  4     1
2025.09.04 14:48:39.832 INFO  [mymaint] HA     100  B     |  A        2025.09.04D14:48:39.822580756 5  11    10
2025.09.04 14:48:39.832 INFO  [mymaint] HA     100  A     |  A        2025.09.04D14:48:39.822580756 5  12    10
2025.09.04 14:48:39.832 INFO  [mymaint] HA     0    B     |  B        2025.09.04D14:23:44.850117027 2  4     1
2025.09.04 14:48:39.832 INFO  [mymaint] HA     0    A     |  A        2025.09.04D14:23:44.814655581 2  3     1
2025.09.04 14:48:39.847 DEBUG [mymaint] DISC Received update for /config/haState
2025.09.04 14:48:39.847 INFO  [mymaint] HA New registry HA state:
2025.09.04 14:48:39.848 INFO  [mymaint] HA     feed node|  sndNode ts                          cc mode event
2025.09.04 14:48:39.848 INFO  [mymaint] HA     ----------|  ------------------------------------------------
2025.09.04 14:48:39.848 INFO  [mymaint] HA     200  A     |  A        2025.09.04D14:23:42.376031010 2  3     1
2025.09.04 14:48:39.848 INFO  [mymaint] HA     200  B     |  B        2025.09.04D14:23:42.755643715 2  4     1
2025.09.04 14:48:39.848 INFO  [mymaint] HA     100  B     |  A        2025.09.04D14:48:39.838373592 6  4     10
2025.09.04 14:48:39.848 INFO  [mymaint] HA     100  A     |  A        2025.09.04D14:48:39.838373592 6  3     10
2025.09.04 14:48:39.848 INFO  [mymaint] HA     0    B     |  B        2025.09.04D14:23:44.850117027 2  4     1
2025.09.04 14:48:39.848 INFO  [mymaint] HA     0    A     |  A        2025.09.04D14:23:44.814655581 2  3     1
.maint.ha[`sysstate;`feed1]
2025.09.04 14:48:46.324 DEBUG [mymaint] DISC Fetching key /config/haState
feed node|  feedName sndNode ts                          cc mode        event
---------|  -----------------------------------------------------------------
100  A    |  feed1     A       2025.09.04D14:48:39.838373592 6  PRIMARY     MANFBKFEED
100  B    |  feed1     A       2025.09.04D14:48:39.838373592 6  SECONDARY   MANFBKFEED

Mark a node as failed

The nodefail command allows you to manually mark a node as failed. Marking a node as failed will automatically trigger a failover from that node if there is a suitable target node.

Syntax `nodefail;`<node>
Optional N/A
Arming required Yes
Confirmation required No
q).maint.ha[`sysstate]
2025.09.04 14:50:42.841 DEBUG [mymaint] DISC Fetching key /config/haState
feed node|  feedName sndNode ts                          cc mode        event
---------|  -----------------------------------------------------------------
0    A    |  MD        A       2025.09.04D14:23:44.814655581 2  PRIMARY     HACTL
0    B    |  MD        B       2025.09.04D14:23:44.850117027 2  SECONDARY   HACTL
100  A    |  feed1     A       2025.09.04D14:48:39.838373592 6  PRIMARY     MANFBKFEED
100  B    |  feed1     A       2025.09.04D14:48:39.838373592 6  SECONDARY   MANFBKFEED
200  A    |  REPL      A       2025.09.04D14:23:42.376031010 2  PRIMARY     HACTL
200  B    |  REPL      B       2025.09.04D14:23:42.755643715 2  SECONDARY   HACTL
q).maint.arm`ha
Now armed: `ha
q).maint.ha[`nodefail;`A]
2025.09.04 14:50:58.213 DEBUG [mymaint] DISC Fetching key /config/nodeState
Setting node A to failed
2025.09.04 14:50:58.216 INFO  [mymaint] DISC Updating key /config/nodeState
q)2025.09.04 14:50:58.221 DEBUG [mymaint] DISC Received update for /config/nodeState
2025.09.04 14:50:58.221 INFO  [mymaint] SYS New registry node state:
2025.09.04 14:50:58.221 INFO  [mymaint] SYS     node| host                       user  baseport  ts
2025.09.04 14:50:58.221 INFO  [mymaint] SYS     ----|-----------------------------------------------------------
2025.09.04 14:50:58.221 INFO  [mymaint] SYS     A    | kxsa.firstderivatives.com  root  20000
2025.09.04 14:50:58.221 INFO  [mymaint] SYS     B    | kxsb.firstderivatives.com  root  20000
2025.09.04 14:50:58.221 INFO  [mymaint] SYS     C    | kxsc.firstderivatives.com  root  20000
2025.09.04 14:50:58.221 INFO  [mymaint] SYS Node state set, nodes=`A`B enabled=`B
2025.09.04 14:50:58.300 INFO  [mymaint] HA     100  A    |  B        2025.09.04D14:50:58...
2025.09.04 14:50:58.300 INFO  [mymaint] HA     0    B    |  B        2025.09.04D14:50:58...
2025.09.04 14:50:58.300 INFO  [mymaint] HA     0    A    |  B        2025.09.04D14:50:58...
q).maint.ha[`sysstate]
2025.09.04 14:51:02.021 DEBUG [mymaint] DISC Fetching key /config/haState
feed node|  feedName sndNode ts                          cc mode        event
---------|  -----------------------------------------------------------------
0    A    |  MD        B       2025.09.04D14:50:58.251989309 4  SECONDARY   MANFOVPROC
0    B    |  MD        B       2025.09.04D14:50:58.251989309 4  PRIMARY     MANFOVPROC
100  A    |  feed1     B       2025.09.04D14:50:58.263699680 8  SECONDARY   MANFOVPROC
100  B    |  feed1     B       2025.09.04D14:50:58.263699680 8  PRIMARY     MANFOVPROC
200  A    |  REPL      A       2025.09.04D14:50:58.274844004 4  SECONDARY   MANFOVPROC
200  B    |  REPL      A       2025.09.04D14:50:58.274844004 4  PRIMARY     MANFOVPROC

Mark a node as healed

The nodeheal command allows you to manually mark a node as healed. It reverses the previous action of marking a node as failed.

A healed node becomes a candidate for normal failovers and failbacks. It is typically used when a node fails but parts of the cluster remain operational. Examples include a network interface failure on a node, a hardware failure on a node detected by the HwMon process, DBW failure on a node or RS failure on a node.

Syntax `nodeheal;`<node>
Optional N/A
Arming required Yes
Confirmation required No
q).maint.arm`ha
Now armed: `ha
q).maint.ha[`nodeheal;`A]
2025.09.04 14:54:39.044 DEBUG [mymaint] DISC Fetching key /config/nodeState
Setting node A to enabled
2025.09.04 14:54:39.047 INFO  [mymaint] DISC Updating key /config/nodeState
q)2025.09.04 14:54:39.054 DEBUG [mymaint] DISC Received update for /config/nodeState
2025.09.04 14:54:39.054 INFO  [mymaint] SYS New registry node state:
2025.09.04 14:54:39.055 INFO  [mymaint] SYS     node| host                       user  baseport
2025.09.04 14:54:39.055 INFO  [mymaint] SYS     ----|-----------------------------------------------------------
2025.09.04 14:54:39.055 INFO  [mymaint] SYS     A    | kxsa.firstderivatives.com  root  20000
2025.09.04 14:54:39.055 INFO  [mymaint] SYS     B    | kxsb.firstderivatives.com  root  20000
2025.09.04 14:54:39.055 INFO  [mymaint] SYS     C    | kxsc.firstderivatives.com  root  20000
2025.09.04 14:54:39.055 INFO  [mymaint] SYS Node state set, nodes=`A`B enabled=`A`B

Force a mode change to a feed

In certain rare cases, file corruption of a feed's mode may make it necessary to manually force a mode change on a feed. For example, feed1 is primary on both nodes A and B and you want to force one of the feeds to be secondary. Or feed1 is secondary on both nodes and you want to force one of the feeds to be primary.

Syntax `forcesysstate;`<feed>;`<node>;.enumfeedMode.mode
Optional N/A
Arming required Yes
Confirmation required Yes
q).maint.ha[`sysstate]
2025.09.04 15:00:50.867 DEBUG [mymaint] DISC Fetching key /config/haState
feed node|  feedName sndNode ts                          cc mode        event
---------|  -----------------------------------------------------------------
0    A    |  MD        B       2025.09.04D14:50:58.251989309 4  SECONDARY   MANFOVPROC
0    B    |  MD        B       2025.09.04D14:50:58.251989309 4  PRIMARY     MANFOVPROC
100  A    |  feed1     B       2025.09.04D14:50:58.263699680 9  SECONDARY   MANFOVPROC
100  B    |  feed1     B       2025.09.04D14:50:58.263699680 9  SECONDARY   MANFOVPROC
200  A    |  REPL      A       2025.09.04D14:50:58.274844004 4  SECONDARY   MANFOVPROC
200  B    |  REPL      A       2025.09.04D14:50:58.274844004 4  PRIMARY     MANFOVPROC
q).maint.arm`ha
Now armed: `ha
q).maint.ha[`forcesysstate;`feed1;`B;.enum.feedMode.PRIMARY]
2025.09.04 15:01:17.755 DEBUG [mymaint] DISC Fetching key /config/haState
Forcing state:
feed node|  feedName sndNode ts                          cc mode        event
---------|  -----------------------------------------------------------------
100  A    |  feed1     B       2025.09.04D14:50:58.263699680 10 SECONDARY   MANFOVPROC
100  B    |  feed1     B       2025.09.04D14:50:58.263699680 10 PRIMARY     MANFOVPROC
Continue? y
2025.09.04 15:01:19.210 DEBUG [mymaint] DISC Fetching key /config/haState
2025.09.04 15:01:19.217 INFO  [mymaint] DISC Updating key /config/haState
State updated
q)2025.09.04 15:01:19.225 DEBUG [mymaint] DISC Received update for /config/haState
2025.09.04 15:01:19.225 INFO  [mymaint] HA New registry HA state:
2025.09.04 15:01:19.225 INFO  [mymaint] HA     feed node|  sndNode ts
2025.09.04 15:01:19.225 INFO  [mymaint] HA     ----------|  ---------------------------
2025.09.04 15:01:19.225 INFO  [mymaint] HA     200  A     |  A        2025.09.04D14:50:58.274844004
2025.09.04 15:01:19.225 INFO  [mymaint] HA     200  B     |  A        2025.09.04D14:50:58.274844004
2025.09.04 15:01:19.225 INFO  [mymaint] HA     100  B     |  B        2025.09.04D14:50:58.263699680
2025.09.04 15:01:19.225 INFO  [mymaint] HA     100  A     |  B        2025.09.04D14:50:58.263699680
2025.09.04 15:01:19.225 INFO  [mymaint] HA     0    B     |  B        2025.09.04D14:50:58.251989309
2025.09.04 15:01:19.225 INFO  [mymaint] HA     0    A     |  B        2025.09.04D14:50:58.251989309
2025.09.04 15:01:19.382 INFO  [mymaint] DISC Discovery changes, scs=`SDL deleted=`kxsSDL_B1
2025.09.04 15:01:21.229 INFO  [mymaint] DISC Discovery changes, scs=`SDL added=`kxsSDL_B1
.maint.ha[`sysstate]
2025.09.04 15:02:07.580 DEBUG [mymaint] DISC Fetching key /config/haState
feed node|  feedName sndNode ts                          cc mode        event
---------|  -----------------------------------------------------------------
0    A    |  MD        B       2025.09.04D14:50:58.251989309 4  SECONDARY   MANFOVPROC
0    B    |  MD        B       2025.09.04D14:50:58.251989309 4  PRIMARY     MANFOVPROC
100  A    |  feed1     B       2025.09.04D14:50:58.263699680 10 SECONDARY   MANFOVPROC
100  B    |  feed1     B       2025.09.04D14:50:58.263699680 10 PRIMARY     MANFOVPROC
200  A    |  REPL      A       2025.09.04D14:50:58.274844004 4  SECONDARY   MANFOVPROC
200  B    |  REPL      A       2025.09.04D14:50:58.274844004 4  PRIMARY     MANFOVPROC

Clear your system state entries

The clearsysstate command clears all system state entries for the specified feed or feeds. This action affects the active state of all running processes that participate in HA activity and should be performed only on a quiesced system.

Danger

This command is for emergency use only and should never be executed on an active system.

Syntax `clearsysstate;`<feeds>
Optional N/A
Arming required Yes
Confirmation required Yes
q).maint.ha[`sysstate]
2025.09.04 15:11:37.211 DEBUG [mymaint] DISC Fetching key /config/haState
feed node|  feedName sndNode ts                          cc mode        event
---------|  -----------------------------------------------------------------
0    A    |  MD        B       2025.09.04D14:50:58.251989309 4  SECONDARY   MANFOVPROC
0    B    |  MD        B       2025.09.04D14:50:58.251989309 4  PRIMARY     MANFOVPROC
100  A    |  feed1     B       2025.09.04D14:50:58.263699680 10 SECONDARY   MANFOVPROC
100  B    |  feed1     B       2025.09.04D14:50:58.263699680 10 PRIMARY     MANFOVPROC
200  A    |  REPL      A       2025.09.04D14:50:58.274844004 4  SECONDARY   MANFOVPROC
200  B    |  REPL      A       2025.09.04D14:50:58.274844004 4  PRIMARY     MANFOVPROC
q).maint.arm`ha
Now armed: `ha
q).maint.ha[`clearsysstate]
*** Warning: active processes will be affected ***
Removing:
2025.09.04 15:12:11.959 DEBUG [mymaint] DISC Fetching key /config/haState
feed node|  feedName sndNode ts                          cc mode        event
---------|  -----------------------------------------------------------------
0    A    |  MD        B       2025.09.04D14:50:58.251989309 4  SECONDARY   MANFOVPROC
0    B    |  MD        B       2025.09.04D14:50:58.251989309 4  PRIMARY     MANFOVPROC
100  A    |  feed1     B       2025.09.04D14:50:58.263699680 10 SECONDARY   MANFOVPROC
100  B    |  feed1     B       2025.09.04D14:50:58.263699680 10 PRIMARY     MANFOVPROC
200  A    |  REPL      A       2025.09.04D14:50:58.274844004 4  SECONDARY   MANFOVPROC
200  B    |  REPL      A       2025.09.04D14:50:58.274844004 4  PRIMARY     MANFOVPROC
Continue? y
2025.09.04 15:12:21.247 INFO  [mymaint] DISC Updating key /config/haState
State cleared
q)2025.09.04 15:12:21.251 DEBUG [mymaint] DISC Received update for /config/haState
2025.09.04 15:12:21.252 INFO  [mymaint] HA New registry HA state:
2025.09.04 15:12:21.252 INFO  [mymaint] HA     feed node|  sndNode ts  cc mode event
2025.09.04 15:12:21.252 INFO  [mymaint] HA     ----------|  ----------------------
.maint.ha[`sysstate]
2025.09.04 15:12:45.212 DEBUG [mymaint] DISC Fetching key /config/haState
feed node|  feedName sndNode ts  cc mode event
---------|  -----------------------------------
q)

Stop/restart your entire cluster

When stopping/restarting a cluster, never start a single node by itself with the intention of forcing it to become primary unless you are certain that node was primary for all feeds and no other nodes took over since stopping the cluster. Note, the minimum number of nodes to reach a quorum would be required.

  1. Start all available nodes at the same time and allow them to reach consensus during the initialization phase. They will be able to decide which one should be primary for each of the feeds based on previous states.
  2. If in doubt, contact KXS support for advice before taking further action. It is critical to avoid faulty application states that arise due to improper operation.

Troubleshooting

Look up your failover events

You can find out why a particular failover event took place by reviewing the logs for any data loader (SDL, MDL, etc). It shows the results of the chkFeed function. This internal function is run at the frequency that you define in the haChkFreq column in systemParams.yaml and generates a trace message published by a data loader.

Line A

2023.08.11 05:47:22.512 DEBUG [kxsSDL_D2] HA chkFeed: id=110 nodes=`C`D phi=NY lts=YY wts=YY cur=YY rsp=YY ems=YY ns=YY qr=NY

Line B

2023.08.11 05:47:22.513 WARN [kxsSDL_D2] HA Fail-over of feed feed10 from node C:
phi=251.301 0.000 hts=2023.08.11D05:47:14.054320000 2023.08.11D05:47:22.353276000
pts=2023.08.11D05:47:12.053753000 2023.08.11D05:47:13.942399000
lts=2023.08.11D05:47:14.050757000 2023.08.11D05:47:22.408395000
wts=2023.08.11D05:47:14.002252000 2023.08.11D05:47:15.906492000
ets=2023.08.11D05:47:12.017318771 2023.08.11D05:47:22.017470333
flts=2023.08.11D05:47:14.644530000 age=8.45868 0.159724 nok=`A`B`C`D qr=NY

2023.08.11 05:47:22.513 TRACE [kxsSDL_D2] HA phi: 2.000547 2.000008 2.000683 2.000005 2.000215 2.000427
2.000414 2.000398 2.000578 2.000871 2.000306 2.000957 2.002111 2.00001 2.000292
2.000274 2.000278 2.002856
2.002246 2.000591 2.000931 2.000209 2.000112 2.000466 2.000455 2.000219 2.000053
2.000438 2.000603 2.000133
2.000669 2.001256 2.00158 2.001809 2.000462 2.000587 2.000134 2.000646 2.000149
2.000001 2.000754 2.000163

2023.08.11 05:47:22.513 WARN [kxsSDL_D2] HA Triggering AUTOFOVFEED of feed feed10 from node C

Line A: Reports the outcome of the evaluation when something anomalous is detected. The message gives the feed ID, node list, and the heuristic outcomes for each node and each of the key HA indicators described earlier. Nodes are listed starting with the primary, through nodes in decreasing preference according to the feed's configuration, up to and ending with the local node.

In this example, nodes C and D support feed 110 and C is primary; the node sequence viewed on D is CD. A Y indicates that the condition for the heuristic on the associated node is nominal. The message is not displayed unless at least one anomalous condition is detected; however, the emission of the line does not necessarily imply that a mode change event will occur.

Line B: Reports that a mode change event from the current primary (node C) to this node has been initiated. Most properties displayed on this line represent a pair of values referring to the primary node and to this node, respectively. The interpretation of the values is described in more detail below.

Property Description
phi Shows the phi calculation based on the heartbeat interarrival times for the primary node and for this node. Higher values suggest an increased probability that the associated node is encountering difficulty.
hts Shows the timestamps of the last heartbeat message received by this node from the primary, and the last heartbeat message sent by this node to peer HA clients. This property influences the phi calculation.
pts Shows the timestamps of the penultimate heartbeat message received by this node from the primary, and the penultimate heartbeat message sent by this node to peer HA clients. This property influences the phi calculation.
lts Shows the timestamps of the last raw message reported to have been received by the primary, and the last raw message received by this node.
wts Shows the timestamps of the last formatted message reported to have been published by the primary, and the last formatted message received by this node
ets Shows the timestamps of the last successful EMS check reported by the primary, and the last successful EMS check on this node
flts Shows the time base used to assess feed delinquency. This value is updated on the transition of incoming feed data to nominal, at which time it is set to the latest incoming message time reported by any node. This property influences the lts calculation.
age Shows the age of the feed data received by this node from the primary, and the age of the feed data sent by this node to peer HA clients. This property influences the lts and wts calculations, and is reflective of the cur and rsp properties shown on the preceding output line.
nok Shows the global list of nodes that are enabled and not in a hardware error state. A node that is not running may show in this list unless administrator action or HWMON has flagged it. A node that does not share common feed interest with this process may also show in the list.
qr Shows whether nodes are reporting having a server registry quorum present. This is the same determination as shown on the preceding output line.

The chkFeed trace message is only generated if one of the indicators = N. If all indicators = Y, there is no trace message and your feed and node are in a healthy state.

TRACE [kxsSDL_B1] HA chkFeed: id=50 nodes=`A`B phi=NY lts=YN wts=YY cur=YY rsp=YY ems=YY ns=YY qr=YY

The order of the nodes (AB) indicates the order of the indicators (Y or N). In the above example, you are reporting on nodes A and B and the lts indicator for node B is N, which means that the last raw message received by node B exceeded the configured threshold of what you consider to be lagging.

Check your log files

The log files of importance for countDB mismatch and unintended failover errors are generated on both nodes:

  • SDL (a pair that share the same feed where the mismatch was seen)
  • DBW

You can retrieve and view the log file from the command line by using kxsctl log. You name the process, and kxsctl lists the logs available for it so you can select one and page through it. You can view the logs of a process on any node, so you don't have to run the command on each node involved in the issue.

[user@kxsa logs]# kxsctl log kxsSDL_A1

For the full set of flags, see Look up log files.

Check your configuration settings

There are numerous configurations that need to be considered when deploying an HA Sensors system. Efforts have been made to set reasonable defaults in our KX Sensors packages. However, additional tuning and configuration changes may be required.

  • The Performance Tuning Guide contains several kernel settings that are important for systems requiring low failover times.
  • Check your disk buffer settings.
  • Check your router settings.
  • Make sure that your server times are in sync and review your ntp settings.
  • Review whether feeds are correctly set to be periodic/aperiodic. This optimizes failover times for periodic.

Next steps