Skip to content

Network conditioning and fast failure detection

This page explains how KX Sensors adjusts TCP socket properties to improve network failure detection speed and support high-traffic connections.

KX Sensors adjusts TCP socket properties for two main purposes:

  • Improve detection of network failures. Especially in high-availability deployments, rapid detection of network failures is important to ensure that fail-over and recovery procedures are initiated as quickly as possible when a fault occurs. The default TCP settings that influence detection of communication failures (keep-alive and timeout settings) are designed for slow failure rather than fast failure.
  • Support high-traffic connections. Certain communication channels in KX Sensors (notably those involved in the ingestion pipeline) need to support very high volumes of data. The default TCP buffer sizes that influence bandwidth management are designed for much lower volumes of data than KX Sensors is expected to accommodate.

KX Sensors v2 included a library to provide this "network conditioning" functionality: netcon. Sensors v3 no longer requires netcon. TCP socket options are now set by the core code, both on the server and client side, rather than an additional library.

For more information, see connOpenTO, connRetryInt, connRTT (equivalent to the dist value in netcon config), and connRate (equivalent to the rate value in netcon config) in System parameters (systemParams.yaml).

On the client side, see the SAPI documentation for configuration parameters pertaining to socket conditioning.

To set connRTT:

  1. Run the following command:

    # This command executes ping every 0.2 seconds for 60 seconds.
    ping -w 60 -i 0.2 <ip address>
    
  2. Take the maximum value returned by the command and multiply it by two. If that value is higher than 500, use it as the connRTT value. If it is lower than 500, use 500.

Complete this exercise between all nodes of the cluster, and between all nodes hosting external clients and each node of the cluster, for all networks (in cases where nodes are configured to connect via multiple networks). Use the highest value from the results.

Next steps