Skip to content

Configuration objects

This page describes how KX Sensors configuration objects are stored as YAML files, and details the structure of the core configuration objects.

All configuration objects, tables, and APIs are stored in YAML files that you can view and edit using any text editor. Changes to a YAML file require a dynamic upgrade to take effect. Refer to the metadata files for details about the expected structure, field data types, and whether or not a field is required.

For example, the following metadata file defines the process configuration object:

process:
  m-description: 'This configuration parameter stores the process wide level configurations, including how to intialize and receive data, what to publish and subscribe to, how to react to EOX events, and where to store in memory tables.

 In any package the top row can be left blank in the Process column to define new global defaults at the package level'
  values:
    m-type: list
    m-required: true
    m-keys: [ process, svcClass ]
    m-value:
      process:
        m-type: symbol
        m-description: Process name

      parent:
        m-type: symbol
        m-description: Row above to use as base configuration

      svcClass:
        m-type: symbol
        m-description: Service class of the process

      node:
        m-type: symbol
        m-description: Node of the process in HA system

      analytics:
        m-type: symbols
        m-description: Analytics/Analytic groups to load into process #! Remove field when conversion of most client packages is complete

      libraries:
        m-type: symbols
        m-description: Instructions to load into process, in order
        m-overlay: merge
...

The meta/ directory of kxs-core holds the metadata file for every configuration object in the system:

$ ls meta/
alert.yaml            process.yaml
alertEmail.yaml       protectedWords.yaml
api.yaml              repl.yaml
dbSum.yaml            rLogPubSub.yaml
dbSumParams.yaml      sapi.yaml
enum.yaml             scaleDefaults.yaml
estDefaults.yaml      schema.yaml
feed.yaml             sdl.yaml
flags.yaml            streams.yaml
hdbTiers.yaml         svcClassLimits.yaml
hosts.yaml            systemParams.yaml
javaEmail.yaml        tableDefaults.yaml
logLevels.yaml        tblGen.yaml
logWatch.yaml         uomConversion.yaml
masterNodes.yaml      valDefaults.yaml
mdbTables.yaml        veeDefaults.yaml
mon.yaml              volumes.yaml
msProxy.yaml          watch.yaml
nodeConfig.yaml
nodes.yaml
nodeState.yaml
opsReportEmailList.yaml
pen.yaml

Validate changes

To test the syntax of your edited YAML object, run your YAML code through a YAML validator. To test the logic of your YAML object, restart a KXS process that uses the configuration object or table.

Metadata, config, and override files

There is a minimum of two different YAML files for each configuration object: a metadata file and a core config file. If required, there may be one or more package override files; for example, a utilities config object or IoT config object.

The metadata file defines each column in the configuration object; for each column the datatype (short, string, symbol, etc.) is defined as well as an optional description. Unlike configuration values, which can be overridden in overlay packages, metadata files for objects defined in KX packages such as KXS Core should never be modified.

For example, here's the metadata for the feed configuration object, with no overlay type specified:

feed:
  m-description: Specifies data feeds being processed by the system, including sufficient details for automatic failover in high-availability (HA) systems
  values:
    m-type: dict
    m-required: true
    m-value:
      id:
        m-type: short
        m-description: Numeric identifier, unique within system
        m-required: true

      isCoherent:
        m-type: boolean
        m-description: Indicates if feed is coherent (i.e. can be safely processed simultaneously by separate nodes)

      isMediated:
        m-type: boolean
        m-description: Indicates if feed is actively mediated by non-coherent secondary nodes

      isActiveSec:
        m-type: boolean
        m-description: Indicates if secondary is expected to receive feed

      compFn:
        m-type: symbol
        m-description: Comparison function used to compare relative position within the stream of two fingerprints. See `.ha.comp` for details
        m-required: true

The core config object contains the default values for each column in the metadata file. Core config files should never be modified; if a default value is not acceptable, create an override with the desired value(s). For example, here's the core config file for feed.yaml:

feed:
  values:
    <default>:
      isCoherent: false
      isMediated: true
      isActiveSec: true
      compFn: .ha.comp
      nodes: <all>
      skewThr: 10
      phiThr: 2.5
      phiMeanOfs: 2200
      phiMinDev: 100
      desc: Default entry

    MD:
      id: 0
      sc: mdl
      streamIn: mutreqs
      topicIn: MD
      streamsOut: emsint
      topicsOut: MD
      desc: Master data feed (API calls to MDL)

The utilities 1-node config object contains overrides for the utilities package. For example, a utilities-layer override for feed.yaml might look like this:

feed:
  values:
    feed1:
      id: 100
      sc: sdl
      nodes: [ A ]
      procs: [ kxsSDL_A1 ]
      streamIn: feed1
      topicIn: feed1
      streamsOut: emsint
      topicsOut: [ preSD1 ]
      dbIgnoreTopics: [ preSD1 ]
    AGG:
      id: 200
      sc: agg
      nodes: [ A ]
      procs: [ kxsAGG_A1 ]
      streamIn: emsint
      topicIn: SD1
      streamsOut: emsint
      topicsOut: [ AGG1 ]
    MD:
      nodes: [ A ]
      procs: [ kxsMDL_A ]
    CD:
      nodes: [ A ]
      procs: [ kxsCDL_A ]
      id: 52
      sc: cdl
      streamIn: mutreqs
      streamsOut: emsint
      topicIn: CD
      topicsOut: CD
      desc: Calculation data feed (API calls to CDL)

process.yaml

In this configuration object, you configure your default process parameters including how to initialize and receive data, how to react to EOX events and where to store in-memory tables. You define your process parameters at the service class level; all processes assigned that service class will inherit its parameters.

For example, suppose you have three SDLs running on node A. The service class is called SDL while the service class instances or processes will be called kxsSDL_A1, kxsSDL_A2 and kxsSDL_A3. You define the prefix and numbering for your service class instances in systemParams.yaml.

process:
  values:
   - process: <default>
     load: Version
     mdbNS: .
     mountNS: .mdbtmp
     destroyMountNS: true
     rcv: .mdb.rcv

   - svcClass: DBW
     libraries: [ src/proc/storage/dbw/dbw.q ]
     init: .dbw.init
     sub: "func:.dbw.subs"
     load: [ Version, .dbw.loads ]
     eoia: .dbw.eoia
     eopa: .dbw.eopa
     eoda: .dbw.eoda
     lrz: .dbw.lrz
Attribute Description Example
process If you specify a specific process, the parameters will apply to that process only. If you leave this column blank, the parameters will apply to all processes assigned the designated service class. initNoder
parent This column for grouping similar processes has been largely replaced by the svcClass column.
svcClass Service class of the process. rdb
node Node that the process will run on or blank for all nodes.
libraries q code files to load into process in the sequence specified. src/sapi/sapimon.q
init Initialization functions to execute in the sequence specified. Must be a niladic function. .rdb.init
pub EMS topics and tables to publish to on the emsint stream.
Warning! Topics beginning with "_" are for internal KXS use and should never be used in extensions to the core packages of Sensors.
MdUpdate
sub EMS topics and tables to subscribe to on the emsint stream.
Warning! Topics beginning with "_" are for internal KXS use and should never be used in extensions to the core packages of Sensors.
[ MdUpdate, QueryLog, MD, SD ]
load List of non-subscription tables to keep in memory. Version
tables Tables or table groups of interest from EMS subscriptions. MDB handlers will be registered for these tables. [ SD, ImportLog ]
mdbNS Namespace in which to store in-memory tables. .mdb
mountNS Where to mount any on-disk database for initialization. .mdbrmp
destroyMountNS Boolean denoting whether or not to keep database mounted. true
rcv Defines the update function to be executed when receiving table data. .mdb.applyRcv
eoia Functions to run at the end of your end-of-interval signal from EMS.
eoiz Functions to run at the end of your end-of-interval signal from DBW.
eopa Functions to run at the start of your EOP signal. eoia
eopz Functions to run at end of your EOP finish signal from DBW.
eoda Functions to run at the end of your EOD start signal from EMS.
eodz Functions to run at the end of your EOD finish signal from DBW.
lra Functions to run prior to process catching up to the live point in an EMS stream; that is, the process is replaying a log file.
lrz Functions to run once a process has caught up to the live point in an EMS stream; that is, the process is receiving new data. .dbw.lrx

mdbTables.yaml

In this configuration object, you link your processes to your in-memory tables and specify which functions to run during and after the EOD, EOI and EOP signals from EMS.

mdbTables:
  values:
   - process: <default>
     table: <default>
     isvp: false

    #
    # Set splitType default as NONE for all tables.
    #
   - table: <any>
     splitType: NONE

    #
    # Defaults by table category.
    #
   - table: basic
     init: [ loadSnapshot, applySchemaAttrs ]
     rcv: upd

   - table: partitionedDeltaMem
     init: [ loadEmpty, unenum, applySchemaAttrs ]
     rcv: upd

   - table: splayed
     init: [ loadSnapshot, unenum, applySchemaAttrs ]
     rcv: upd

   - table: partitionedDelta
     init: [ loadEmpty, unenum, applySchemaAttrs ]
     rcv: upd

   - table: memOnly
     init: [ loadEmpty, applySchemaAttrs ]
     rcv: upd

   - table: memOnlyRepl
     init: [ loadEmpty, applySchemaAttrs ]
...

Note

For changes to take effect, a restart of the affected process is required.

Attribute Description
process Process name. Can be blank for all processes, the name of a specific process, or a parent process group.
table Table name. Can be blank for all tables, a schema group, a table category, or a specific table name.
isvp Whether or not the table should be virtually partitioned in the process.
colnames Column names to maintain.
init How the process initializes this table on startup.
rcv Receive function (real-time and log replay).
splitType May be set to NONE (default), MEMORYMAP, or MEMORY. See Split master data tables for more details.

The process field may contain one of the following values (from low to high precedence):

Value Example Description
literal <default> Entry applies to any process that is not otherwise specified by service class or name.
literal <any> Entry applies to any process that is not otherwise specified by service class or name.
service class RDB Entry applies to any process of the service class that is not otherwise specified by name.
process name RDB_A1 Entry applies to the specified process.
literal <all> Entry applies to all processes.

Likewise, the table field may contain one of the following values (from low to high precedence):

Value Example Description
literal <default> Entry applies to any table that is not otherwise specified by category, group, function, or name.
literal <any> Entry applies to any table that is not otherwise specified by category, group, function, or name.
table category splayed Entry applies to tables of the category that is not otherwise specified by group, function, or name.
table group MD Entry applies to tables of the group that is not otherwise specified by function or name.
table name ns.getTbls Entry applies the tables whose names are returned by the given niladic function.
literal <all> Entry applies to all tables.

Split master data tables

If a master data table is very large, it may be neither feasible nor desirable to load the entire table in memory. In such cases, KXS offers the following splitType options for use in mdbTables.yaml:

Value Description
NONE (default) Entire table is loaded into memory. This offers the best query performance.
MEMORYMAP Split table with the base table memory-mapped from the on-disk snapshot and the delta table in memory. The delta table retains only recent updates. This configuration offers faster in-memory updates as they are applied to the delta table which is generally smaller than the base table. It offloads the rest to disk, reducing the memory footprint at the cost of query performance.
MEMORY Split table into a base and delta table, with both maintained in memory. The main benefit is faster in-memory updates as they are applied to the delta table which is generally smaller than the base table. It does impact query performance but not as much as the MEMORYMAP option. It does not reduce the memory footprint.

Note

If you change your memory split configuration, the processes affected will require a restart. If multiple process instances are available, this can be done in a rolling manner.

Diagram comparing the NONE, MEMORYMAP, and MEMORY splitType configurations across physical memory, virtual memory, and disk for a master data table Diagram comparing the NONE, MEMORYMAP, and MEMORY splitType configurations across physical memory, virtual memory, and disk for a master data table

Configuration

To enable a splitType for a table, set it in mdbTables.yaml. For example, if all processes should use the MEMORYMAP option for the Sensor table, add the following:

#
# Table customizations.
#
mdbTables:
  values:
   - table: Sensor
     splitType: MEMORYMAP

On the other hand, if only the RDB service class should use the MEMORYMAP option for the Sensor table, add the following:

#
# Service class customizations.
#
mdbTables:
  values:
   - process: RDB
     table: Sensor
     splitType: MEMORYMAP

Use of functions in the .mdb namespace

If APIs make use of the .mdb functions, the detail of memory mapping is abstracted away from the user. All the .mdb functions can be used seamlessly between the different configurations. For example, an API that calls .mdb.sel can simply continue to do so; the framework takes care of resolving this function to the appropriate variant depending on the configuration.

Choose splitType configurations

To decide which tables and service classes should use the MEMORYMAP option, consider:

  • Table size: tables can be sorted by size to identify candidates for memory mapping to yield the largest memory footprint reductions.
  • Use in queries: memory mapping tables that are heavily used in query joins has a larger adverse effect on query performance. The impact is proportional to the size of the tables.
  • Query load by service class: different service classes perform different queries, so the sets of tables to be memory mapped may vary across service classes.
  • Service class instance count: memory footprint reductions are realized by each process instance, so multiply accordingly to project memory footprint reductions.

uomConversion.yaml

This configuration object contains your seed unit of measure conversions (feet to inches, inches to centimeters, gallons to litres, kilograms to pounds, etc.).

uomConversion:
  values:
   - from: THERM
     to: BTU
     scaling: 99976.1

   - from: GAL
     to: L
     scaling: 3.78541

   - from: GPM
     to: GPH
     scaling: 60

   - from: LPM
     to: LPH
     scaling: 60

   - from: GPM
     to: LPM
     scaling: 3.78541

   - from: FT
     to: IN
     scaling: 12

   - from: FT
     to: CM
     scaling: 30.48
...

Warning

This configuration object is static rather than dynamic. Any changes to it require a complete system shutdown and restart.

Attribute Definition Example
from The source unit of measure before conversion. FT (feet)
to The target unit of measure after conversion. IN (inches)
scaling The value by which to scale the source unit of measure. 12

pen.yaml

In this configuration object, you define your “penning” (suspending) rules, if any, for each service class. Penned data is sensor data from SDL and VEE awaiting a certain event such as completion of a request or arrival of other data and cannot be processed immediately. For example, VEE receives a set of readings with an unknown channel and “pens” or defers their processing until MDL creates the missing channel.

If you do not define overrides for individual service classes in this configuration object, the system defaults will apply to all service classes.

pen:
  values:
   - proc: <default>
     rlfreq: 15
     stto: 30
     arqfreq: 2
     arqth: 2
     isstr: false
     strdef: 0
     strmax: 0

   - proc: VEE
     stto: 60
     isstr: true
     strdef: 15000
     strmax: 10

   - proc: PROF
     rlfreq: 5

Note

If you change a record in this configuration object, you must restart the relevant process (VEE, SDL, etc.).

Attribute Definition Example
proc The name of the service class. <default>
rlfreq The frequency in minutes that the recovery log snapshot is generated. 15
stto Short-term request timeout in seconds for GW operations. 30
arqfreq The frequency in minutes to check for aged requests. 60
arqth The age in minutes of a request at which point it is considered “aged” and an alert is generated in the client application. 1440
isstr Flag indicating if short-term requests should be retried if the request times out or fails to reach a destination target. N
strdef Retry deferral in milliseconds for short-term requests that result in NO_DEST (timeout requests will be retried immediately). 0
strmax Maximum number of retries for short-term requests that result in TIMEOUT or NO_DEST. 0

protectedWords.yaml

In this configuration object, you define your programming language specific aliases for enumerations where the default KXS enumeration conflicts with protected words in that language.

For example, the enumeration IN within KXS refers to the inch unit of measure. However, in a language such as plsql, the word IN is a protected word and therefore causes compilation issues. Hence, for the plsql-specific exports of KXS enumerations, we use the INCH alias instead.

protectedWords:
  values:
   - language: plsql
     word: IN
     replacement: INCH

   - language: cpp
     word: KG
     replacement: KGR

   - language: cpp
     word: ERANGE
     replacement: ERANGECHK

   - language: c
     word: KG
     replacement: KGR

   - language: c
     word: ERANGE
     replacement: ERANGECHK
Attribute Definition Example
language The programming language. plsql
word The word to be protected. IN
replacement The replacement word or alias. INCH

sdl.yaml

In this configuration object, you define your custom ingestion functions and micro-batching rules for application-specific message types.

sdl:
  values:
    MD_REISSUE:
      isBatch: false
      ingestFn: .sdl.ingestMdRepub
      batchLenThr: 0
      maxHoldTm: 0D00:00:00.000000000
      maxInactTm: 0D00:00:00.000000000
      expTm: 0D00:00:00.000000000
Attribute Definition Example
isBatch Flag indicating whether or not the message type will be micro-batched. true, false
splitFn Dyadic function that is responsible for dividing the incoming message into one or more segments that will be microbatched separately in each batch. .sdl.splitTable
preIngestFn Dyadic pre-ingestion processing function for the message type. .sdl.valData
ingestFn Function used to perform ingestion of incoming sensor data, typically transforming the data into one or more database tables (including ImportLog and ImportException if required) to be forwarded to the next step in the ingestion pipeline (e.g. VEE). .sdl.transform / .example.ingFn
batchLenThr Maximum batch length threshold to release a batch. You define this application-specific length (number of records, number of bytes, etc.) in your splitting function. If you enter zero as your batch length threshold, the decision to release a batch will be based on time rather than batch size. 1000
maxHoldTm Maximum age threshold for a batch. Batches older than the specified age will be ingested. 0D00:00:10
maxInactTm Optional. Maximum inactivity threshold for a batch. Batches that have not received a new segment for longer than the specified threshold will be ingested. 0D00:00:10
expTm Expiration threshold for a batch in the event that a publishing client fails to terminate a batch. You should set this parameter to a value larger than the expected life of a batch. 4D

tblGen.yaml

In this configuration object, you configure the functions that can be used to generate initial seed versions of database tables created by the DBW process during the initial start-up of an environment.

Each entry in the configuration specifies the name of the function to be run and its associated q file. Note that multiple entries for the same table are supported and each entry will be executed. The execution sequence for each table is determined by the repository-based sequence for the configuration overrides and the row order within each.

tblGen:
  values:
   - table: TimeZone
     file: tblGen
     function: genTZ

   - table: User
     file: tblGen
     function: genUser

   - table: Organization
     file: tblGen
     function: genOrg

   - table: Watermark
     file: tblGen
     function: genWatermark

   - table: CalcWatermark
     file: tblGen
     function: genCalcWatermark

Note

Changes to this configuration object only take effect when you are performing an initial install of KX Sensors or you are rebuilding your database.

Attribute Definition Example
table The name of the table to be created. TimeZone
file The q file containing the function. TblGen
function The name of the function to be run. Functions specified in the parameter should take three parameters:
ts: timestamp of table creation, of type timestamp
tn: table name, of type symbol
t: table data, of type table (initially empty, but may have had rows added by previous invocations)
These functions should return an augmented version of t (with rows added, modified, or removed, but honoring the configured schema of the table).
genTZ

rLogPubSub.yaml

In this configuration object, you specify the publishers and subscribers of recovery logs as well as the directory path of the write-down folder and the rollover frequency.

rLogPubSub:
  values:
    <default>:
      dir: "${OPDATA}/rlog" # See opDataDir in systemParams.yaml
      lfreq: 60
      rfreq: 60
      mincut: 5

Note

For changes to take effect, a restart of the affected process is required.

Attribute Definition Example
name Name of the folder that the process writes recovery logs to. SDL
topic Name of the table for the process to listen for. ReadingVal
dir Directory path to write-down folder. /data/kxs/rlog
lfreq The log rollover frequency in minutes. 60
rfreq The subscriber reporting frequency in seconds. 60
mincut The minimum number of messages in a log file before it is cut. 5
pub
subs Comma-separated list of service classes expected to subscribe to the recovery log.
rcv Name of the dyadic application-defined function to receive control on receipt of new data.
rcvm Name of the dyadic application-defined function to receive control on receipt of new batched data.

The sample values in the table above indicate that the subscriber is listening for the topic ReadingVal within the SDL folder of the directory and is checking every 60 seconds to see whether it has caught up with everything published here by the publishing process. This log is rolled over to a new file every 60 minutes, and SDL doesn't write a file to it until it has at least five messages. On receiving this topic (table), it executes the function defined in rcv (.vee.recv).

Next steps