MSSQL in Windows Server Failover Cluster, can't find instances

CMK version: CEE 2.3.0p43
OS version: Ubuntu 20.04 LTS

We are trying to configure the new MSSQL Plugin ( Monitoring Microsoft SQL Server ) to monitor MSSQL Server Instances in a Failover Cluster.
Our problem is, no instances are found by the plugin. As seen in the logs the agent tries to find instances only on localhost, but the instances are running on a dedicated cluster node IPs.

What we are trying to achieve is to use one rule for all cluster hosts with the option “Try to detect MS SQL instances present on server”.

The deprecated plugin mssql.vbs solves this issue by querying the registry (lines 467- 474 and 489-499).

Error message (new plugin):

2026-03-11 14:38:56.571 +01:00 [INFO] [mk_sql::setup]: Using config file: C:\ProgramData\checkmk\agent\config\mk-sql.yml
2026-03-11 14:38:56.576 +01:00 [INFO] [mk_sql::config::ms_sql]: localhost is defined, adding registry instances
2026-03-11 14:38:56.576 +01:00 [ERROR] [mk_sql::platform::registry]: Failed to open registry key: Os { code: 2, kind: NotFound, message: "Das System kann die angegebene Datei nicht finden." }
2026-03-11 14:38:56.577 +01:00 [INFO] [mk_sql::config::ms_sql]: Found 2 SQL server instances in REGISTRY: [ SQL99:Port(63262), SQL98:Port(57026) ]
2026-03-11 14:38:56.577 +01:00 [INFO] [mk_sql::setup]: Cache dir exists "C:\\ProgramData\\checkmk\\agent\\state\\mk-sql-cache\\mssql-990DD514E6A69BCC"
2026-03-11 14:38:56.578 +01:00 [INFO] [mk_sql::ms_sql::instance]: Using cache dir "C:\\ProgramData\\checkmk\\agent\\state\\mk-sql-cache\\mssql-990DD514E6A69BCC"
2026-03-11 14:38:56.578 +01:00 [INFO] [mk_sql::ms_sql::instance]: Generating main data
2026-03-11 14:38:56.578 +01:00 [INFO] [mk_sql::ms_sql::instance]: Finding instances...
2026-03-11 14:38:56.578 +01:00 [INFO] [mk_sql::ms_sql::client]: Local connection by port `localhost:1433`
2026-03-11 14:38:56.578 +01:00 [INFO] [mk_sql::ms_sql::client]: Connecting to addr 'localhost:1433'...
2026-03-11 14:38:57.728 +01:00 [WARN] [mk_sql::ms_sql::client]: Timeout: deadline has elapsed when creating client from config
2026-03-11 14:38:57.728 +01:00 [ERROR] [mk_sql::ms_sql::instance]: Failed to create main client: Timeout: deadline has elapsed when creating client from config
2026-03-11 14:38:57.728 +01:00 [INFO] [mk_sql::ms_sql::instance]: Finding instances by SQL Browser
2026-03-11 14:38:57.728 +01:00 [WARN] [mk_sql::ms_sql::instance]: Error discovering instances: Impossible to connect
2026-03-11 14:38:57.728 +01:00 [INFO] [mk_sql::ms_sql::instance]: Found 0 instances by discovery: [  ]
2026-03-11 14:38:57.728 +01:00 [INFO] [mk_sql::ms_sql::instance]: Add custom instance SQL98 
2026-03-11 14:38:57.728 +01:00 [INFO] [mk_sql::ms_sql::instance]: Add custom instance SQL99 
2026-03-11 14:38:57.729 +01:00 [ERROR] [mk_sql::platform::registry]: Failed to open registry key: Os { code: 2, kind: NotFound, message: "Das System kann die angegebene Datei nicht finden." }
2026-03-11 14:38:57.729 +01:00 [INFO] [mk_sql::ms_sql::instance]: Connecting using port from endpoint 57026
2026-03-11 14:38:57.729 +01:00 [INFO] [mk_sql::ms_sql::client]: Local connection by port `localhost:57026`
2026-03-11 14:38:57.729 +01:00 [INFO] [mk_sql::ms_sql::client]: Connecting to addr 'localhost:57026'...
2026-03-11 14:38:58.739 +01:00 [WARN] [mk_sql::ms_sql::client]: Timeout: deadline has elapsed when creating client from config
2026-03-11 14:38:58.739 +01:00 [ERROR] [mk_sql::ms_sql::instance]: Error creating client for `SQL98`: Timeout: deadline has elapsed when creating client from config
2026-03-11 14:38:58.740 +01:00 [INFO] [mk_sql::ms_sql::instance]: Instance `SQL98` at port 57026 not found. Try to use named connection.
2026-03-11 14:38:58.740 +01:00 [INFO] [mk_sql::ms_sql::client]: Browse connection at port `localhost:`
2026-03-11 14:38:58.740 +01:00 [INFO] [mk_sql::ms_sql::client]: Named connection to addr localhost:1434

Hey @e71fc63353b1

the root cause is clear from the log: mk-sql always connects to localhost, even after finding the instances in the registry. In a WSFC FCI, the SQL instances listen on the Cluster VNN/IP — not localhost — so every connection attempt times out.

This is actually a known limitation, documented in Werk #15842:

“If several databases are running on a system, each using their own IP addresses, these must be explicitly specified in the configuration of the agent plug-in, as the addresses and ports are currently not yet found automatically.”

This applies to both 2.3 and 2.4

Workaround: Explicitly define the instances in mk-sql.yml with the Cluster VNN/IP:

yaml

mssql:
  main:
    authentication:
      username: ""
      type: integrated
  instances:
    - sid: SQL98
      auth:
        username: ""
        type: integrated
      connection:
        hostname: <your-cluster-vnn-or-ip>
        port: 57026
    - sid: SQL99
      auth:
        username: ""
        type: integrated
      connection:
        hostname: <your-cluster-vnn-or-ip>
        port: 63262

Auto-detection with “one rule for all cluster hosts” is not possible this way — but it gets monitoring working. The proper fix needs to happen in mk-sql itself (FCI detection via registry/WMI + connect to VNN).
Worth opening a support ticket over your CMK-Partner especially since mssql.vbs is being deprecated.

Other Idea is deployment …

Greetz Bernd

Depper dive why this is happening at the SQL Server /
Windows level:

Why localhost never works for a FCI

A SQL Server Failover Cluster Instance is by design not bound to any
physical node hostname. Instead, it uses a single Virtual Network Name
(VNN)
tied to a virtual IP address that moves with the active node on
failover. Clients must always connect via this VNN — connecting to the
physical node name or localhost will never reach the instance:

Windows Server Failover Cluster with SQL Server - SQL Server Always On | Microsoft Learn

Where the VNN is stored

The VNN is persisted in the cluster registry under:
HKLM\Cluster\Resources\{GUID}\Parameters → value: VirtualServerName

This is exactly the key mk-sql would need to read in order to detect the
correct connection endpoint for a FCI — instead of falling back to
localhost. The old mssql.vbs handled this correctly.

Microsoft documents these cluster registry parameters here:

Manually re-create registry keys for cluster resources - SQL Server | Microsoft Learn

So the fix in mk-sql would conceptually be straightforward: detect
whether a discovered instance is a FCI (e.g. via the Cluster subkey
under HKLM\SOFTWARE\Microsoft\Microsoft SQL Server\<instance>), read
the VirtualServerName from the cluster registry, and use that as the
connection hostname instead of localhost.

Greetz Bernd

Thanks Bernd, we’ll try to request the missing feature.
Fabian

Try the workaround quite easy and quick

After opening a ticket, the suggested solution was to enable the “Connection > Monitoring Backend > Use Odbc…” option.

With this option selected, FCI are detected automatically.

Thank you for posting this solution. I had exactly the same problem on my MSSQL-Windows-Clusters.

In my opinion the plugin should automatically detect the case, when the odbc backend has to be used on windows machines. I think this shouldn’t be too hard to implement.

The old plugin detected everything automatically, so the new one should also…