CMK version: Enterprise edition 2.1.0p16 OS version: Redhat 7.9
Hi all,
I am using enterprise edition 2.1.0p16, and I want to know the best practice for setup the maximum number of check attempt for service and maximum number of check attempt . Ideally I want the service down and host down alerts to be sent after 5 mins of down , should I just set both maximum number of check attempt for host and maximum number of check attempt for service to 5mins ? Because the enterprise edition use smart ping for host ping, so should I setup the rule for setting for host check via Smart PING as below to 5 min ? Will there be conflicts with Maximum number of check attempts for Host ? Please advise. I am really confused about what I should configure.
In community edition, I set both maximum number of check attemt for host and maximum number of check attemt for service to 5mins, and it works perfectly. But I am not sure what I should set in Enterprise edition because it uses smart ping instead.
Short answer - No
Leave the smart ping configuration as it is in the default settings (one packet in 15 seconds).
Calculate the check attempts for services and hosts as you want the state to be a hard state.
Recommendation is - the host should reach it hard state earlier than the service.
With smart ping and normal service check interval of 1 minute, it would look like this.
12 host check attempts = 3 minutes until host hard state = notification sent
4 service check attempts = 4 minutes until service hard state = notification sent
Leave the smart ping configuration as it is in the default settings (one packet in 15 seconds).
– I don’t see this one packet in 15 second setting, I only see the normal check interval for host check is 6 sec, and retry check interval host host checks is 6 seconds. Please check and advise if those setting need to be amended .
Alternatively, to make it simple to configure and understand , I would set the maximum number check attempt for service to 6 (1 check in every minute) , and leave the smart ping configuration to 300s , so the host will reach hard state in 300s (5 mins) , and service will reach hard state in 6 min, is it ok ?
The number of check attempts for host checks. The default setting is 1 and this means at the first error your host state will enter the hard failure state and sent a notification.
This i described in my previous post.
but i have been using this default number of check attempt for host check setting as 1 and the smart ping configuration to 300s since day one, which works ok, the notification is sent out after 300s rather than send out when first check attempt fails.