checkMK SNMP performance on one NODE, your experience, how to debug it? SNMP fetcher timeout very often!

Hello, what is recommended configuration/best practice for monitoring hunderds hosts via SNMP?

In my usecase I have on CMA 1.4.7 and CMK 2.0.0p18 about 1100 hosts/50K services.
There is about 300 hosts using SNMP, and aboe 210 host are ILOs.
I have a 5 minut check/3 minutes retry interval on ILOs, 2/2 minutes on other SNMP, some aplications and appliances. And i very often gets fetcher timeouts see cmc.log bellow , I increased number of fetcher to 64 than to 96 (using decreased from 80% to about 20% but problems remains]. I tried to create new site, to same CMA connected it via distributed monitoring and moved about 100 snmp host there, but problem remains? But load of CMA host generally is no problem. /cpus, memory, IOs. etc./

When I exported list of these 100 hosts, and add it to n new CMA in same network segment (not connect via distrubred monitorin yet), it happens app. one a day for one random host. And I doubled SNMP reading from host (main site, and this test sites) co it seems no problem on monitored hosts or network path…

Is it limit of network stack in CMA or generally How to identify it?

How many SNMP host you have on your sites? How is your experince with it and how to tune it?

cmc log:

2022-01-20 11:21:27 [3] [fetcher pool] [service "hostnameXXX044-ilo;Check_MK"] [helper 4000] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "hostnameXXX044-ilo" timed out after 60 seconds')
2022-01-20 11:21:31 [2] [fetcher pool] [service "hostnameXXX-ilo;Check_MK"] [helper 4054] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "hostnameXXX-ilo" timed out after 60 seconds')
2022-01-20 11:35:57 [2] [fetcher pool] [service "brocade-VL2;Check_MK"] [helper 21192] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "brocade-VL2" timed out after 60 seconds')
2022-01-20 11:35:57 [2] [fetcher pool] [service "hostnameXXX-ilo;Check_MK"] [helper 10396] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "hostnameXXX-ilo" timed out after 60 seconds')
2022-01-20 11:36:00 [2] [fetcher pool] [service "hostnameXXX-ilo;Check_MK"] [helper 21865] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "hostnameXXX-ilo" timed out after 60 seconds')
2022-01-20 11:36:04 [2] [fetcher pool] [service "hostnameXXX137rhel;Check_MK"] [helper 14406] [log] Error in PROGRAM fetcher: MKTimeout('Fetcher for host "hostnameXXX137rhel" timed out after 60 seconds')
2022-01-20 11:36:05 [2] [fetcher pool] [service "hostnameXXX135ms-ilo;Check_MK"] [helper 21636] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "hostnameXXX135ms-ilo" timed out after 60 seconds')
2022-01-20 11:36:05 [2] [fetcher pool] [service "hostnameXXX144;Check_MK"] [helper 28975] [log] Error in PROGRAM fetcher: MKTimeout('Fetcher for host "hostnameXXX144" timed out after 60 seconds')
2022-01-20 11:36:05 [2] [fetcher pool] [service "hostnameXXX-Cho;Check_MK"] [helper 3602] [log] Error in TCP fetcher: MKTimeout('Fetcher for host "hostnameXXX-Cho" timed out after 60 seconds')
2022-01-20 11:36:06 [2] [fetcher pool] [service "hostnameXXX;Check_MK"] [helper 30717] [log] Error in TCP fetcher: MKTimeout('Fetcher for host "hostnameXXX" timed out after 60 seconds')
2022-01-20 11:36:09 [3] [fetcher pool] [service "hostnameXXX-ilo;Check_MK"] [helper 9006] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "hostnameXXX-ilo" timed out after 60 seconds')
2022-01-20 11:36:14 [2] [fetcher pool] [service "hostnameXXXp-idrac;Check_MK"] [helper 4046] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "hostnameXXXp-idrac" timed out after 60 seconds')
2022-01-20 11:36:17 [2] [fetcher pool] [service "hostnameXXX;Check_MK"] [helper 4052] [log] Error in TCP fetcher: MKTimeout('Fetcher for host "hostnameXXX" timed out after 60 seconds')
2022-01-20 11:36:17 [2] [fetcher pool] [service "hostnameXXX;Check_MK"] [helper 4061] [log] Error in TCP fetcher: MKTimeout('Fetcher for host "hostnameXXX" timed out after 60 seconds')
2022-01-20 11:36:19 [2] [fetcher pool] [service "hostnameXXX142-ilo;Check_MK"] [helper 4074] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "hostnameXXX142-ilo" timed out after 60 seconds')
2022-01-20 11:36:21 [2] [fetcher pool] [service "hostnameXXX-idrac;Check_MK"] [helper 4058] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "hostnameXXX-idrac" timed out after 60 seconds')
2022-01-20 11:36:23 [2] [fetcher pool] [service "hostnameXXX-ilo;Check_MK"] [helper 4065] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "hostnameXXX-ilo" timed out after 60 seconds')
2022-01-20 11:36:26 [2] [fetcher pool] [service "hostnameXXX046-ilo;Check_MK"] [helper 4088] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "hostnameXXX046-ilo" timed out after 60 seconds')
2022-01-20 11:36:28 [2] [fetcher pool] [service "brocade-kz1;Check_MK"] [helper 4028] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "brocade-kz1" timed out after 60 seconds')
2022-01-20 11:36:31 [2] [fetcher pool] [service "brocade-VL1;Check_MK"] [helper 5780] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "brocade-VL1" timed out after 60 seconds')
2022-01-20 11:44:09 [3] [fetcher pool] [service "hostnameXXX;Check_MK"] [helper 25933] [log] [cycle 10036, command "342;hostnameXXX;checking;60"] memory usage increased from 70.14 MB to 119.25 MB, exiting
2022-01-20 11:44:11 [4] [fetcher pool] cannot send request: Broken pipe
2022-01-20 11:44:13 [3] [fetcher pool] [helper 25933] exited with status 14
2022-01-20 11:44:13 [5] [fetcher pool] [helper 26806] started, commandline: /omd/sites/icinga1/bin/fetcher
2022-01-20 11:50:50 [2] [fetcher pool] [service "hostnameXXX138rhel-ilo;Check_MK"] [helper 21636] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "hostnameXXX138rhel-ilo" timed out after 60 seconds')
2022-01-20 11:51:03 [2] [fetcher pool] [service "hostnameXXX-ilo;Check_MK"] [helper 19916] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "hostnameXXX-ilo" timed out after 60 seconds')
2022-01-20 11:51:05 [2] [fetcher pool] [service "hostnameXXX-ilo;Check_MK"] [helper 28975] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "hostnameXXX-ilo" timed out after 60 seconds')
2022-01-20 11:51:05 [3] [fetcher pool] [service "hostnameXXX-ilo;Check_MK"] [helper 27841] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "hostnameXXX-ilo" timed out after 60 seconds')
2022-01-20 11:51:06 [2] [fetcher pool] [service "hostnameXXX;Check_MK"] [helper 14406] [log] Error in TCP fetcher: MKTimeout('Fetcher for host "hostnameXXX" timed out after 60 seconds')
2022-01-20 11:51:07 [2] [fetcher pool] [service "hostnameXXX140-ilo;Check_MK"] [helper 9006] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "hostnameXXX140-ilo" timed out after 60 seconds')
2022-01-20 11:51:17 [2] [fetcher pool] [service "hostnameXXX-ilo;Check_MK"] [helper 4049] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "hostnameXXX-ilo" timed out after 60 seconds')
2022-01-20 11:51:21 [2] [fetcher pool] [service "hostnameXXX-CHD;Check_MK"] [helper 4019] [log] Error in TCP fetcher: MKTimeout('Fetcher for host "hostnameXXX-CHD" timed out after 60 seconds')
2022-01-20 11:51:23 [2] [fetcher pool] [service "hostnameXXX;Check_MK"] [helper 4059] [log] Error in TCP fetcher: MKTimeout('Fetcher for host "hostnameXXX" timed out after 60 seconds')
2022-01-20 11:51:24 [2] [fetcher pool] [service "hostnameXXX-ilo;Check_MK"] [helper 4067] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "hostnameXXX-ilo" timed out after 60 seconds')
2022-01-20 11:51:25 [2] [fetcher pool] [service "hostnameXXX041-ilo;Check_MK"] [helper 4046] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "hostnameXXX041-ilo" timed out after 60 seconds')
2022-01-20 11:51:28 [2] [fetcher pool] [service "hostnameXXX;Check_MK"] [helper 21865] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "hostnameXXX" timed out after 60 seconds')
2022-01-20 11:51:29 [2] [fetcher pool] [service "hostnameXXX;Check_MK"] [helper 4029] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "hostnameXXX" timed out after 60 seconds')
2022-01-20 11:51:32 [2] [fetcher pool] [service "brocade-VL2;Check_MK"] [helper 10621] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "brocade-VL2" timed out after 60 seconds')
2022-01-20 11:51:32 [2] [fetcher pool] [service "hostnameXXX033ms-ilo;Check_MK"] [helper 1026] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "hostnameXXX033ms-ilo" timed out after 60 seconds')
2022-01-20 11:51:37 [2] [fetcher pool] [service "hostnameXXX;Check_MK"] [helper 4066] [log] Error in TCP fetcher: MKTimeout('Fetcher for host "hostnameXXX" timed out after 60 seconds')
2022-01-20 11:51:38 [2] [fetcher pool] [service "hostnameXXX145-ilo;Check_MK"] [helper 4000] [log] Error in SNMP fetcher: MKTimeout('Fetcher for host "hostnameXXX145-ilo" timed out after 60 seconds')

Fetcher timeout has nothing to do with the load on the system. I think you need to investigate the hosts.
First some questions. How long takes a normal query to your iLO interface?
Same for the normal SNMP queries.
If you see here in the CheckMK service graphs regular spikes then you have a answer performance problem.

I have single monitoring instances with over 1k SNMP only devices (switches, routers and firewalls) there i have only timeouts from time to time on some well known devices (old and industrial switches). All others working there without problem.

Thank you for reply, sorry for my english. As I wrote:

“When I exported list of these 100 hosts, and add it to n new CMA in same network segment (not connect via distrubred monitorin yet), it happens app. one a day for one random host. And I doubled SNMP reading from host (main site, and this test sites) co it seems no problem on monitored hosts or network path…”

so it doesnt look like a problem on hosts… maybe some snmp setting like connection timeouts etc, but not sure :o( I tried to make modifications setting recommended in Monitoring via SNMP - Monitoring of SNMP devices with Checkmk article 5 but I was not successfull, so I get suspicion in some global setting, fetchers, load etc

What is your setting for SNMP rules like [Fetch intervals for SNMP sections] , [Bulk walk: Number of OIDs per bulk] etc, or using default settings?

Also:

Seems problem solved.

  • there was problem with many historical rules, and missing default for SNMP rules
  • there was problem on network (temporarily peak utilization) for monitoring devices on specific segment

Thank you for response and hints.

This topic was automatically closed 365 days after the last reply. New replies are no longer allowed. Contact an admin if you think this should be re-opened.