Failover nodes continuously switching Active/Standby after rebuilding secondary server

Failover nodes continuously switching Active/Standby after rebuilding secondary server

Hello ManageEngine Support Team,

We are experiencing a critical issue with OpManager Failover after rebuilding the secondary server.

Our environment:

  • Primary server:
    • OpManager Build: 12.8.709
    • Applications Manager Build: 178206
  • Secondary server:
    • OpManager Build: 12.8.709
    • Applications Manager Build: 178206

Previously, these two servers were configured as a working Failover pair.

During the recent upgrade, the secondary OpManager server stopped starting. We therefore completely uninstalled OpManager and Applications Manager from the secondary server and performed a clean installation.

After the clean installation, we followed the Failover configuration procedure:

  1. Installed the same OpManager and Applications Manager builds on both servers.
  2. Configured the required shared OpManager folders between the servers.
  3. Used the provided .bat script to import/synchronize the required files to the secondary server.
  4. Started the Failover configuration.

After enabling Failover, the system entered an unstable state.

The servers continuously attempt to change their roles:

  • the secondary server switches from Standby to Active;
  • shortly afterward, the roles change again;
  • the servers repeatedly alternate between Active and Standby;
  • during these role changes, the OpManager web interface and monitoring system are frequently unavailable.

The role switching appears to continue indefinitely, and the Failover pair does not reach a stable Primary/Standby state.

Could you please help us determine the cause of this behaviour and restore a stable Failover configuration?

Please also specify exactly which diagnostic data we should provide to speed up the investigation, including:

  • required log file names and paths from both servers;
  • Applications Manager Failover logs, if they are required;
  • output of any diagnostic scripts or commands;
  • Support Information File or complete logs archive, if required.
  • etc.

Please let us know whether we should temporarily stop one of the OpManager servers before collecting the logs or performing any further actions.

At the moment, the monitoring system is frequently unavailable because of the continuous Active/Standby role switching, so we would appreciate your assistance as soon as possible.

                        New to ADSelfService Plus?