# SLES 11 SP1 CLUSTER - NODE OFFLINE

**URL:** <https://forums.suse.com/t/sles-11-sp1-cluster-node-offline/21862>\
**Category:** SLES Configure-Administer\
**Created:** [November 10, 2011, 1:06pm UTC](https://forums.suse.com/t/sles-11-sp1-cluster-node-offline/21862 "2011-11-10T13:06:02Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![System1](https://avatars.discourse-cdn.com/v4/letter/s/51bf81/32.png) [@System1](https://forums.suse.com/u/System1)\
**Post date:** [November 10, 2011, 1:06pm UTC](https://forums.suse.com/t/sles-11-sp1-cluster-node-offline/21862/1 "2011-11-10T13:06:02Z")

</div>

Hello,

After I have configured the cluster with 2 nodes, both shows in their  
status as DC’s and the other node as offline (dirty). Please help me to  
troubleshoot this.

Node 1

## Code:

# ============ Last updated: Fri Nov 11 05:39:25 2011 Stack: openais Current DC: cluster1 - partition WITHOUT quorum Version: 1.1.2-2e096a41a5f9e184a1c1537c82c6da1093698eb5 2 Nodes configured, 2 expected votes 0 Resources configured.

Online: [cluster1]  
OFFLINE: [cluster2]

* * *

Node2

## Code:

# ============ Last updated: Fri Nov 11 05:07:45 2011 Stack: openais Current DC: cluster2 - partition WITHOUT quorum Version: 1.1.2-2e096a41a5f9e184a1c1537c82c6da1093698eb5 2 Nodes configured, 2 expected votes 0 Resources configured.

Node cluster1: UNCLEAN (offline)  
Online: [cluster2]

* * *

How I start troubleshooting this?

regards,

ccilleperuma.

## – ccilleperuma

ccilleperuma’s Profile: [http://forums.novell.com/member.php?userid=97239](http://forums.novell.com/member.php?userid=97239)  
View this thread: [http://forums.novell.com/showthread.php?t=448042](http://forums.novell.com/showthread.php?t=448042)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/flex022/uploads/suse/original/2X/5/5012ba89e3ffb5220dac47d5ea0ba032e2fe1cb6.png) [@system](https://forums.suse.com/u/system)\
**Post date:** [November 11, 2011, 1:26pm UTC](https://forums.suse.com/t/sles-11-sp1-cluster-node-offline/21862/2 "2011-11-11T13:26:02Z")

</div>

Hi ccilleperuma,

looks like some sort of communication problem to me. I recommend to  
have a closer look at the log output (probably syslog) - it can be  
rather verbose and tells you step by step what both nodes are trying to  
attempt.

Common suggestions are to check for IP connectivity, firewalling issues  
and alike. And double-check your config, ie concerning the cluster IP  
ports defined on both nodes…

Regards  
Jens

## – from the times when today’s “old school” was “new school” :eek:

jmozdzen’s Profile: [http://forums.novell.com/member.php?userid=32246](http://forums.novell.com/member.php?userid=32246)  
View this thread: [http://forums.novell.com/showthread.php?t=448042](http://forums.novell.com/showthread.php?t=448042)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/flex022/uploads/suse/original/2X/5/5012ba89e3ffb5220dac47d5ea0ba032e2fe1cb6.png) [@system](https://forums.suse.com/u/system)\
**Post date:** [November 18, 2011, 9:56am UTC](https://forums.suse.com/t/sles-11-sp1-cluster-node-offline/21862/3 "2011-11-18T09:56:01Z")

</div>

I will post a config, which I posted in a nother forum to get a clear  
idea of the situation.

Hi,

Sorry for the replying delay, I had another issue to solve.

In my cluster servers there is no ha.cf probably because of following

## Code:

## Ultimately it will change in SLES11, HA will be replaced with OpenAIS and follow the same packaging and naming convention according to the recent changes in the project.

and the openais.conf file says

## Code:

# This configuration file is not used any more

# Please refer to /etc/corosync/corosync.conf

* * *

So I list the corosync file here of cluster1

## Code:

aisexec {  
#Group to run aisexec as. Needs to be root for Pacemaker

group: root

#User to run aisexec as. Needs to be root for Pacemaker

user: root

}  
service {  
#Default to start mgmtd with pacemaker

use\_mgmtd: yes

ver: 0

name: pacemaker

}  
totem {  
#The mode for redundant ring. None is used when only 1 interface specified, otherwise, only active or passive may be choosen

rrp\_mode: none

#How long to wait for join messages in membership protocol. in ms

join: 60

#The maximum number of messages that may be sent by one processor on receipt of the token.

max\_messages: 20

#The virtual synchrony filter type used to indentify a primary component. Change with care.

vsftype: none

#The fixed 32 bit value to indentify node to cluster membership. Optional for IPv4, and required for IPv6. 0 is reserved for other usage

nodeid: 1

#How long to wait for consensus to be achieved before starting a new round of membership configuration.

consensus: 4000

#HMAC/SHA1 should be used to authenticate all message

secauth: on

#How many token retransmits should be attempted before forming a new configuration.

token\_retransmits\_before\_loss\_const: 10

#How many threads should be used to encypt and sending message. Only have meanings when secauth is turned on

threads: 1

#Timeout for a token lost. in ms

token: 3000

#The only valid version is 2

version: 2

interface {  
#Network Address to be bind for this interface setting

bindnetaddr: 192.168.30.0

#The multicast address to be used

mcastaddr: 226.0.1.5

#The multicast port to be used

mcastport: 5454

#The ringnumber assigned to this interface setting

ringnumber: 0

}  
#To make sure the auto-generated nodeid is positive

clear\_node\_high\_bit: no

}  
logging {  
#Log to a specified file

to\_logfile: no

#Log to syslog

to\_syslog: yes

#Whether or not turning on the debug information in the log

debug: off

#Log timestamp as well

timestamp: on

#Log to the standard error output

to\_stderr: yes

#Logging file line in the source code as well

fileline: off

#Facility in syslog

syslog\_facility: daemon

}  
amf {  
#Enable or disable AMF

mode: disable

## }

hosts file of cluster1

## Code:

# 

# hosts This file describes a number of hostname-to-address

# mappings for the TCP/IP subsystem. It is mostly

# used at boot time, when no name servers are running.

# On small systems, this file can be used instead of a

# “named” name server.

# Syntax:

# 

# IP-Address Full-Qualified-Hostname Short-Hostname

# 

127.0.0.1 localhost

# special IPv6 addresses

::1 localhost ipv6-localhost ipv6-loopback

fe00::0 ipv6-localnet

## ff00::0 ipv6-mcastprefix ff02::1 ipv6-allnodes ff02::2 ipv6-allrouters ff02::3 ipv6-allhosts 127.0.0.2 cluster1.cbl cluster1 192.168.30.71 cluster2.cbl cluster2 192.168.30.70 cluster1.cbl cluster1

hosts file of cluster2

## Code:

# 

# hosts This file describes a number of hostname-to-address

# mappings for the TCP/IP subsystem. It is mostly

# used at boot time, when no name servers are running.

# On small systems, this file can be used instead of a

# “named” name server.

# Syntax:

# 

# IP-Address Full-Qualified-Hostname Short-Hostname

# 

127.0.0.1 localhost

# special IPv6 addresses

::1 localhost ipv6-localhost ipv6-loopback

fe00::0 ipv6-localnet

## ff00::0 ipv6-mcastprefix ff02::1 ipv6-allnodes ff02::2 ipv6-allrouters ff02::3 ipv6-allhosts 127.0.0.2 cluster2.cbl cluster2 192.168.30.70 cluster1.cbl cluster1 192.168.30.71 cluster2.cbl cluster2

This is the output of crm\_mon -1  
Cluster1

## Code:

# ============ Last updated: Sat Nov 19 02:07:42 2011 Stack: openais Current DC: cluster1 - partition WITHOUT quorum Version: 1.1.2-2e096a41a5f9e184a1c1537c82c6da1093698eb5 2 Nodes configured, 2 expected votes 0 Resources configured.

Node cluster2: UNCLEAN (offline)  
Online: [cluster1]

* * *

Cluster2

## Code:

# Last updated: Sat Nov 19 02:07:35 2011 Stack: openais Current DC: cluster2 - partition WITHOUT quorum Version: 1.1.2-2e096a41a5f9e184a1c1537c82c6da1093698eb5 2 Nodes configured, 2 expected votes 0 Resources configured.

Node cluster1: UNCLEAN (offline)  
Online: [cluster2]

* * *

I have e0 of both servers configured to 192.100.100.70 & 71 which  
connects to the out side network  
and e1 of both configured as 192.168.30.70 & 71 and connect through a  
cross cable

Yes, I have installed heartbeat and peacemaker both but not yet  
configured the peacemaker as i want to check the cluster connectivity  
first  
In heartbeat communication channels, I have given 192.168.30.0 as Bind  
network address as it only shows subnets (192.100.100.0/192.168.30.0) to  
select.  
Also it asks for mulicats address and port which i assigned  
226.0.1.5:5454 and 226.0.1.6:5454 respectively  
Cluster1 node id is 1 and cluster2 is 2  
rrp mode none for both as i dont have redundant channels

I hope this will help you to get an idea of my setup. Thanks very much  
for the interest you shown in this issue.

Regards,

ccilleperuma.

## – ccilleperuma

ccilleperuma’s Profile: [http://forums.novell.com/member.php?userid=97239](http://forums.novell.com/member.php?userid=97239)  
View this thread: [http://forums.novell.com/showthread.php?t=448042](http://forums.novell.com/showthread.php?t=448042)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/flex022/uploads/suse/original/2X/5/5012ba89e3ffb5220dac47d5ea0ba032e2fe1cb6.png) [@system](https://forums.suse.com/u/system)\
**Post date:** [November 18, 2011, 3:36pm UTC](https://forums.suse.com/t/sles-11-sp1-cluster-node-offline/21862/4 "2011-11-18T15:36:01Z")

</div>

Hi ccilleperuma,  
[color=blue]

> Also it asks for mulicats address and port which i assigned[/color]  
> 226.0.1.5:5454 and 226.0.1.6:5454 respectively

You mean you have told the nodes to communicate via different multicast  
addresses? That already could be the cause of the split - all nodes of a  
cluster use the same multicast channel to communiate.

With regards  
Jens

## – from the times when today’s “old school” was “new school” :eek:

jmozdzen’s Profile: [http://forums.novell.com/member.php?userid=32246](http://forums.novell.com/member.php?userid=32246)  
View this thread: [http://forums.novell.com/showthread.php?t=448042](http://forums.novell.com/showthread.php?t=448042)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/flex022/uploads/suse/original/2X/5/5012ba89e3ffb5220dac47d5ea0ba032e2fe1cb6.png) [@system](https://forums.suse.com/u/system)\
**Post date:** [November 21, 2011, 11:26am UTC](https://forums.suse.com/t/sles-11-sp1-cluster-node-offline/21862/5 "2011-11-21T11:26:02Z")

</div>

I have assigned the same multicast address and restarted the cluster  
service. But the issue is still the same.  
Time scync server is a must for the cluster servers?

## – ccilleperuma

ccilleperuma’s Profile: [http://forums.novell.com/member.php?userid=97239](http://forums.novell.com/member.php?userid=97239)  
View this thread: [http://forums.novell.com/showthread.php?t=448042](http://forums.novell.com/showthread.php?t=448042)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/flex022/uploads/suse/original/2X/5/5012ba89e3ffb5220dac47d5ea0ba032e2fe1cb6.png) [@system](https://forums.suse.com/u/system)\
**Post date:** [November 21, 2011, 12:36pm UTC](https://forums.suse.com/t/sles-11-sp1-cluster-node-offline/21862/6 "2011-11-21T12:36:02Z")

</div>

Hi ccilleperuma,  
[color=blue]

> Time scync server is a must for the cluster servers?[/color]

I’m not sure, but would recommend it anyhow - debugging distributed  
problems is a mess when time stamps differ.

Have you found any indicators in the cluster logs?

With regards  
Jens

## – from the times when today’s “old school” was “new school” :eek:

jmozdzen’s Profile: [http://forums.novell.com/member.php?userid=32246](http://forums.novell.com/member.php?userid=32246)  
View this thread: [http://forums.novell.com/showthread.php?t=448042](http://forums.novell.com/showthread.php?t=448042)

---

<div class="post-metadata">

**Author:** ![devanand\_s](https://avatars.discourse-cdn.com/v4/letter/d/a3d4f5/32.png) [@devanand\_s](https://forums.suse.com/u/devanand_s)\
**Post date:** [November 26, 2014, 6:13am UTC](https://forums.suse.com/t/sles-11-sp1-cluster-node-offline/21862/7 "2014-11-26T06:13:54Z")

</div>

Hi,

You can sync the config from the active node using Csync2 -xv

Regards  
Dev

[QUOTE=jmozdzen;1015]Hi ccilleperuma,  
[color=blue]

> Time scync server is a must for the cluster servers?[/color]

I’m not sure, but would recommend it anyhow - debugging distributed  
problems is a mess when time stamps differ.

Have you found any indicators in the cluster logs?

With regards  
Jens

## – from the times when today’s “old school” was “new school” :eek:

jmozdzen’s Profile: [http://forums.novell.com/member.php?userid=32246](http://forums.novell.com/member.php?userid=32246)  
View this thread: [http://forums.novell.com/showthread.php?t=448042](http://forums.novell.com/showthread.php?t=448042)[/QUOTE]

---

<div class="post-metadata">

**Author:** ![Jens-U](https://avatars.discourse-cdn.com/v4/letter/j/d78d45/32.png) [@Jens-U](https://forums.suse.com/u/Jens-U)\
**Post date:** [November 26, 2014, 1:44pm UTC](https://forums.suse.com/t/sles-11-sp1-cluster-node-offline/21862/8 "2014-11-26T13:44:53Z")

</div>

Hi Dev,

> > Time scync server is a must for the cluster servers?  
> > You can sync the config from the active node using Csync2 -xv

this thread is rather old (2011) and while csync2 is fine for synchronizing files in a cluster, it will not help concerning drifting system time bases. Setting up a proper ntp configuration might be more helpful to solve this issue.

Regards,  
Jens
