# Cluster agent in endless crash loop

**URL:** <https://forums.suse.com/t/cluster-agent-in-endless-crash-loop/16784>\
**Category:** SUSE Rancher Prime\
**Created:** [March 6, 2020, 3:07pm UTC](https://forums.suse.com/t/cluster-agent-in-endless-crash-loop/16784 "2020-03-06T15:07:05Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![seimic](https://sea2.discourse-cdn.com/flex022/user_avatar/forums.suse.com/seimic/32/6370_2.png) [@seimic](https://forums.suse.com/u/seimic)\
**Post date:** [March 6, 2020, 3:07pm UTC](https://forums.suse.com/t/cluster-agent-in-endless-crash-loop/16784/1 "2020-03-06T15:07:05Z")

</div>

Hi,

any idea how to investigate the root cause of the following error?  
It’s a user cluster registered as “Custom” RKE cluster in Rancher HA 2.3.5  
Nodes are created externally, no provider like AWS etc. used.

 ![image](https://us1.discourse-cdn.com/flex022/uploads/suse/original/2X/b/bd8b86bea69bce7dbaf49649fbfa14904ac5f98b.png)

After this message nothing else follows and the cluster agent crashes.

 ![image](https://us1.discourse-cdn.com/flex022/uploads/suse/original/2X/3/359a862f77d4c27ea3fc3a4f45eb58cfe2750c4d.png)

Setup:

Rancher HA Cluster: v2.3.5  
Nginx as external loadbalancer (in same network)  
User cluster: rancher/rancher-agent:v2.3.5  
Both Kubernetes v1.16.6-rancher1-2  
Everything in same network

Any help appreciated.

Kind regards,  
Michael

---

<div class="post-metadata">

**Author:** ![seimic](https://sea2.discourse-cdn.com/flex022/user_avatar/forums.suse.com/seimic/32/6370_2.png) [@seimic](https://forums.suse.com/u/seimic)\
**Post date:** [March 7, 2020, 11:06pm UTC](https://forums.suse.com/t/cluster-agent-in-endless-crash-loop/16784/2 "2020-03-07T23:06:55Z")

</div>

Installed the user cluster again, without to change the cidrs… now all pods in 10.42.0.0/16 subnet and services in 10.43.0.0/16…  
The error remains same with different IP:

level=fatal msg=“Get [https://10.43.0.1:443/apis/apiextensions.k8s.io/v1beta1/customresourcedefinitions:](https://10.43.0.1:443/apis/apiextensions.k8s.io/v1beta1/customresourcedefinitions:) dial tcp 10.43.0.1:443: i/o timeout”

Did the DNS checks from here: [https://rancher.com/docs/rancher/v2.x/en/troubleshooting/dns/](https://rancher.com/docs/rancher/v2.x/en/troubleshooting/dns/)  
There are no errors in coredns, upstream dns is reachable too from all nodes (in this example only one control and one worker)

cattle-node-agents work fine.  
Can’t find the reason. 😬

---

<div class="post-metadata">

**Author:** ![seimic](https://sea2.discourse-cdn.com/flex022/user_avatar/forums.suse.com/seimic/32/6370_2.png) [@seimic](https://forums.suse.com/u/seimic)\
**Post date:** [March 9, 2020, 12:34pm UTC](https://forums.suse.com/t/cluster-agent-in-endless-crash-loop/16784/3 "2020-03-09T12:34:28Z")

</div>

OK, replaced Weave CNI trough Canal and it works… but I would prefer Weave because of encryption.

This failes as mentioned above

```
network:
    options:
      flannel_backend_type: vxlan
      plugin: weave
        weave_network_provider:
          password: ...

```

Canal works directly

```
network:
    options:
      flannel_backend_type: vxlan
      plugin: canal

```

Its a very basic setup with one control plane and one worker to evaluate the whole automated setup.  
([https://rancher.com/docs/rke/latest/en/config-options/add-ons/network-plugins/](https://rancher.com/docs/rke/latest/en/config-options/add-ons/network-plugins/))

Kind regards,  
Michael
