# Some problems with network

**URL:** <https://forums.suse.com/t/some-problems-with-network/8473>\
**Category:** Rancher 1.x\
**Created:** [February 1, 2018, 9:26am UTC](https://forums.suse.com/t/some-problems-with-network/8473 "2018-02-01T09:26:16Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![chelius](https://avatars.discourse-cdn.com/v4/letter/c/e9c0ed/32.png) [@chelius](https://forums.suse.com/u/chelius)\
**Post date:** [February 1, 2018, 9:26am UTC](https://forums.suse.com/t/some-problems-with-network/8473/1 "2018-02-01T09:26:16Z")

</div>

Hello! There is a problem that I can not solve and I do not understand why it arises.  
I have test environment on Hyper-V, 4 hosts on RancherOS and Rancher v.1.6.14 (Cattle orchestration) with clear installation. Before RancherOS i’m truing with Ubuntu and CoreOS, problem is the same in all systems.  
My steps: In node with name ROS-MASTER i’m runned Rancher server with internal MySQL DB. After that i’m running rancer agent on ROS-MASTER and on 3 nodes ROS-01/02/03. I see how all infrastructure services are deployed and have a “green status”, network and healthchecs are work. Next, for example, I leave everything for the night and see the next picture in the morning

 ![image](https://us1.discourse-cdn.com/flex022/uploads/suse/original/2X/b/b2b0b4bd6031a68b43113bbd209b002a8fe6ae4c.png)  
As I understand it for myself. Services are in “initializing” state because they can’t be healchecked → Healthcheck containers can’t see other nodes because there are problems with IPsec…

ipsec-ipsec-router-2 logs [01.02.2018 11:17:3408[JOB] CHILD\_SA ESP/0x00000000/10.0.20.70 not found for dele - Pastebin.com](https://pastebin.com/qt7Y5agp)  
healthcheck-healthcheck-3 logs [01.02.2018 11:22:17time="2018-02-01T09:22:17Z" level=info msg="Starting haproxy - Pastebin.com](https://pastebin.com/CaJrCKVW)

I will be glad to any help! Thx!

---

<div class="post-metadata">

**Author:** ![superseb](https://sea2.discourse-cdn.com/flex022/user_avatar/forums.suse.com/superseb/32/2424_2.png) [@superseb](https://forums.suse.com/u/superseb)\
**Post date:** [February 1, 2018, 5:55pm UTC](https://forums.suse.com/t/some-problems-with-network/8473/2 "2018-02-01T17:55:40Z")

</div>

What is the status of the hosts (Infrastructure -\> Hosts)? From the logging it looks like there is a network interruption between the hosts, but ipsec should be able to recover from this.

---

<div class="post-metadata">

**Author:** ![chelius](https://avatars.discourse-cdn.com/v4/letter/c/e9c0ed/32.png) [@chelius](https://forums.suse.com/u/chelius)\
**Post date:** [February 2, 2018, 8:36am UTC](https://forums.suse.com/t/some-problems-with-network/8473/3 "2018-02-02T08:36:35Z")

</div>

Usualy status of Hosts is Active, this morning they were Disconected. I checked the connection between the hosts, it is present, but with high latency (20-400 ms). After rebooting the host, the latency becomes normal. What would it be, the performance of the host or Hyper-V?

UPD:  
I looked at the load average on the host, with 4 virtual cores and 2 GB of RAM, the load average varies from 20 to 30. I will investigate what causes such a load. Host with such parameters should be enough or need to increase resources?

---

<div class="post-metadata">

**Author:** ![nexcode](https://avatars.discourse-cdn.com/v4/letter/n/e99b99/32.png) [@nexcode](https://forums.suse.com/u/nexcode)\
**Post date:** [February 5, 2018, 9:26am UTC](https://forums.suse.com/t/some-problems-with-network/8473/4 "2018-02-05T09:26:36Z")

</div>

We decided to use it on production. Approximately every 12 hours falls ipsec.  
Now we are very sorry that we spent time on this not a stable solution.

---

<div class="post-metadata">

**Author:** ![nexcode](https://avatars.discourse-cdn.com/v4/letter/n/e99b99/32.png) [@nexcode](https://forums.suse.com/u/nexcode)\
**Post date:** [February 5, 2018, 9:30am UTC](https://forums.suse.com/t/some-problems-with-network/8473/5 "2018-02-05T09:30:51Z")

</div>

Now we just reboot the server about every 12 hours. I believe that you need to get rid of this product. It breaks, there is no good support. Bad choice. : (

---

<div class="post-metadata">

**Author:** ![nexcode](https://avatars.discourse-cdn.com/v4/letter/n/e99b99/32.png) [@nexcode](https://forums.suse.com/u/nexcode)\
**Post date:** [February 5, 2018, 9:35am UTC](https://forums.suse.com/t/some-problems-with-network/8473/6 "2018-02-05T09:35:51Z")

</div>

We use bare metal. Network between servers is always working perfectly.

---

<div class="post-metadata">

**Author:** ![nexcode](https://avatars.discourse-cdn.com/v4/letter/n/e99b99/32.png) [@nexcode](https://forums.suse.com/u/nexcode)\
**Post date:** [February 5, 2018, 10:14am UTC](https://forums.suse.com/t/some-problems-with-network/8473/7 "2018-02-05T10:14:26Z")

</div>

I think the problem is the following:  
When it breaks, agent on first host changes ip to 172.17.0.1 (or some such, I don’t remember exactly)  
After reboot it restore to normal ip addr. On other hosts this does not happen.

**At the moment I don’t know why the agent to change the ip and how can I prevent him to do it**.

---

<div class="post-metadata">

**Author:** ![nexcode](https://avatars.discourse-cdn.com/v4/letter/n/e99b99/32.png) [@nexcode](https://forums.suse.com/u/nexcode)\
**Post date:** [February 5, 2018, 10:38am UTC](https://forums.suse.com/t/some-problems-with-network/8473/8 "2018-02-05T10:38:41Z")

</div>

On your server agents are changing the IPs?

---

<div class="post-metadata">

**Author:** ![nexcode](https://avatars.discourse-cdn.com/v4/letter/n/e99b99/32.png) [@nexcode](https://forums.suse.com/u/nexcode)\
**Post date:** [February 5, 2018, 10:08pm UTC](https://forums.suse.com/t/some-problems-with-network/8473/9 "2018-02-05T22:08:36Z")

</div>

In this section there is a dirty solution to this problem:

> <https://github.com/rancher/rancher/issues/7817>
>
> \*\*Rancher Versions:\*\*
> Server: 1.4.1 (same with 1.3.x)
> healthcheck: 0.2.3
> ipse…c: 0.0.4
> scheduler: 0.4.0
> 
> \*\*Docker Version:\*\*
> 1.12.6 & 1.13.1
> \*\*OS and where are the hosts located? (cloud, bare metal, etc):\*\*
> Fresh install of Debian 8 bare metal (OVH) 
> \*\*Setup Details: (single node rancher vs. HA rancher, internal DB vs. external DB)\*\*
> single node rancher
> \*\*Environment Type: (Cattle/Kubernetes/Swarm/Mesos)\*\*
> Cattle
> \*\*Steps to Reproduce:\*\*
> \* Clear install of Debian 8 on dedicated-server (Debian 8.7 stable (Jessie) Server HOST-64-H - 64G Xeon D-1540) 
> \* Install docker-engine 1.12.6 or 1.13.1
> \* \`docker run -d --restart=unless-stopped -p 8080:8080 rancher/server\`
> \* \`docker run -d --privileged -v /var/run/docker.sock:/var/run/docker.sock -v /var/lib/rancher:/var/lib/rancher rancher/agent:v1.2.0 http://\*\*\*:8080/v1/scripts/\*\*\*\*\*\*\*\`
> 
> \*\*Results:\*\* 
> "Timeout getting IP address" on stack "healthcheck", "ipsec" and "scheduler".
> If I create a standalone container "redis" with network "bridge" it's start and run perfectly, but if I create one with network "managed" I have a error "Timeout getting IP address"
> 
> If you could give me any solution to get around this problem I will be very grateful 😀
> 
> (I have already tried \`{ "dns": \["8.8.8.8", "8.8.4.4"\], "dns-search": \["example.org"\] }\` in daemon.json, and change DNS in resolv.conf
> I also try to put the rancher host on another server)

The developers just advised to update… ☹

---

<div class="post-metadata">

**Author:** ![nexcode](https://avatars.discourse-cdn.com/v4/letter/n/e99b99/32.png) [@nexcode](https://forums.suse.com/u/nexcode)\
**Post date:** [February 5, 2018, 10:09pm UTC](https://forums.suse.com/t/some-problems-with-network/8473/10 "2018-02-05T22:09:58Z")

</div>

I use rancher/server 1.6.14 and docker 17.06.2-ce (rancheros 1.1.3)  
And what should I update?

---

<div class="post-metadata">

**Author:** ![chelius](https://avatars.discourse-cdn.com/v4/letter/c/e9c0ed/32.png) [@chelius](https://forums.suse.com/u/chelius)\
**Post date:** [February 6, 2018, 8:16am UTC](https://forums.suse.com/t/some-problems-with-network/8473/11 "2018-02-06T08:16:48Z")

</div>

In my case it turned out that the problem is in the containers. We migrate to ASP .Net Core, and the problem turned out to be in the applications that live in the containers. Applications consumed resources in a geometric progression, as a result of which the average load grew. We found a decision to turn off ServerGarbageCollector and the system has been working steadily for a couple of days. But we have not gotten to the production yet)

---

<div class="post-metadata">

**Author:** ![leodotcloud](https://sea2.discourse-cdn.com/flex022/user_avatar/forums.suse.com/leodotcloud/32/3103_2.png) [@leodotcloud](https://forums.suse.com/u/leodotcloud)\
**Post date:** [April 28, 2018, 8:32pm UTC](https://forums.suse.com/t/some-problems-with-network/8473/12 "2018-04-28T20:32:42Z")

</div>

@nexcode Are you still having trouble with IPSec? Where are your hosts running? Cloud/Datacenter? Can you check the output of `cat /proc/net/xfrm_stats` inside the ipsec containers? Do you see the errors going up? What version of rancher/server are you running?
