# SSD - failure

**URL:** https://forums.suse.com/t/ssd-failure/35470
**Category:** General Discussion
**Created:** [February 7, 2021, 12:18pm UTC](https://forums.suse.com/t/ssd-failure/35470 "2021-02-07T12:18:34Z")
**Posts on this page:** 16
**Page:** 1

<div class="post-metadata">

### Author: ![HANS-CHRISTOPH](https://avatars.discourse-cdn.com/v4/letter/h/5daacb/32.png) [@HANS-CHRISTOPH](https://forums.suse.com/u/HANS-CHRISTOPH)
#### Post date: [February 7, 2021, 12:18pm UTC](https://forums.suse.com/t/ssd-failure/35470/1 "2021-02-07T12:18:34Z")

</div>

I using SLED over some years. On my laptop that runs every day for some hours I use a SSD. Some years ago we had to exchange one disk - ok. shit happen. Now I get failures again.  
The laptop freeze totally and the HD controller is running continuously. Checking GSmartControl, no disk failures are found.  
Can there be a issue with SSD and linux or specific SUSE?

---

<div class="post-metadata">

### Author: ![HANS-CHRISTOPH](https://avatars.discourse-cdn.com/v4/letter/h/5daacb/32.png) [@HANS-CHRISTOPH](https://forums.suse.com/u/HANS-CHRISTOPH)
#### Post date: [February 7, 2021, 12:37pm UTC](https://forums.suse.com/t/ssd-failure/35470/2 "2021-02-07T12:37:10Z")

</div>

just the data as well:  
SLED15:/home/hans-christoph # hdparm -I /dev/sda

/dev/sda:

ATA device, with non-removable media  
Model Number: KINGSTON SA400S37240G  
Serial Number: 50026B767B01BD02  
Firmware Revision: SBFK71E0  
Transport: Serial, ATA8-AST, SATA 1.0a, SATA II Extensions, SATA Rev 2.5, SATA Rev 2.6, SATA Rev 3.0  
Standards:  
Supported: 11 10 9 8 7 6 5  
Likely used: 11  
Configuration:  
Logical max current  
cylinders 16383 16383  
heads 16 16  
sectors/track 63 63  
–  
CHS current addressable sectors: 16514064  
LBA user addressable sectors: 268435455  
LBA48 user addressable sectors: 468862128  
Logical Sector size: 512 bytes  
Physical Sector size: 512 bytes  
Logical Sector-0 offset: 0 bytes  
device size with M = 1024_1024: 228936 MBytes  
device size with M = 1000_1000: 240057 MBytes (240 GB)  
cache/buffer size = unknown  
Form Factor: 2.5 inch  
Nominal Media Rotation Rate: Solid State Device  
Capabilities:  
LBA, IORDY(can be disabled)  
Queue depth: 32  
Standby timer values: spec’d by Standard, no device specific minimum  
R/W multiple sector transfer: Max = 16 Current = 16  
DMA: mdma0 mdma1 mdma2 udma0 udma1 udma2 udma3 udma4 udma5 \*udma6  
Cycle time: min=120ns recommended=120ns  
PIO: pio0 pio1 pio2 pio3 pio4  
Cycle time: no flow control=120ns IORDY flow control=120ns  
Commands/features:  
Enabled Supported:  
\* SMART feature set  
Security Mode feature set  
\* Power Management feature set  
\* Write cache  
\* Look-ahead  
\* Host Protected Area feature set  
\* WRITE\_BUFFER command  
\* READ\_BUFFER command  
\* NOP cmd  
\* DOWNLOAD\_MICROCODE  
SET\_MAX security extension  
\* 48-bit Address feature set  
\* Device Configuration Overlay feature set  
\* Mandatory FLUSH\_CACHE  
\* FLUSH\_CACHE\_EXT  
\* SMART error logging  
\* SMART self-test  
\* General Purpose Logging feature set  
\* WRITE\_{DMA|MULTIPLE}\_FUA\_EXT  
\* 64-bit World wide name  
\* WRITE\_UNCORRECTABLE\_EXT command  
\* {READ,WRITE}\_DMA\_EXT\_GPL commands  
\* Segmented DOWNLOAD\_MICROCODE  
\* Gen1 signaling speed (1.5Gb/s)  
\* Gen2 signaling speed (3.0Gb/s)  
\* Gen3 signaling speed (6.0Gb/s)  
\* Native Command Queueing (NCQ)  
\* Phy event counters  
\* READ\_LOG\_DMA\_EXT equivalent to READ\_LOG\_EXT  
\* DMA Setup Auto-Activate optimization  
Device-initiated interface power management  
\* Software settings preservation  
\* DOWNLOAD MICROCODE DMA command  
\* SET MAX SETPASSWORD/UNLOCK DMA commands  
\* WRITE BUFFER DMA command  
\* READ BUFFER DMA command  
\* DEVICE CONFIGURATION SET/IDENTIFY DMA commands  
\* Data Set Management TRIM supported (limit 8 blocks)  
Security:  
Master password revision code = 65534  
supported  
not enabled  
not locked  
not frozen  
not expired: security count  
supported: enhanced erase  
20min for SECURITY ERASE UNIT. 60min for ENHANCED SECURITY ERASE UNIT.  
Logical Unit WWN Device Identifier: 50026b767b01bd02  
NAA : 5  
IEEE OUI : 0026b7  
Unique ID : 67b01bd02  
Checksum: correct

---

<div class="post-metadata">

### Author: ![malcolmlewis](https://sea2.discourse-cdn.com/flex022/user_avatar/forums.suse.com/malcolmlewis/32/11375_2.png) [@malcolmlewis](https://forums.suse.com/u/malcolmlewis)
#### Post date: [February 7, 2021, 3:17pm UTC](https://forums.suse.com/t/ssd-failure/35470/3 "2021-02-07T15:17:21Z")

</div>

@HANS-CHRISTOPH Hi, what about output from `smartctl -a /dev/sda`. Are you running btrfs, perhaps needs defrag/balance etc?

---

<div class="post-metadata">

### Author: ![HANS-CHRISTOPH](https://avatars.discourse-cdn.com/v4/letter/h/5daacb/32.png) [@HANS-CHRISTOPH](https://forums.suse.com/u/HANS-CHRISTOPH)
#### Post date: [February 9, 2021, 8:00pm UTC](https://forums.suse.com/t/ssd-failure/35470/4 "2021-02-09T20:00:08Z")

</div>

Vendor Specific SMART Attributes with Thresholds:  
ID# ATTRIBUTE\_NAME FLAG VALUE WORST THRESH TYPE UPDATED WHEN\_FAILED RAW\_VALUE  
1 Raw\_Read\_Error\_Rate 0x000a 100 100 000 Old\_age Always - 0  
9 Power\_On\_Hours 0x0012 100 100 000 Old\_age Always - 3910  
12 Power\_Cycle\_Count 0x0012 100 100 000 Old\_age Always - 1581  
148 Unknown\_Attribute 0x0000 255 255 000 Old\_age Offline - 8  
149 Unknown\_Attribute 0x0000 255 255 000 Old\_age Offline - 2  
167 Unknown\_Attribute 0x0022 100 100 000 Old\_age Always - 0  
168 SATA\_Phy\_Error\_Count 0x0012 100 100 000 Old\_age Always - 0  
169 Unknown\_Attribute 0x0000 100 100 000 Old\_age Offline - 6  
170 Bad\_Blk\_Ct\_Erl/Lat 0x0013 100 100 010 Pre-fail Always - 0/7  
172 Unknown\_Attribute 0x0032 100 100 000 Old\_age Always - 0  
173 MaxAvgErase\_Ct 0x0000 100 100 000 Old\_age Offline - 21 (Average 13)  
181 Program\_Fail\_Cnt\_Total 0x0012 100 100 000 Old\_age Always - 0  
182 Erase\_Fail\_Count\_Total 0x0000 255 255 000 Old\_age Offline - 1  
187 Reported\_Uncorrect 0x0032 100 100 000 Old\_age Always - 1  
192 Unsafe\_Shutdown\_Count 0x0012 100 100 000 Old\_age Always - 60  
194 Temperature\_Celsius 0x0023 066 048 000 Pre-fail Always - 34 (Min/Max 11/52)  
196 Not\_In\_Use 0x0000 100 100 000 Old\_age Offline - 2  
199 CRC\_Error\_Count 0x0032 100 100 000 Old\_age Always - 0  
218 CRC\_Error\_Count 0x0000 100 100 000 Old\_age Offline - 0  
231 SSD\_Life\_Left 0x0013 100 100 000 Pre-fail Always - 98  
233 Flash\_Writes\_GiB 0x0013 100 100 000 Pre-fail Always - 3219  
241 Lifetime\_Writes\_GiB 0x0012 100 100 000 Old\_age Always - 3142  
242 Lifetime\_Reads\_GiB 0x0012 100 100 000 Old\_age Always - 1505  
244 Average\_Erase\_Count 0x0000 100 100 000 Old\_age Offline - 13  
245 Max\_Erase\_Count 0x0000 100 100 000 Old\_age Offline - 21  
246 Total\_Erase\_Count 0x0000 100 100 000 Old\_age Offline - 160584

SMART Error Log Version: 1  
No Errors Logged

SMART Self-test log structure revision number 1  
Num Test\_Description Status Remaining LifeTime(hours) LBA\_of\_first\_error

# 1 Short offline Completed without error 00% 3907 -

# 2 Short offline Completed without error 00% 3905 -

# 3 Short offline Completed without error 00% 3890 -

# 4 Extended offline Completed without error 00% 3889 -

# 5 Short offline Completed without error 00% 3877 -

# 6 Short offline Completed without error 00% 3871 -

# 7 Short offline Completed without error 00% 3867 -

# 8 Short offline Completed without error 00% 3861 -

# 9 Short offline Completed without error 00% 3858 -

#10 Short offline Completed without error 00% 3854 -  
#11 Short offline Completed without error 00% 3850 -  
#12 Short offline Completed without error 00% 3842 -  
#13 Short offline Completed without error 00% 3834 -  
#14 Short offline Completed without error 00% 3819 -  
#15 Short offline Completed without error 00% 3815 -  
#16 Short offline Completed without error 00% 3814 -  
#17 Short offline Completed without error 00% 3808 -  
#18 Short offline Completed without error 00% 3795 -  
#19 Short offline Completed without error 00% 3781 -  
#20 Short offline Completed without error 00% 3762 -  
#21 Short offline Completed without error 00% 3752 -

SMART Selective self-test log data structure revision number 0  
Note: revision number not 1 implies that no selective self-test has ever been run  
SPAN MIN\_LBA MAX\_LBA CURRENT\_TEST\_STATUS  
1 0 0 Not\_testing  
2 0 0 Not\_testing  
3 0 0 Not\_testing  
4 0 0 Not\_testing  
5 0 0 Not\_testing  
Selective self-test flags (0x0):  
After scanning selected spans, do NOT read-scan remainder of disk.  
If Selective self-test is pending on power-up, resume after 0 minute delay.

---

<div class="post-metadata">

### Author: ![HANS-CHRISTOPH](https://avatars.discourse-cdn.com/v4/letter/h/5daacb/32.png) [@HANS-CHRISTOPH](https://forums.suse.com/u/HANS-CHRISTOPH)
#### Post date: [February 9, 2021, 8:03pm UTC](https://forums.suse.com/t/ssd-failure/35470/5 "2021-02-09T20:03:27Z")

</div>

I use BTRFS as native file system - as SUSE suggest.

---

<div class="post-metadata">

### Author: ![malcolmlewis](https://sea2.discourse-cdn.com/flex022/user_avatar/forums.suse.com/malcolmlewis/32/11375_2.png) [@malcolmlewis](https://forums.suse.com/u/malcolmlewis)
#### Post date: [February 9, 2021, 11:06pm UTC](https://forums.suse.com/t/ssd-failure/35470/6 "2021-02-09T23:06:22Z")

</div>

@HANS-CHRISTOPH Hi, that looks ok, what scheduler is in use?

```auto
cat /sys/block/sda/queue/scheduler

```

---

<div class="post-metadata">

### Author: ![HANS-CHRISTOPH](https://avatars.discourse-cdn.com/v4/letter/h/5daacb/32.png) [@HANS-CHRISTOPH](https://forums.suse.com/u/HANS-CHRISTOPH)
#### Post date: [February 10, 2021, 2:38pm UTC](https://forums.suse.com/t/ssd-failure/35470/7 "2021-02-10T14:38:28Z")

</div>

hans-christoph@SLED15:~\> cat /sys/block/sda/queue/scheduler  
[mq-deadline] kyber bfq none  
This is the result. The disc are not full, have space. It is strange. And this issues happen randomly. Sometimes I think it happens when using several programs and Internet? But I couldn’t find any issue here.

---

<div class="post-metadata">

### Author: ![malcolmlewis](https://sea2.discourse-cdn.com/flex022/user_avatar/forums.suse.com/malcolmlewis/32/11375_2.png) [@malcolmlewis](https://forums.suse.com/u/malcolmlewis)
#### Post date: [February 10, 2021, 3:11pm UTC](https://forums.suse.com/t/ssd-failure/35470/8 "2021-02-10T15:11:04Z")

</div>

@HANS-CHRISTOPH Hi, that’s the correct one for a SSD. Can you confirm that the likes of fstrim, btrfs defrag, balance etc has been run?

---

<div class="post-metadata">

### Author: ![malcolmlewis](https://sea2.discourse-cdn.com/flex022/user_avatar/forums.suse.com/malcolmlewis/32/11375_2.png) [@malcolmlewis](https://forums.suse.com/u/malcolmlewis)
#### Post date: [February 10, 2021, 4:15pm UTC](https://forums.suse.com/t/ssd-failure/35470/9 "2021-02-10T16:15:09Z")

</div>

@HANS-CHRISTOPH Hi forgot to add the command to check 😉 `systemctl list-timers` will show the info when it ran etc…

---

<div class="post-metadata">

### Author: ![HANS-CHRISTOPH](https://avatars.discourse-cdn.com/v4/letter/h/5daacb/32.png) [@HANS-CHRISTOPH](https://forums.suse.com/u/HANS-CHRISTOPH)
#### Post date: [February 11, 2021, 6:12am UTC](https://forums.suse.com/t/ssd-failure/35470/10 "2021-02-11T06:12:01Z")

</div>

NEXT LEFT LAST PASSED UNIT ACTIVATES  
Thu 2021-02-11 08:00:00 CET 49min left Thu 2021-02-11 07:00:07 CET 10min ago snapper-timeline.timer snapper-timeline.service  
Thu 2021-02-11 23:05:47 CET 15h left Tue 2021-02-09 10:08:38 CET 1 day 21h ago snapper-cleanup.timer snapper-cleanup.service  
Thu 2021-02-11 23:11:04 CET 16h left Tue 2021-02-09 10:13:56 CET 1 day 20h ago systemd-tmpfiles-clean.timer systemd-tmpfiles-clean.service  
Fri 2021-02-12 00:00:00 CET 16h left Thu 2021-02-11 06:50:48 CET 20min ago logrotate.timer logrotate.service  
Fri 2021-02-12 00:00:00 CET 16h left Thu 2021-02-11 06:50:48 CET 20min ago mandb.timer mandb.service  
Fri 2021-02-12 00:15:52 CET 17h left Thu 2021-02-11 06:50:48 CET 20min ago check-battery.timer check-battery.service  
Fri 2021-02-12 01:39:26 CET 18h left Thu 2021-02-11 06:50:48 CET 20min ago backup-sysconfig.timer backup-sysconfig.service  
Fri 2021-02-12 01:58:12 CET 18h left Thu 2021-02-11 06:50:48 CET 20min ago backup-rpmdb.timer backup-rpmdb.service  
Mon 2021-02-15 00:00:00 CET 3 days left Mon 2021-02-08 08:06:54 CET 2 days ago btrfs-balance.timer btrfs-balance.service  
Mon 2021-02-15 00:00:00 CET 3 days left Mon 2021-02-08 08:06:54 CET 2 days ago fstrim.timer fstrim.service  
Mon 2021-03-01 00:00:00 CET 2 weeks 3 days left Mon 2021-02-01 07:39:12 CET 1 weeks 2 days ago btrfs-scrub.timer btrfs-scrub.service

---

<div class="post-metadata">

### Author: ![HANS-CHRISTOPH](https://avatars.discourse-cdn.com/v4/letter/h/5daacb/32.png) [@HANS-CHRISTOPH](https://forums.suse.com/u/HANS-CHRISTOPH)
#### Post date: [February 11, 2021, 6:20am UTC](https://forums.suse.com/t/ssd-failure/35470/11 "2021-02-11T06:20:49Z")

</div>

It looks like all run well.  
See, I have 2 SLED 15.2 running. One on my old desktop with a HD. There are no problems with disc.  
One on my laptop with a SSD. I had exchanged one disk in past. The problems seems to be the same as last time. Random freezing of laptop. I that case, hard reboot is the only option. What I can see in that case is the SSD controller (LED) is running.  
Searching for “Linux SSD freeze” I can see other had same problem.

---

<div class="post-metadata">

### Author: ![malcolmlewis](https://sea2.discourse-cdn.com/flex022/user_avatar/forums.suse.com/malcolmlewis/32/11375_2.png) [@malcolmlewis](https://forums.suse.com/u/malcolmlewis)
#### Post date: [February 11, 2021, 2:04pm UTC](https://forums.suse.com/t/ssd-failure/35470/12 "2021-02-11T14:04:32Z")

</div>

Hi  
So if you look at the times when the above services ran, does it correspond with when you get a freeze?

I also suggest enable the magic sysrq key… [https://en.wikipedia.org/wiki/Magic\_SysRq\_key](https://en.wikipedia.org/wiki/Magic_SysRq_key) use `cat /proc/sys/kernel/sysrq` to see the current value, can also setup a `/etc/sysctl.d/10-magic-sysrq.conf` file…

---

<div class="post-metadata">

### Author: ![HANS-CHRISTOPH](https://avatars.discourse-cdn.com/v4/letter/h/5daacb/32.png) [@HANS-CHRISTOPH](https://forums.suse.com/u/HANS-CHRISTOPH)
#### Post date: [February 15, 2021, 7:02pm UTC](https://forums.suse.com/t/ssd-failure/35470/13 "2021-02-15T19:02:42Z")

</div>

Hi  
I’m not sure. The services run often and at night as can see on time stamp.  
The sysrq value is 184 - ??  
I can see I have to work with it a little bit more, but seems to have a lot of function. When freeze, ALT+F2 doesn’t work.

---

<div class="post-metadata">

### Author: ![malcolmlewis](https://sea2.discourse-cdn.com/flex022/user_avatar/forums.suse.com/malcolmlewis/32/11375_2.png) [@malcolmlewis](https://forums.suse.com/u/malcolmlewis)
#### Post date: [February 15, 2021, 7:42pm UTC](https://forums.suse.com/t/ssd-failure/35470/14 "2021-02-15T19:42:41Z")

</div>

@HANS-CHRISTOPH Hi, have a read here: [https://www.kernel.org/doc/html/latest/admin-guide/sysrq.html](https://www.kernel.org/doc/html/latest/admin-guide/sysrq.html)

Try the key combo and press the keys in the required order press alt and hold, press sys rq key and release, then press the following keys one at a time R E I S U B and then release the alt key. System should reboot…

---

<div class="post-metadata">

### Author: ![HANS-CHRISTOPH](https://avatars.discourse-cdn.com/v4/letter/h/5daacb/32.png) [@HANS-CHRISTOPH](https://forums.suse.com/u/HANS-CHRISTOPH)
#### Post date: [March 4, 2021, 9:38am UTC](https://forums.suse.com/t/ssd-failure/35470/15 "2021-03-04T09:38:14Z")

</div>

Now I found same failure on my Desktop as well. this has an regularly HD. I think I can relate that to Firefox. It happens wehn many windows are open and Firefox works in longer time.  
I always happens when working with Firefox

---

<div class="post-metadata">

### Author: ![malcolmlewis](https://sea2.discourse-cdn.com/flex022/user_avatar/forums.suse.com/malcolmlewis/32/11375_2.png) [@malcolmlewis](https://forums.suse.com/u/malcolmlewis)
#### Post date: [March 5, 2021, 2:42am UTC](https://forums.suse.com/t/ssd-failure/35470/16 "2021-03-05T02:42:53Z")

</div>

@HANS-CHRISTOPH Hi, sounds like you might need to look at firefox tweaks, eg disk cache to reduce that.  
How much system RAM?

Maybe reduce swappiness?

```auto
cat /etc/sysctl.d/98-swap.conf

#Reduce swappiness
vm.swappiness=1
vm.vfs_cache_pressure=50

```
