Skip to content

Fix issues with SMART monitoring script - #2837

Merged
enoch85 merged 5 commits into
nextcloud:mainfrom
pq2:patch-6
Sep 26, 2026
Merged

enoch85 merged 5 commits into
nextcloud:mainfrom
pq2:patch-6

Conversation

@pq2

@pq2 pq2 commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

See commit messages

🤖 AI (if applicable)

  • The content of this PR was partly or fully generated using AI

pq2 and others added 3 commits August 10, 2026 13:21
Add automatic -d device type detection for smartctl

Some drives (especially USB enclosures) require an explicit -d option
(e.g. sat, usbjmicron) for smartctl to communicate with them, since the
USB bridge is not always recognized automatically.

- Add get_smart_device_option() to probe common -d types (sat, sat,12,
  sat,16, usbjmicron, usbsunplus, usbcypress) and detect the correct
  one per drive during setup
- Run `update-smart-drivedb` before detection to increase the chance
  smartctl recognizes the bridge without needing a manual -d option
- Store detected options per drive in DRIVE_OPTS associative array and
  use them for the health check and manual troubleshooting hints
- Replace the generic DEVICESCAN line in /etc/smartd.conf with explicit
  per-drive lines carrying the correct -d option (Direct mode)
- Embed the detected DRIVE_OPTS directly into the generated weekly
  notification script as a static array, assuming USB enclosures don't
  change between runs (no runtime re-detection, no external file)

Signed-off-by: pq2 <github@nhelmschmidt.de>
NVMe drives can have many harmless "Invalid Field in Command" entries
in their Error Information Log without any real issue, so the existing
check for "No Errors Logged" incorrectly flagged healthy NVMe drives
(e.g. Samsung 970 EVO) as unhealthy.

- Detect NVMe drives by device name (nvme*) and only check the overall
  "PASSED" health status for them
- Keep the existing combined check (No Errors Logged + PASSED) for
  SATA/ATA drives unchanged
@enoch85

enoch85 commented Sep 26, 2026 •

Copy link
Copy Markdown
Member

Thanks @pq2! I pushed one commit on top of yours (8457b26). Both features stay: auto-detecting a -d type for USB-bridged drives, and no false "unhealthy" for NVMe drives. It fixes a few issues found in review:

  • Direct mode / smartd.conf: DEVICESCAN <directives> is back as the last line. Only drives that need a -d type get an explicit /dev/<kname> -d removable -d <type> ... line before it. With one line per drive, smartd exits at startup if a listed drive can't be registered (see smartd.cpp: "Unable to register device ... (no Directive -d removable). Exiting."), so an unplugged USB disk stopped all monitoring, and drives added later were never picked up. DEVICESCAN skips drives that are already listed, so systems without USB-bridged drives get the same smartd.conf as before.
  • Direct notification script: now runs smartctl -a -d $SMARTD_DEVICETYPE $SMARTD_DEVICE, so USB drives show real data instead of a device-type error.
  • Detection: inline loop, so smartctl runs once per normal drive. Extra types are only tried for USB drives (lsblk TRAN=usb). sat,16 is dropped because it is the same as sat. The "doesn't support smart monitoring" case prints the smartctl output again.
  • Weekly script: DRIVE_OPTS only holds drives that need a type. It is written into the script with declare -p, and smartctl falls back to -d auto.
  • NVMe: now one condition in the existing elif.

Known limitation: types are stored per kernel name (/dev/sdX), so re-run the script if USB drives are added or re-ordered.

- Direct mode: keep 'DEVICESCAN <directives>' as the last line of
  smartd.conf and only prepend explicit '/dev/<kname> -d removable -d <type>'
  lines for drives that need a -d type. An explicitly listed device that
  can't be registered makes smartd exit unless '-d removable' is set, and
  DEVICESCAN skips devices that are already listed, so normal drives
  (and drives added later) are monitored exactly like before.
- Direct notification script: pass -d $SMARTD_DEVICETYPE (the -d type
  from smartd.conf, or 'auto') to smartctl.
- Probe the drive types inline instead of a helper, so smartctl runs only
  once per normal drive, only try the extra types for USB drives, drop
  'sat,16' (same as 'sat'), and print the smartctl output again if a
  drive doesn't support smart monitoring.
- Only store drives that need a -d type in DRIVE_OPTS and write it to the
  weekly script with 'declare -p', falling back to '-d auto'.
- Simplify the NVMe check to one condition: NVMe drives are only checked
  for PASSED since their error log often has harmless entries.

Like the weekly DRIVE_OPTS, the explicit smartd.conf lines use the kernel
name (/dev/sdX), so re-run the script if USB drives are added or re-ordered.

Signed-off-by: enoch85 <mailto@danielhansson.nu>
@enoch85
enoch85 merged commit 9bdf650 into nextcloud:main Sep 26, 2026
1 check failed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants