Skip to content

fix(roles): keep the Icinga API credentials away from local users - #420

Open
NavidSassan wants to merge 7 commits into
mainfrom
fix/icinga-credentials-out-of-reach
Open

NavidSassan wants to merge 7 commits into
mainfrom
fix/icinga-credentials-out-of-reach

Conversation

@NavidSassan

Copy link
Copy Markdown
Member

Extracted from #396, which it does not depend on.

The password of the Icinga API user was readable by every local user on hosts deployed with setup_basic, and borg-backup and do-reboot passed it to curl on the command line, where it shows up in the process list. The Rocket.Chat webhook of schedule_reboot, which carries its token, had the same problem.

Changes

  • role:tools: schedule-icinga-downtime moves from a shell function in /etc/profile.d/alias.sh, which every login shell sources, to /usr/local/sbin/schedule-icinga-downtime, mode 0700. The reboot alias calls it through root's PATH. The function also ran exit 1 on a missing argument, which closed the calling shell.
  • role:borg_local: borg-backup is 0700 instead of 0755, and curl reads the credentials as a header from a pipe that the printf builtin fills instead of getting them with --user. The clamd@scan downtime request had a stray quote and a trailing comma, so Icinga answered 400 and the downtime was never set.
  • role:schedule_reboot: do-reboot is 0700 instead of 0744. The credentials reach curl the same way, the webhook URL through --config <(printf ...).
  • molecule schedule_reboot/no_reboot: asserts do-reboot is 0700 root.
  • molecule setup_basic: deploys schedule-icinga-downtime and asserts it is 0700 root, and that a login shell of nobody sees the reboot alias but neither the password nor a working script.

Admins who ran schedule-icinga-downtime as an unprivileged user need root now.

Tests

  • As part of Nextcloud: hardening, Unix sockets for Redis/Valkey, and credentials out of reach #396, on Rocky 10 against Icinga 2.14.6: the script sets the downtime for the host and its services, strace shows the password in no execve argument, the clamd@scan request returns 200 instead of 400. --header @file and --config verified with curl 7.61.1 on Rocky 8.
  • Molecule schedule_reboot/no_reboot and setup_basic have not been run on this branch yet.

markuslf and others added 7 commits October 2, 2026 16:11
The filter had a stray quote before the second match() and the body a
trailing comma, so the Icinga API answered 400 and borg-backup, which runs
curl with --silent and discards the output, never set the downtime
(verified against Icinga 2.14.6: 400 before, 200 now).
…l users

borg-backup was deployed 0755 and passed the credentials to curl with
--user. It is now 0700, since only root runs it (systemd timers), and the
credentials reach curl as a header read from a pipe that the printf builtin
fills.
…cal users

do-reboot was deployed 0744 and passed the Icinga API credentials with
--user and the Rocket.Chat webhook URL, which carries its token, as curl
arguments. It is now 0700, since only schedule-reboot.service runs it, the
credentials reach curl as a header from a pipe, and the URL through
--config <(printf ...), verified with curl 7.61.1 on Rocky 8. The Molecule
reboot scenario still sees the same Authorization header.
schedule-icinga-downtime was a shell function in /etc/profile.d/alias.sh,
which every login shell reads, so every local user could read the
credentials and set downtimes with them. It is now
/usr/local/sbin/schedule-icinga-downtime, readable and executable by root
only, which the reboot alias calls through root's PATH. The function also
ran `exit 1` on a missing argument, which closed the calling shell. Verified
on Rocky 10: the script sets the downtime for the host and all its services
on Icinga 2.14.6, and strace shows the password in no execve argument.
…o root

Set tools__icinga2_api_* so that schedule-icinga-downtime gets deployed.
Assert that it is 0700 root, and that a login shell of nobody sees the
reboot alias but not the password, and cannot run the script.
@NavidSassan

Copy link
Copy Markdown
Member Author

Molecule

Latest VM per distro, ansible-core 2.16:

  • schedule_reboot/no_reboot and schedule_reboot/reboot: green on Debian 13, Rocky 10 and Ubuntu 26.04.
  • setup_basic: on Rocky 10, verify.yml stops at the AIDE check, which reports changed __pycache__ files of the monitoring-plugins venv. That comes from main and is unrelated to this PR. The credential play of this PR, run on its own against the same VM, is green. Debian 13 and Ubuntu 26.04 could not be tested: the VMs were not reachable over SSH after creation (Molecule infrastructure, not this PR).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants