Skip to content

Nextcloud: hardening, Unix sockets for Redis/Valkey, and credentials out of reach - #396

Open
markuslf wants to merge 27 commits into
feat/redis-valkey-unixsocketfrom
fix/nextcloud
Open

markuslf wants to merge 27 commits into
feat/redis-valkey-unixsocketfrom
fix/nextcloud

Conversation

@markuslf

Copy link
Copy Markdown
Member

Follows the Nextcloud hardening guide and brings the nextcloud role up to the state the wordpress role reached, after comparing a hand-built production installation against what setup_nextcloud deploys.

What changes

Redis / Valkey over a Unix socket. Nextcloud talked to the cache over TCP on the loopback. The redis and valkey roles now open a socket at 0770 owned by their own group, the nextcloud role adds the web server user to that group, and setup_nextcloud configures SELinux before the services, since on RHEL 10 the shipped policy does not know Valkey yet and the daemon would otherwise keep running unconfined.

A dedicated database user. The installer received the database administrator and created an oc_ user that could connect from any host. nextcloud__database_login is now mandatory and setup_nextcloud creates that user with access from localhost only. Existing installations keep their user.

Requirements are checked before anything is touched. The role reads the installed Nextcloud, PHP and MariaDB versions and aborts with a message naming what to upgrade. The table covers Nextcloud 30 to 35.

The PHP module list matches what Nextcloud actually loads. APCu, bcmath, IMAP and memcached are removed: every cache lives in Redis or Valkey, and the remaining entries carry a comment saying why they are there. Checked against OC_Util::checkServer() and apps/settings/lib/SetupChecks/PhpModules.php.

Credentials no longer sit where local users can read them. Passwords reach occ on stdin instead of the command line, the Icinga API credentials move out of /etc/profile.d into a root-only script, and borg-backup, do-reboot, schedule-icinga-downtime and nextcloud-update are 0700. The borg_local downtime request also sent malformed JSON, which Icinga rejected with a 400.

fail2ban, opt-in. Behind a reverse proxy a ban would hit the proxy, so setup_nextcloud__skip_fail2ban defaults to true. Where it is switched on, the jail reads X-Forwarded-For and the proxies listed in the Nextcloud setting trusted_proxies are never banned. The log moves to /var/log/nextcloud/, because SELinux does not let fail2ban read the data directory.

Smaller fixes. The SMTP certificate is verified, the LDAP remnants report runs from its own timer, nextcloud-update adds missing primary keys and runs the expensive repairs, the data directory gets its ownership and SELinux label, and a failed installer attempt is retried instead of being skipped forever.

Breaking changes

Twelve entries under ### Breaking Changes in the CHANGELOG, each naming the variable to set to keep the previous behaviour. The ones that need action in the inventory before a run: nextcloud__version and nextcloud__database_login are mandatory now.

Testing

Every run was made against a Rocky 10 VM with Nextcloud 35, PHP 8.5 from Remi, MariaDB and Valkey, from an empty machine to a working installation, and repeated until a second run reported no change. The Molecule scenarios for valkey, schedule_reboot, fail2ban and setup_nextcloud were extended to assert the socket, the config ownership, the root-only scripts and the fail2ban counting through X-Forwarded-For. pre-commit run --all-files is clean, ansible-lint reports only findings that predate this branch.

Not tested: RHEL 8 and 9 with Redis from Remi, Debian and Ubuntu.

markuslf and others added 27 commits October 2, 2026 16:22
nextcloud-ldap-show-remnants.timer pointed to nextcloud-app-update.service,
so the monthly report never ran and the apps were updated once more instead.
…_datadir

mkdir, restorecon and the httpd_sys_rw_content_t fcontext were hardcoded to
/data, so a different nextcloud__datadir got neither.
…irs in nextcloud-update

nextcloud-app-update already ran db:add-missing-primary-keys, the server
update did not. maintenance:repair --include-expensive applies the mimetype
migrations that `occ setupchecks` reports after an upgrade, since Nextcloud
leaves them to the administrator.
--database-pass and --admin-pass were readable in the process list. Without
them occ maintenance:install asks for the database password first and the
admin password second, one line of stdin each when there is no TTY (verified
with Nextcloud 35.0.0 on Rocky 10, PHP 8.3 and 8.5). The stdin is a block
scalar: '{{ a ~ "\n" ~ b }}' in single quotes reaches the command as a
literal backslash-n with ansible-core 2.16.
The script carries the credentials of the Icinga API user and was deployed
with 0755. Only root can run it anyway, since it restarts services and
switches to the web server user.
The role set mail_smtpstreamoptions ssl allow_self_signed=true,
verify_peer=false and verify_peer_name=false on every host. The entries are
now state: absent, so existing hosts get them removed; the empty 'ssl' array
left behind is a no-op, since the Mailer merges the options into its own
(verified against Nextcloud 35).
…xtcloud__app_configs

The force and state subkeys of nextcloud__apps were listed under
nextcloud__app_configs, and nextcloud__apps claimed present/absent with a
default of present, while the task defaults to enabled and the module knows
four states. Also drop the trailing spaces from the keys of the reverse proxy
example.
…ud forbids anyway

'\n', '\r' and '\u0000' in single quotes reached Nextcloud as literal
backslash sequences, not as control characters. Nextcloud always forbids the
backslash and the characters 0 to 31 (lib/private/Files/FilenameValidator.php,
OCP\Constants::FILENAME_INVALID_CHARS), so the entries never had an effect.
'|' moves up to index 6 and indexes 7 to 9 are removed from existing hosts.
The filter is the one from the Nextcloud hardening guide. Verified with
fail2ban-regex 1.1.0 on Rocky 10 against the log of Nextcloud 35.0.0: failed
logins via web, WebDAV and OCS, disabled accounts, failed two-factor
challenges, IPv4 and IPv6. The trusted domain error is logged at info level
and only matches with loglevel 1 or lower. The regular expressions contain
'{%', so the template renders them raw.
fail2ban_t may read files of the logfile attribute, but not
httpd_sys_rw_content_t in the data directory: the jail found no log file
until fail2ban_t was made permissive (fail2ban 1.1.0 with fail2ban-selinux,
selinux-policy 42.1.18 on Rocky 10). /var/log/nextcloud is labeled
httpd_log_t, where httpd_t may create and append, and the log file is created
by the role, since Nextcloud silently falls back to the data directory if it
cannot create it itself. Rotation runs in nextcloud-jobs.service, outside
httpd_t.
…he trusted proxies

fail2ban belongs on hosts that clients reach directly, so it stays off by
default (setup_nextcloud__skip_fail2ban). When enabled, the playbook runs it
after nextcloud, since fail2ban does not start while the log of a jail is
missing, and injects the nextcloud jail. The addresses from the trusted_proxies
entries of nextcloud__sysconfig go to the new
fail2ban__jail_default_ignoreip__dependent_var, so no jail bans the proxy.
…alled Nextcloud

The requirements per major version come from the server source at the
release tags v30.0.0 to v35.0.0: the PHP range from lib/versioncheck.php,
outside of which Nextcloud refuses to run, and MIN_MARIADB from
apps/settings/lib/SetupChecks/SupportedDatabase.php. The major version is read
from version.php of the unpacked code, since nextcloud__version may just say
'latest'. Nextcloud sets no minimum for the Redis or Valkey server, only
phpredis >= 4.0.0.
Installing the policy build dependencies of the selinux role can update the
SELinux policy, and a daemon started under the older policy keeps its domain
until it is restarted. On a Rocky 10.0 host (selinux-policy 40.13.26), the
first start of Valkey labeled /usr/bin/valkey-server bin_t and ran it in
unconfined_service_t; the later update to 42.1.18 relabeled the binary to
redis_exec_t, but the process stayed unconfined, and httpd_t may not connect
to its socket there. setup_basic already runs selinux right after
policycoreutils.
…cket

nextcloud__redis_unixsocket and nextcloud__redis_group default to Valkey on
RHEL 10 and Redis elsewhere, as setup_nextcloud installs them. The web server
user joins the group; SELinux allows httpd_t to connect to redis_t and to
write redis_var_run_t sockets without a boolean (selinux-policy 42.1.18 on
Rocky 10).

'redis port' is removed instead of set to 0: RedisFactory picks 6379 for a
host name and the socket for a path when the port is missing, so removing the
port first and changing the host second keeps every step valid. Setting the
host to the socket path first left port 6379 next to it, after which every
occ call failed at bootstrap, including the one meant to fix the port
(reproduced on Rocky 10 with Nextcloud 35.0.0).
…ange

Both PHP-FPM restarts ran on every run, and occ background:cron had no
changed_when. restorecon now runs without --force, which reset the SELinux
user of the files the installer and Nextcloud created since the last run; the
targeted policy ignores the SELinux user. Apps installed during the run get a
restorecon of their own, since the notify_push binary needs bin_t and the
notify_push block that relabels it is skipped with nextcloud__skip_notify_push.
Nextcloud writes config.php at the start of maintenance:install, so a failed
attempt made `creates: config.php` skip the installer on every later run. The
role asks `occ status` instead: it reports installed=false for an empty
config.php (the notice goes to stderr), and the installer may overwrite the
config as long as config/CAN_INSTALL from the tarball exists, which only a
successful installation removes (lib/private/Config.php, verified with
Nextcloud 35.0.0).
… list

nextcloud-update passed the credentials to curl with --user. They now reach
curl as a header file read from a pipe that the printf builtin fills
(--header @file since curl 7.55, verified with curl 7.61.1 on Rocky 8).
CONTRIBUTING: software versions must always be mandatory variables. The
variable only picks the tarball of a new installation; updates of an
installed Nextcloud run through nextcloud-update. The Molecule scenario pins
'latest-35'.
The role handed the database administrator to occ maintenance:install, which
then created 'oc_<admin>'@'%' (lib/private/Setup/MySQL.php). setup_nextcloud
now creates the database and nextcloud__database_login@localhost through
mariadb_server, with the charset, collation and privileges the installer
would use, and the installer gets that user. Without the privilege to read
mysql.user, createSpecificUser() falls back to the provided credentials, so
config.php holds nextcloud__database_login and no 'oc_' user is created
(verified with Nextcloud 35.0.0 and MariaDB 11.4 on Rocky 10).
nextcloud__mariadb_login is gone; the version check queries MariaDB as the
new user as well.
The playbook installs Collabora on the Nextcloud host but leaves the reverse
proxy and the richdocuments settings to the inventory. The walkthrough
follows the setup in production: a hostname of its own for Collabora on the
proxy, forwarded to port 9980, and the proxy in wopi_allowlist.
The role keeps all Nextcloud caches (local, distributed, locking) in Redis
or Valkey, so APCu was installed but never used. Hosts that already have it
keep it: an inventory that sets memcache.local to APCu would break if the
role removed the package. Verified on Rocky 10: without php-pecl-apcu the
role reports no change and the Molecule verify passes.
The module list now names every module the role used to install with the
reason it stays or goes. present: what Nextcloud 35 requires or recommends
(OC_Util::checkServer(), SetupChecks/PhpModules.php), redis, ldap for
user_ldap, smbclient for SMB external storage, opcache and imagick, which
have setup checks of their own. absent: apcu and memcached (all caches live
in Redis / Valkey), bcmath (gmp covers WebAuthn and SFTP) and imap (only the
IMAP backend of user_external). php-json is dropped from the list instead of
set absent, since on RHEL it is only a name that php-common provides.
Verified on Rocky 10 with Remi PHP 8.5: the first run removes the four
packages, the second reports no change, and the Molecule verify passes.
CONTRIBUTING: no upstream version numbers in texts that stay in the repo,
they age with the next release. The role checks the requirements itself and
names the versions it expects in the error message, so the README points at
that instead of repeating "10.11+ for Nextcloud 35". The CHANGELOG entries
outside Breaking Changes are one sentence again.
The version check connected through /var/lib/mysql/mysql.sock on every
platform, which does not exist on Debian and Ubuntu
(/run/mysqld/mysqld.sock). Take the path from vars/<os>.yml, as
icingaweb2 does.
…version description

fail2ban moves to the roles that setup_nextcloud leaves off by default.
The Troubleshooting entries use the bold heading of roles/example and
get two blank lines above the section. argument_specs and the task
comment no longer offer 'latest' for nextcloud__version, matching the
README.
… in setup_nextcloud

Stat the socket Nextcloud uses (mode 0770, group redis or valkey) and
check that the config file keeps the <server>:root ownership and 0640
that the package's tmpfiles.d rule restores at every boot. This covers
the redis role on Rocky 8 and 9 and valkey on Rocky 10. Also re-wraps
one comment.
@NavidSassan

Copy link
Copy Markdown
Member

This PR was too large to review in one go, so two parts that are not specific to Nextcloud now have their own PRs:

The collabora commit is dropped: main already supports CODE 26.04.4 (b6cbeb1).

fix/nextcloud was rebuilt on top of #421 with the same commits, minus the ones that moved. c481b719 was split by scenario; its setup_nextcloud part stays here. The conflicts with main were in the CHANGELOG and in roles/fail2ban (portscan changes on main), both resolved by keeping both sides. The previous tip is cf5368eb.

For the review, in this order:

  1. Small fixes: LDAP remnants timer, datadir ownership, primary keys / expensive repairs, PHP-FPM restart only on change, installer retry, MariaDB socket path
  2. Credentials: occ passwords on stdin, nextcloud-update 0700, Icinga password out of the process list
  3. Breaking changes: nextcloud__version and nextcloud__database_login mandatory, version check, PHP modules, SMTP certificate verification
  4. Redis / Valkey socket on the Nextcloud side, SELinux before the services
  5. Logging to /var/log/nextcloud and fail2ban
  6. Docs

@NavidSassan
NavidSassan added this pull request to stack #422 October 2, 2026 14:31
@NavidSassan

Copy link
Copy Markdown
Member

Must fix

The nextcloud jail does not see failed WebDAV, OCS or app-password logins on Nextcloud 35.0.1. Upstream commit nextcloud/server@d25d5665a7c ("chore: Add more debug output", first released in v35.0.1) lowered the message in OC\User\Session::handleLoginFailed() (lib/private/User/Session.php) from warning to debug. With the role's loglevel 2 it never reaches nextcloud.log. Only the login form (LoggedInCheckCommand) still logs at warning. Reproduced in Molecule on Rocky 10 with 35.0.1: the request in setup_nextcloud/verify.yml:280 gets its 401 and Apache logs the X-Forwarded-For address, but there is no "Login failed" line, and fail2ban-client status nextcloud stays at 0. The filter itself is fine.

As it stands, the jail protects the login form only, while the filter comment, the fail2ban README and the CHANGELOG promise WebDAV and the OCS API as well. Options, roughly in order of preference:

  • report it upstream as a regression (fail2ban per the hardening guide relies on the warning)
  • send the failed login in verify.yml through the login form, and document that WebDAV/OCS logins are not seen on 35.0.1 and later until upstream changes it
  • loglevel 0 would bring the line back, but makes the log very verbose, so not a good default

fail2ban/install: the nextcloud filter test fails with ansible-core < 2.19. The sample lines are JSON and are passed as '{{ item["line"] }}' (extensions/molecule/fail2ban/install/verify.yml, from line 172). ansible-core 2.16 turns a template result that looks like a dict into a dict, so fail2ban-regex gets the Python repr with single quotes, and all three lines are reported as missed. ansible-core 2.19 no longer does that, which is probably why it passed for you. Writing the sample lines to a file and passing its path to fail2ban-regex works with every version.

Same backup problem as in #420. nextcloud-update carries the Icinga API credentials and was 0755. backup: true (roles/nextcloud/tasks/main.yml:543) copies the old file with its mode, so a 0755 copy with the credentials stays next to the new 0700 script, as do older backups. The same fix as in #420 applies.

Molecule

Rocky 10, Nextcloud 35.0.1, ansible-core 2.16: fail2ban/install and setup_nextcloud fail as described above. Every play before the fail2ban play in setup_nextcloud/verify.yml is green: installation, the dedicated database user, the SELinux domain of mariadbd, the Redis / Valkey socket, the root-only nextcloud-update, the SMTP certificate check and the LDAP remnants timer. The idempotence step did not run, since verify failed first.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants