diff --git a/CHANGELOG.md b/CHANGELOG.md index 7133197c5..bce2c0570 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -63,6 +63,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ### Changed +* LFOps requires community.general 7.0.0 or newer (still below 9.0.0), which `ansible-galaxy collection install linuxfabrik.lfops` pulls in, while a manually maintained collection list has to be raised. +* **plugin:bitwarden_item**: The lookup syncs the Bitwarden vault once per Ansible run instead of every 60 seconds, which makes runs with many lookups considerably faster, since each sync makes `bw serve` list the whole vault. Before it creates a missing item, it syncs again, so an item created elsewhere during the run is not created a second time. * **role:icingaweb2_module_generictts**: Downloads the module from Linuxfabrik, who maintain it since Icinga archived the original repository. The tarballs of v2.1.0 are identical. * **role:duplicity**: `/var/lib/aide` is backed up by default, so that the AIDE database can be compared with a copy outside the host. Hosts without AIDE are not affected. * **role:monitoring_plugins**: The source install removes plugins that an earlier run deployed and the checked-out version no longer carries. @@ -86,6 +88,13 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ### Fixed +* **plugin:bitwarden_item, module:bitwarden_item**: A failed sync of the Bitwarden vault, such as an "HTTP Error 400: Bad Request" or a timeout of `bw serve`, is tried again after 10, 30 and 60 seconds instead of aborting the run right away. +* **role:aide**: Before it creates the database, the role also waits for running `dnf-automatic` jobs on the Red Hat family and for the update jobs of the system_update role on every platform, not only for the apt jobs on Debian and Ubuntu. An update during the initialisation left files in the database that the first check then reported. +* **playbook:setup_basic**: With `setup_basic__skip_duplicity` or `setup_basic__skip_glances`, the playbook no longer builds the Python venv of the skipped role, which could abort the run with a pip error on hosts that do not back up with duplicity. +* **role:kernel_settings**: `sunrpc.*` settings, such as the `sunrpc.tcp_slot_table_entries` the mariadb_server role sets, survive a reboot. Until now the `sunrpc` module was not loaded again after a reboot on hosts without NFS, so TuneD could not apply the setting and the next run of the role failed in `tuned-adm verify`. +* **role:chrony**: The role aborts if neither `chrony__ntp_pools` nor `chrony__ntp_servers` is set, instead of leaving the host without a time source. +* **role:repo_baseos**: The Rocky Linux `security` repository works on Rocky 8 releases before 8.5, where dnf failed to download its metadata. +* **role:monitoring_plugins**: The source install deploys the OID lists and MIBs of the `snmp` plugin, which until now failed with "No such file or directory" on every host installed this way. * **role:bind**: A secondary zone with `type: 'slave'` is saved to its file again, so the secondary answers it after a restart without waiting for the primary. * **role:bind**: Reverse lookups for private and special-use addresses, such as `10.0.0.0/8` or `fd00::/8`, are answered locally, as BIND does by default, instead of waiting for the forwarders, which also no longer see the internal addressing. * **role:kernel_settings**: The role works with fedora.linux_system_roles 2.5.0 and later, which a fresh installation of LFOps pulls in. Until now the run aborted with "kernel_settings_transparent_hugepages must be null, one of always, madvise, never" unless `kernel_settings__transparent_hugepages__*_var` and `kernel_settings__transparent_hugepages_defrag__*_var` were set. diff --git a/COMPATIBILITY.md b/COMPATIBILITY.md index 2157079f0..8e3ddf745 100644 --- a/COMPATIBILITY.md +++ b/COMPATIBILITY.md @@ -35,7 +35,7 @@ Which Ansible role is proven to run on which OS? | dnf_makecache | | | x | x | x | | | | | | dnf_versionlock | | | x | x | (x) | | | | | | docker | | | x | (x) | (x) | | | | | -| duplicity | x | x | x | x | x | x | x | x | Fedora 35 | +| duplicity | x | x | x | x | x | x | x | x | | | elastic_agent | (x) | (x) | (x) | x | (x) | (x) | x | (x) | | | elastic_agent_fleet_server | (x) | (x) | (x) | x | (x) | (x) | x | (x) | | | elasticsearch | (x) | (x) | x | x | (x) | (x) | x | (x) | | diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 771d9f59d..17baa1783 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -980,11 +980,11 @@ __mariadb_server__python__modules__dependent_var: - name: 'python3-PyMySQL' mariadb_server__python__modules__dependent_var: '{{ __mariadb_server__python__modules__dependent_var - | linuxfabrik.lfops.platform_select(ansible_facts) + | linuxfabrik.lfops.platform_select(ansible_facts, default=[]) }}' ``` -`vars/main.yml` is auto-loaded at play parse, visible to every role in the play, and non-overridable from inventory like other `vars/`. Jinja evaluation is lazy, so the filter only runs when a consumer actually references the public variable. +`vars/main.yml` is auto-loaded at play parse, visible to every role in the play, and non-overridable from inventory like other `vars/`. The filter only runs when a consumer templates the public variable, but see below for what counts as that. Consumers stay simple - they reference the public variable directly, with no awareness of the selection mechanism: @@ -995,7 +995,13 @@ Consumers stay simple - they reference the public variable directly, with no awa - role: 'linuxfabrik.lfops.mariadb_server' ``` -The filter mirrors the precedence of `shared/tasks/platform-variables.yml` (least to most specific: `os_family`, `os_family + distribution_major_version`, `os_family + distribution_version`, `distribution`, `distribution + distribution_major_version`, `distribution + distribution_version`) and returns the value of the most specific present key. Pass `default=[]` (or whatever the consumer expects) when the value is optional on platforms not listed in the dict; otherwise an unmatched call raises an error. +The filter mirrors the precedence of `shared/tasks/platform-variables.yml` (least to most specific: `os_family`, `os_family + distribution_major_version`, `os_family + distribution_version`, `distribution`, `distribution + distribution_major_version`, `distribution + distribution_version`) and returns the value of the most specific present key. Without a `default`, a platform that matches no key raises an error. + +A published `__dependent_var` has to give a valid value on every host, whether or not its own role runs there. ansible-core up to 2.18 resolves every variable name of an expression before it evaluates it, so a consumer templates the value even in the untaken branch of a `ternary()` or an inline `if`. Gating the injection on a skip variable in the playbook therefore does not keep an error out. ansible-core 2.19 evaluates these lazily, but LFOps still has to run on older releases. In addition, `skip_injections: false` (see "`skip_role` Variables in Playbooks") uses the injection of a skipped role on purpose, on a host that role does not run on. So: + +* Always pass `default=[]` (or the empty value the consumer expects) to `platform_select` in a `__dependent_var`. +* Guard a value that depends on runtime state, as `roles/nextcloud/vars/main.yml` does below. +* If the publishing role cannot work without the value, it asserts its supported platforms in its own validation block, tagged `always`. The run then aborts only where that role runs, with a message that names it, instead of in the parameters of some other role. `roles/duplicity` is the reference. A `__dependent_var` has to be computable before the consuming role starts. Do not derive one from a variable that the consuming role itself only sets at runtime. A consuming role with a `meta/argument_specs.yml` templates every declared role parameter at role entry, before any of its own tasks run, and an undefined value anywhere inside the platform-keyed dictionary collapses the whole dictionary, so `platform_select` aborts the play with `input must be a dict keyed by platform identifier, got AnsibleUndefined`. diff --git a/extensions/molecule/kernel_settings/sysctl/inventory/group_vars/systems_under_test.yml b/extensions/molecule/kernel_settings/sysctl/inventory/group_vars/systems_under_test.yml index 7d59dd932..264104ae2 100644 --- a/extensions/molecule/kernel_settings/sysctl/inventory/group_vars/systems_under_test.yml +++ b/extensions/molecule/kernel_settings/sysctl/inventory/group_vars/systems_under_test.yml @@ -1,4 +1,6 @@ kernel_settings__sysctl__group_var: + - name: 'sunrpc.tcp_slot_table_entries' + value: 128 - name: 'vm.overcommit_memory' value: 1 diff --git a/extensions/molecule/kernel_settings/sysctl/verify.yml b/extensions/molecule/kernel_settings/sysctl/verify.yml index f07f630ca..8598c77a9 100644 --- a/extensions/molecule/kernel_settings/sysctl/verify.yml +++ b/extensions/molecule/kernel_settings/sysctl/verify.yml @@ -24,3 +24,29 @@ - name: 'Assert that transparent hugepages are set to madvise' ansible.builtin.assert: that: '"[madvise]" in (__molecule__transparent_hugepage_enabled_result["content"] | ansible.builtin.b64decode)' + + +# The sunrpc sysctls only exist while the sunrpc module is loaded. Nothing else loads it on a host +# without NFS, so check after a reboot that the module comes back on its own and TuneD applies the +# value again. +- name: 'Verify sunrpc.tcp_slot_table_entries survives a reboot' + hosts: 'systems_under_test' + gather_facts: false + tasks: + - name: 'systemctl reboot' + ansible.builtin.reboot: # yamllint disable-line rule:empty-values + become: true + + - name: 'Read sysctl sunrpc.tcp_slot_table_entries from procfs' + ansible.builtin.slurp: + src: '/proc/sys/sunrpc/tcp_slot_table_entries' + register: '__molecule__sysctl_sunrpc_tcp_slot_table_entries_result' + + - name: 'Assert that sunrpc.tcp_slot_table_entries is set to 128' + ansible.builtin.assert: + that: '__molecule__sysctl_sunrpc_tcp_slot_table_entries_result["content"] | ansible.builtin.b64decode | int == 128' + + - name: 'tuned-adm verify --ignore-missing' + ansible.builtin.command: 'tuned-adm verify --ignore-missing' + become: true + changed_when: false diff --git a/extensions/molecule/monitoring_plugins_source/install/verify.yml b/extensions/molecule/monitoring_plugins_source/install/verify.yml index 569f8a3a3..12dc1fe18 100644 --- a/extensions/molecule/monitoring_plugins_source/install/verify.yml +++ b/extensions/molecule/monitoring_plugins_source/install/verify.yml @@ -83,6 +83,28 @@ loop_control: label: '{{ item["item"] }}' + # The snmp plugin reads its OID lists and MIBs from directories next to itself. Without + # them every call ends in an I/O error before any SNMP request is sent. The manifest has to + # list each file, so that a removal takes them away and leaves the device definitions of the + # admin in place. + - name: 'Stat the default OID list of the snmp plugin' + ansible.builtin.stat: + path: '/usr/lib64/nagios/plugins/device-oids/any-any-any.csv' + register: '__molecule__snmp_device_oids' + + - name: 'Read the install manifest' + ansible.builtin.slurp: + src: '/usr/lib64/linuxfabrik-monitoring-plugins/install-manifest.txt' + register: '__molecule__manifest' + + - name: 'Assert the OID list is deployed, readable and listed in the manifest' + ansible.builtin.assert: + that: + - '__molecule__snmp_device_oids["stat"]["exists"]' + - '__molecule__snmp_device_oids["stat"]["mode"] == "0644"' + - '__molecule__snmp_device_oids["stat"]["pw_name"] == "root"' + - '"/usr/lib64/nagios/plugins/device-oids/any-any-any.csv" in (__molecule__manifest["content"] | b64decode).splitlines()' + # Only the modules and the LICENSE of the library are deployed, the same set the release zip # bundles; the documentation and development files of the repository have no business on a # monitored host. `__pycache__` appears as soon as a plugin has run. diff --git a/extensions/molecule/monitoring_plugins_source/remove/verify.yml b/extensions/molecule/monitoring_plugins_source/remove/verify.yml index 9504325e8..d23428278 100644 --- a/extensions/molecule/monitoring_plugins_source/remove/verify.yml +++ b/extensions/molecule/monitoring_plugins_source/remove/verify.yml @@ -89,6 +89,7 @@ - '/usr/lib64/nagios/plugins/about-me' - '/usr/lib64/nagios/plugins/assets' - '/usr/lib64/nagios/plugins/cpu-usage' + - '/usr/lib64/nagios/plugins/device-oids' register: '__molecule__leftovers' - name: 'Assert none of it is left' diff --git a/extensions/molecule/setup_basic/verify.yml b/extensions/molecule/setup_basic/verify.yml index c5c313fa9..30a3be37c 100644 --- a/extensions/molecule/setup_basic/verify.yml +++ b/extensions/molecule/setup_basic/verify.yml @@ -221,3 +221,14 @@ when: - 'ansible_facts["os_family"] == "RedHat"' + + # The inventory skips duplicity. Its venv must not be built anyway: python_venv used to get the + # duplicity venv regardless of setup_basic__skip_duplicity, and its pip install can fail. + - name: 'stat /opt/python-venv/duplicity' + ansible.builtin.stat: + path: '/opt/python-venv/duplicity' + register: '__molecule__duplicity_venv_stat_result' + + - name: 'Assert that the duplicity venv is absent while duplicity is skipped' + ansible.builtin.assert: + that: 'not __molecule__duplicity_venv_stat_result["stat"]["exists"]' diff --git a/galaxy.yml b/galaxy.yml index 44637d6af..de3388153 100644 --- a/galaxy.yml +++ b/galaxy.yml @@ -59,7 +59,7 @@ dependencies: 'ansible.utils': '*' 'ansible.windows': '*' 'community.crypto': '<3.0.0' # we need python 3.6 support to run against rhel8 - 'community.general': '<9.0.0' # we need python 3.6 support to run against rhel8 + 'community.general': '>=7.0.0,<9.0.0' # >=7.0.0: modprobe `persistent`; <9.0.0: we need python 3.6 support to run against rhel8 'community.grafana': '*' 'community.libvirt': '*' 'community.mongodb': '*' diff --git a/playbooks/setup_basic.yml b/playbooks/setup_basic.yml index 4e3a157c8..5f672331b 100644 --- a/playbooks/setup_basic.yml +++ b/playbooks/setup_basic.yml @@ -174,8 +174,8 @@ # venvs (further down in the Backups section). - role: 'linuxfabrik.lfops.python_venv' python_venv__venvs__dependent_var: '{{ - duplicity__python_venv__venvs__dependent_var + - glances__python_venv__venvs__dependent_var + (not setup_basic__skip_duplicity | d(false)) | ternary(duplicity__python_venv__venvs__dependent_var, []) + + (not setup_basic__skip_glances | d(false)) | ternary(glances__python_venv__venvs__dependent_var, []) }}' when: - 'not setup_basic__skip_python_venv | d(false)' diff --git a/plugins/lookup/bitwarden_item.py b/plugins/lookup/bitwarden_item.py index d0c985c67..12a614deb 100644 --- a/plugins/lookup/bitwarden_item.py +++ b/plugins/lookup/bitwarden_item.py @@ -23,9 +23,10 @@ - On success, the plugin returns the full Bitwarden item object. C(username) and C(password) are additionally lifted to the top level so they can be addressed without going through the C(login) sub-dictionary. - When I(name) is omitted, a title is generated automatically as C(hostname - purpose) (e.g. C(dbserver - MariaDB)) or just C(hostname) when no purpose is given. - Generated passwords use Python's C(secrets) module (cryptographically strong RNG), not the Bitwarden generator. This lifts the 128-character limit and allows arbitrary character sets, including hex. - - Items are read from a local on-disk cache backed by C(bw serve). A cached C(bw sync) is performed at most every 60 seconds, so consecutive lookups in the same play do not hammer the API. + - Items are read from a local on-disk cache backed by C(bw serve). The cache is synced once per Ansible run, since listing all items takes C(bw serve) up to half a minute; where the run cannot be identified (no C(/proc) on the controller), at most every 60 seconds. An item missing from the cache is looked up again after a fresh sync before it is created, so an item created elsewhere during a long run is not created twice. + - A failed sync is tried again after 10, 30 and 60 seconds before the lookup fails. - Lookups on the same controller run one at a time, so hosts that are processed in parallel and need the same missing item create it only once. - - Right after a sync, C(bw serve) can report an empty vault for a few seconds (U(https://github.com/bitwarden/clients/issues/23283)). The plugin then asks again for about ten seconds and fails rather than treat every item as missing. A vault that really is empty needs one item created by hand first. + - Right after a sync, C(bw serve) can report an empty vault for a few seconds (U(https://github.com/bitwarden/clients/issues/23283)). The plugin then asks again for about ten seconds, counts a vault that stays empty as a failed sync, and fails after the last sync attempt rather than treat every item as missing. A vault that really is empty needs one item created by hand first. notes: - Lookups are evaluated by the templating engine on the controller and have no notion of check mode, so a run with C(--check) creates a missing item for real. Set I(create) to C(false) to turn that into a failure. @@ -298,6 +299,8 @@ sample: 'root' """ +import os + from ansible.errors import AnsibleError from ansible.plugins.lookup import LookupBase from ansible.utils.display import Display @@ -311,6 +314,36 @@ # inspired by the lookup plugins lastpass (same topic) and redis (more modern) +def _read_proc(pid, name): + with open(f'/proc/{pid}/{name}', 'rb') as f: + return f.read() + + +def get_run_id(): + """Identify the Ansible run this lookup is evaluated in, or return None. + + Ansible forks its workers from the process of `ansible-playbook` without exec, so + they carry the same command line. Walking up the parents while the command line + stays the same ends at that process; its PID and start time identify the run, the + start time guards against a reused PID. Verified with ansible-core 2.16 on Fedora + 44, with the linear and the mitogen_linear (Mitogen 0.3.44) strategies: all workers + of a run got the same ID, and every run a different one. Without /proc, for example + on a macOS controller, there is no ID, and the caller falls back to a sync interval. + """ + try: + pid = os.getpid() + cmdline = _read_proc(pid, 'cmdline') + while True: + # the fields after the command name, which itself may contain ") " + fields = _read_proc(pid, 'stat').rsplit(b')', 1)[1].split() + ppid = int(fields[1]) + if ppid <= 1 or _read_proc(ppid, 'cmdline') != cmdline: + return f'{pid}:{int(fields[19])}' + pid = ppid + except (OSError, ValueError, IndexError): + return None + + class LookupModule(LookupBase): def run(self, terms, variables=None, **kwargs): self.set_options(var_options=variables, direct=kwargs) @@ -328,7 +361,8 @@ def _run_under_mutex(self, bw, terms): raise AnsibleError(bw.get_not_unlocked_message(status)) display.vvv('lfbwlp - run - bitwarden vault is unlocked') - bw.sync() + run_id = get_run_id() + synced = bw.sync(run_id=run_id) ret = [] for term in terms: @@ -372,6 +406,16 @@ def _run_under_mutex(self, bw, terms): result = bw.get_items( name, username, folder_id, collection_id, organization_id ) + if not result and not synced: + # the cache is only synced once per run, so an item created since then + # outside of this run would be missing from it and created a second time + display.vvv( + 'lfbwlp - run - not in the cache, syncing before creating it' + ) + synced = bw.sync(force=True, run_id=run_id) + result = bw.get_items( + name, username, folder_id, collection_id, organization_id + ) if len(result) > 1: raise AnsibleError( diff --git a/plugins/module_utils/bitwarden.py b/plugins/module_utils/bitwarden.py index e6ceda419..7ec2d7e15 100644 --- a/plugins/module_utils/bitwarden.py +++ b/plugins/module_utils/bitwarden.py @@ -47,6 +47,9 @@ class _NoopDisplay: def vvv(self, msg, **kwargs): pass + def warning(self, msg, **kwargs): + pass + display = _NoopDisplay() @@ -158,8 +161,10 @@ def prepare_multipart_no_base64(fields): CACHE_FILE = os.path.join(CACHE_DIR, 'lfops_bitwarden_cache.json') CACHE_VERSION = 2026032701 -# how long a process waits for another one to finish its sync, search and create -MUTEX_TIMEOUT = 300 +# how long a process waits for another one to finish its sync, search and create. covers a +# sync that needs all attempts of SYNC_RETRY_DELAYS, each of which can run into the timeout +# of open_url. +MUTEX_TIMEOUT = 600 # `bw serve` briefly reports an empty vault right after a sync, see # https://github.com/bitwarden/clients/issues/23283 @@ -168,6 +173,11 @@ def prepare_multipart_no_base64(fields): EMPTY_LIST_RETRIES = 5 EMPTY_LIST_RETRY_DELAY = 2 +# `bw serve` occasionally answers a sync with "HTTP Error 400: Bad Request" or does not +# answer within the timeout at all, and succeeds again on the next attempt. Seconds to wait +# before each further attempt; the error of the last one aborts the run. +SYNC_RETRY_DELAYS = (10, 30, 60) + class BitwardenException(Exception): pass @@ -371,19 +381,49 @@ def get_not_unlocked_message(self, status): '`bw serve`' ) - def sync(self, force=False, interval=60): + def sync( + self, force=False, interval=60, run_id=None, retry_delays=SYNC_RETRY_DELAYS + ): """Pull the latest vault data from server and repopulate the items cache. - Syncs only if the last sync was more than `interval` seconds ago, unless `force` is True. + + With a `run_id` (the lookup passes one that identifies the Ansible run), syncs + once per run, since listing all items takes `bw serve` up to half a minute. + Without one, syncs only if the last sync was more than `interval` seconds ago. + `force` syncs in any case. A failed sync is tried again after each of the + `retry_delays` seconds. Returns whether it synced. """ - if not force and time.time() - self._cache.get('sync_timestamp', 0) < interval: - display.vvv('lfbw - sync skipped, last sync was recent enough') - return - display.vvv(f'lfbw - syncing vault (force={force})') - self._api_call('sync', method='POST') - self._cache['items'] = self._list_items() + if not force and self._cache.get('items') is not None: + if run_id is not None and self._cache.get('sync_run_id') == run_id: + display.vvv('lfbw - sync skipped, already synced in this run') + return False + if ( + run_id is None + and time.time() - self._cache.get('sync_timestamp', 0) < interval + ): + display.vvv('lfbw - sync skipped, last sync was recent enough') + return False + display.vvv(f'lfbw - syncing vault (force={force}, run_id={run_id})') + for delay in (*retry_delays, None): + try: + self._api_call('sync', method='POST') + items = self._list_items() + break + except (BitwardenException, TimeoutError) as e: + # TimeoutError: open_url raises it unwrapped when `bw serve` accepts the + # connection but does not answer within the timeout + if delay is None: + raise + display.warning( + f'lfbw - syncing the Bitwarden vault failed ({to_native(e)}), ' + f'trying again in {delay}s' + ) + time.sleep(delay) + self._cache['items'] = items + self._cache['sync_run_id'] = run_id self._cache['sync_timestamp'] = time.time() display.vvv(f'lfbw - sync complete, cached {len(self._cache["items"])} items') self._save_cache() + return True def _list_items(self, retries=EMPTY_LIST_RETRIES, delay=EMPTY_LIST_RETRY_DELAY): """Return all items of the vault. An empty list is not trusted. diff --git a/plugins/modules/bitwarden_item.py b/plugins/modules/bitwarden_item.py index d9d40b0a6..1fbdc4fbc 100644 --- a/plugins/modules/bitwarden_item.py +++ b/plugins/modules/bitwarden_item.py @@ -23,8 +23,9 @@ - When I(name) is omitted, a title is generated automatically as C(hostname - purpose) (e.g. C(appsrv01 - MariaDB)) or just C(hostname) when no purpose is given. - On success, the module returns the full Bitwarden item object. C(username) and C(password) are additionally lifted to the top level so they can be addressed without going through the C(login) sub-dictionary. - Items are read from a local on-disk cache backed by C(bw serve). A cached C(bw sync) is performed at most every 60 seconds, so consecutive calls in the same play do not hammer the API. + - A failed sync is tried again after 10, 30 and 60 seconds before the module fails. - Module runs on the same host run one at a time, so hosts that are processed in parallel and need the same missing item create it only once. - - Right after a sync, C(bw serve) can report an empty vault for a few seconds (U(https://github.com/bitwarden/clients/issues/23283)). The module then asks again for about ten seconds and fails rather than treat every item as missing. A vault that really is empty needs one item created by hand first. + - Right after a sync, C(bw serve) can report an empty vault for a few seconds (U(https://github.com/bitwarden/clients/issues/23283)). The module then asks again for about ten seconds, counts a vault that stays empty as a failed sync, and fails after the last sync attempt rather than treat every item as missing. A vault that really is empty needs one item created by hand first. notes: - Only login items (Bitwarden type 1) are managed. Cards, secure notes and identities are out of scope. diff --git a/requirements.yml b/requirements.yml index 5c195a649..3940d378a 100644 --- a/requirements.yml +++ b/requirements.yml @@ -7,7 +7,7 @@ collections: - name: 'community.crypto' version: '<3.0.0' # we need python 3.6 support to run against rhel8 - name: 'community.general' - version: '<9.0.0' # we need python 3.6 support to run against rhel8 + version: '>=7.0.0,<9.0.0' # >=7.0.0: modprobe `persistent`; <9.0.0: we need python 3.6 support to run against rhel8 - name: 'community.grafana' - name: 'community.libvirt' - name: 'community.mongodb' diff --git a/roles/aide/README.md b/roles/aide/README.md index f573831b6..cbd96948d 100644 --- a/roles/aide/README.md +++ b/roles/aide/README.md @@ -30,7 +30,7 @@ This role is compatible with the following aide versions: * On Debian and Ubuntu the same applies to the updates of `unattended-upgrades`: a dpkg hook (`/usr/local/sbin/aide-dpkg-hook`, wired in by `/etc/apt/apt.conf.d/z00-linuxfabrik-aide`) runs a check before `unattended-upgrades` installs its first package, and if that check was clean, `aide-dpkg-update.service` updates the database once the upgrade has finished. `unattended-upgrades` installs in several steps, one dpkg run each, but the check and the update run only once per upgrade. The upgrade takes longer by that one check, while apt holds its lock. The hook only acts on dpkg runs of `unattended-upgrades`, so a package installed with apt by hand is reported, as on the Red Hat family. Changes made on the host while the upgrade runs are accepted along with it. * On Debian and Ubuntu the role creates an empty `/etc/apt/sources.list` if there is none, as on hosts with only `/etc/apt/sources.list.d/*.sources`. `unattended-upgrades` creates that file on every run otherwise, and the check before the first upgrade would report it. * The [duplicity](https://github.com/Linuxfabrik/lfops/tree/main/roles/duplicity) role backs up `/var/lib/aide` by default. An attacker with root privileges can replace the local database. The database only changes when it is updated (by this role, `--tags aide:update_db` or `aide:update_db_force`, the system_update role or `aide-dpkg-update.service`), so a local database that differs from the backup of a day without such an update points to tampering. -* Before it creates the database, the role waits for running `apt-daily.service` and `apt-daily-upgrade.service` jobs: a package installation during `aide --init` would leave entries without checksums in the database. +* Before it creates the database, the role waits for running package jobs: `apt-daily.service` and `apt-daily-upgrade.service` on Debian and Ubuntu, `dnf-automatic.service` and `dnf-automatic-install.service` on the Red Hat family, and `security-update.service` and `update-and-reboot.service` of the [system_update](https://github.com/Linuxfabrik/lfops/tree/main/roles/system_update) role everywhere. A package installation during `aide --init` would leave entries without checksums in the database. ## Known Limitations @@ -203,7 +203,7 @@ aide__timer_state: 'started' **The run aborts at a task that waits up to 5 minutes** -* An AIDE check (`aidecheck.service`) or an `apt-daily` job was still running after 5 minutes. A check takes longer on a large file system or a busy disk. Wait until `systemctl is-active aidecheck.service apt-daily.service apt-daily-upgrade.service` no longer reports `active` or `activating`, then run the role again. +* An AIDE check (`aidecheck.service`) or a package job was still running after 5 minutes. A check takes longer on a large file system or a busy disk. Wait until `systemctl is-active aidecheck.service` and the package jobs the task names no longer report `active` or `activating`, then run the role again. **`aidecheck.service` is failed** diff --git a/roles/aide/tasks/main.yml b/roles/aide/tasks/main.yml index ba744c219..ccc31daa3 100644 --- a/roles/aide/tasks/main.yml +++ b/roles/aide/tasks/main.yml @@ -273,15 +273,16 @@ # a package installation that runs while `aide --init` reads the file system leaves entries # without checksums in the new database, and the first check reports them. on Debian and # Ubuntu, apt-daily-upgrade.service (unattended-upgrades) typically does so on a freshly - # booted cloud image. units that do not exist report inactive. 5 minutes at most. - - name: 'systemctl is-active apt-daily.service apt-daily-upgrade.service (wait up to 5 minutes for running package jobs)' # noqa command-instead-of-module (read-only state query in a retry loop) - ansible.builtin.command: 'systemctl is-active apt-daily.service apt-daily-upgrade.service' - register: '__aide__apt_daily_active_result' - until: '"active" not in __aide__apt_daily_active_result["stdout_lines"] and "activating" not in __aide__apt_daily_active_result["stdout_lines"]' + # booted cloud image; on the Red Hat family dnf-automatic, and on every platform the jobs of + # the system_update role. units that do not exist report inactive. 5 minutes at most. + - name: 'systemctl is-active {{ (__aide__package_job_units + __aide__system_update_units) | join(" ") }} (wait up to 5 minutes for running package jobs)' # noqa command-instead-of-module (read-only state query in a retry loop) + ansible.builtin.command: 'systemctl is-active {{ (__aide__package_job_units + __aide__system_update_units) | join(" ") }}' + register: '__aide__package_jobs_active_result' + until: '"active" not in __aide__package_jobs_active_result["stdout_lines"] and "activating" not in __aide__package_jobs_active_result["stdout_lines"]' retries: 30 delay: 10 changed_when: false - failed_when: '"active" in __aide__apt_daily_active_result["stdout_lines"] or "activating" in __aide__apt_daily_active_result["stdout_lines"]' + failed_when: '"active" in __aide__package_jobs_active_result["stdout_lines"] or "activating" in __aide__package_jobs_active_result["stdout_lines"]' when: - 'not __aide__db_stat_result["stat"]["exists"]' diff --git a/roles/aide/vars/Debian.yml b/roles/aide/vars/Debian.yml index 9992d8a36..5c1744b32 100644 --- a/roles/aide/vars/Debian.yml +++ b/roles/aide/vars/Debian.yml @@ -1 +1,6 @@ __aide__binary_path: '/usr/bin/aide' + +# jobs of the distribution that install packages on their own, see the wait before `aide --init` +__aide__package_job_units: + - 'apt-daily-upgrade.service' + - 'apt-daily.service' diff --git a/roles/aide/vars/RedHat.yml b/roles/aide/vars/RedHat.yml index dd718c500..44b86ea47 100644 --- a/roles/aide/vars/RedHat.yml +++ b/roles/aide/vars/RedHat.yml @@ -1 +1,8 @@ __aide__binary_path: '/usr/sbin/aide' + +# jobs of the distribution that install packages on their own, see the wait before `aide --init`. +# dnf-automatic.service installs only with `apply_updates = yes`; the download and notifyonly +# variants install nothing. Verified against dnf-automatic 4.7, 4.14 and 4.20 on Rocky 8, 9 and 10. +__aide__package_job_units: + - 'dnf-automatic-install.service' + - 'dnf-automatic.service' diff --git a/roles/aide/vars/main.yml b/roles/aide/vars/main.yml index 3a25af7fa..713c4e720 100644 --- a/roles/aide/vars/main.yml +++ b/roles/aide/vars/main.yml @@ -42,3 +42,8 @@ __aide__update_db_refused: '{{ }}' __aide__update_db_left_alone_message: 'aide: The AIDE database was not updated, since the last AIDE check had reported changes or no check deployed by this role has run against the database yet. Review /var/log/aide/aide.log or the output of `aide --check`, then run the playbook with `--tags aide:update_db_force` to accept the current state as the new baseline.' + +# the jobs of the system_update role, which install packages on every platform +__aide__system_update_units: + - 'security-update.service' + - 'update-and-reboot.service' diff --git a/roles/chrony/README.md b/roles/chrony/README.md index c5fa8b24a..379e9ef04 100644 --- a/roles/chrony/README.md +++ b/roles/chrony/README.md @@ -12,7 +12,7 @@ This role installs and configures [chrony](https://chrony.tuxfamily.org/), a NTP ## How the Role Behaves * The configuration is fully templated: `/etc/chrony.conf` on the Red Hat family, `/etc/chrony/chrony.conf` on Debian and Ubuntu, each close to the file the distribution ships. Out-of-band edits are overwritten on the next run (a timestamped backup is kept). -* chronyd uses only the sources from `chrony__ntp_pools` and `chrony__ntp_servers`. If neither is set, the host has no time source. The distribution's default pools, time sources from DHCP and, on Debian and Ubuntu, `/etc/chrony/sources.d` are not used. Ubuntu 26.04 ships its default pools in `/etc/chrony/sources.d`, where chronyd would prefer them over every source from the inventory. +* chronyd uses only the sources from `chrony__ntp_pools` and `chrony__ntp_servers`. If neither is set, the role aborts, since the host would have no time source. The distribution's default pools, time sources from DHCP and, on Debian and Ubuntu, `/etc/chrony/sources.d` are not used. Ubuntu 26.04 ships its default pools in `/etc/chrony/sources.d`, where chronyd would prefer them over every source from the inventory. * On Debian and Ubuntu, drop-ins in `/etc/chrony/conf.d` are read at the beginning of the deployed `chrony.conf`, so the role's settings win over a drop-in that sets the same directive. Debian 13 and Ubuntu 26.04 read them at the end of their own `chrony.conf`, where a drop-in would win. Directives that add something instead of replacing it still take effect from a drop-in: a `pool`, `server` or `sourcedir` there adds time sources next to the ones from the inventory, and with the `prefer` option chronyd uses only those. Likewise, `allow` and `deny` add access rules. * The deployed `chrony.conf` loads no key file, so NTP sources are not authenticated with symmetric keys. RHEL 10's own `chrony.conf` does the same, while RHEL 8 and 9, Debian and Ubuntu load a key file that holds no keys (`/etc/chrony.keys`, `/etc/chrony/chrony.keys`). @@ -32,7 +32,7 @@ This role installs and configures [chrony](https://chrony.tuxfamily.org/), a NTP ## Mandatory Role Variables -This role does not have any mandatory variables. However, either `chrony__ntp_pools` or `chrony__ntp_servers` has to be set, otherwise the host has no time source. +This role does not have any mandatory variables. However, either `chrony__ntp_pools` or `chrony__ntp_servers` has to be set, otherwise the role aborts. ## Optional Role Variables diff --git a/roles/chrony/tasks/main.yml b/roles/chrony/tasks/main.yml index bffb66adb..560ea5e40 100644 --- a/roles/chrony/tasks/main.yml +++ b/roles/chrony/tasks/main.yml @@ -6,6 +6,22 @@ - 'always' +- block: + + # the deployed chrony.conf carries no source of the distribution, so without one from the + # inventory chronyd runs without any time source and the clock drifts unnoticed + - name: 'Check variable constraints' + ansible.builtin.assert: + that: + - '(chrony__ntp_pools | length > 0) or (chrony__ntp_servers | length > 0)' + quiet: true + fail_msg: 'Set chrony__ntp_pools or chrony__ntp_servers, otherwise chronyd has no time source.' + + tags: + # use 'always' so the validation runs even when other roles reference these variables. + - 'always' + + - block: - name: 'Install chrony' diff --git a/roles/duplicity/tasks/main.yml b/roles/duplicity/tasks/main.yml index 7fd2b1212..783da37de 100644 --- a/roles/duplicity/tasks/main.yml +++ b/roles/duplicity/tasks/main.yml @@ -9,6 +9,23 @@ - 'always' +- block: + + # duplicity runs from the venv that python_venv builds from this definition. It falls back + # to an empty list on other platforms, so without this check the role would set up backups + # that have no duplicity to run. + - name: 'Assert that this role supports the platform' + ansible.builtin.assert: + that: + - 'duplicity__python_venv__venvs__dependent_var | length > 0' + quiet: true + fail_msg: 'duplicity: {{ ansible_facts["distribution"] }} {{ ansible_facts["distribution_version"] }} is not supported by this role. Supported platforms: {{ __duplicity__python_venv__venvs__dependent_var.keys() | join(", ") }}.' + + tags: + # use 'always' so the validation runs even when other roles reference these variables. + - 'always' + + - block: - name: 'Install required packages' diff --git a/roles/duplicity/vars/main.yml b/roles/duplicity/vars/main.yml index b1fe61152..1f37a73ca 100644 --- a/roles/duplicity/vars/main.yml +++ b/roles/duplicity/vars/main.yml @@ -67,7 +67,9 @@ __duplicity__python_venv__venvs__dependent_var: python_executable: 'python3.13' exposed_binaries: - 'duplicity' +# `default=[]`: consumers template this variable even where duplicity does not run (see +# "OS-specific Dependent Variables" in CONTRIBUTING.md). The role itself asserts the platform. duplicity__python_venv__venvs__dependent_var: '{{ __duplicity__python_venv__venvs__dependent_var - | linuxfabrik.lfops.platform_select(ansible_facts) + | linuxfabrik.lfops.platform_select(ansible_facts, default=[]) }}' diff --git a/roles/example/vars/main.yml b/roles/example/vars/main.yml index 2cdec6c08..6c012c4d9 100644 --- a/roles/example/vars/main.yml +++ b/roles/example/vars/main.yml @@ -21,8 +21,11 @@ __example__supported_versions: # - role: 'linuxfabrik.lfops.python' # python__modules__dependent_var: '{{ example__python__modules__dependent_var }}' # -# Jinja evaluation is lazy, so the filter only runs when a consumer actually reads the -# variable. Internal (`__`-prefixed) dict first, public variable second. +# A consumer templates this variable even where it does not use it, for example in the +# branch of a `ternary()` that a skipped role leaves untaken. It therefore has to give a +# valid value on every platform, hence `default=[]`. If the role cannot work without the +# value, it asserts its supported platforms itself, as `roles/duplicity` does. +# Internal (`__`-prefixed) dict first, public variable second. __example__python__modules__dependent_var: Debian: - name: 'python-example' @@ -30,5 +33,5 @@ __example__python__modules__dependent_var: - name: 'python3-example' example__python__modules__dependent_var: '{{ __example__python__modules__dependent_var - | linuxfabrik.lfops.platform_select(ansible_facts) + | linuxfabrik.lfops.platform_select(ansible_facts, default=[]) }}' diff --git a/roles/kernel_settings/README.md b/roles/kernel_settings/README.md index 2a683656d..e6f40ed29 100644 --- a/roles/kernel_settings/README.md +++ b/roles/kernel_settings/README.md @@ -10,6 +10,7 @@ The role does nothing on its own and relies on the [linux_system_roles.kernel_se ## How the Role Behaves +* If any `sunrpc.*` setting is configured (the mariadb_server role sets `sunrpc.tcp_slot_table_entries`), the role loads the `sunrpc` kernel module and lists it in `/etc/modules-load.d/sunrpc.conf`, since the `sunrpc.*` settings only exist while the module is loaded and nothing else loads it at boot on a host without NFS. `options sunrpc` lines in `/etc/modprobe.d/` are commented out in the process. The file stays in place when the `sunrpc.*` settings are removed later. * On Ubuntu 22.04 the role removes `kernel.sched_min_granularity_ns` and `kernel.sched_wakeup_granularity_ns` from the TuneD profile it builds on (a `drop` entry in its own profile). The TuneD release of Ubuntu 22.04 sets them although its 5.15 kernel has neither, and TuneD's own verification would otherwise fail on every run. diff --git a/roles/kernel_settings/tasks/main.yml b/roles/kernel_settings/tasks/main.yml index e347b441a..05c5495ac 100644 --- a/roles/kernel_settings/tasks/main.yml +++ b/roles/kernel_settings/tasks/main.yml @@ -18,10 +18,14 @@ - 'Combined transparent_hugepages: {{ kernel_settings__transparent_hugepages__combined_var }}' - 'Combined transparent_hugepages_defrag: {{ kernel_settings__transparent_hugepages_defrag__combined_var }}' - # prevent errors like "Failed to read sysctl parameter 'sunrpc.tcp_slot_table_entries', the parameter does not exist" on RHEL + # prevent errors like "Failed to read sysctl parameter 'sunrpc.tcp_slot_table_entries', the parameter does not exist" on RHEL. + # persistent, because nothing else loads sunrpc on a host without NFS: after a reboot, TuneD could + # not apply the setting and `tuned-adm verify` failed until the next run of this role. + # `persistent` writes /etc/modules-load.d/sunrpc.conf (community.general >= 7.0.0). - name: 'modprobe sunrpc' community.general.modprobe: name: 'sunrpc' + persistent: 'present' state: 'present' when: 'kernel_settings__sysctl__combined_var | selectattr("name", "match", "^sunrpc\.") | list | length > 0' # if there is at least one key starting with "sunrpc." diff --git a/roles/mariadb_server/vars/main.yml b/roles/mariadb_server/vars/main.yml index ff4f7d516..14ab600f9 100644 --- a/roles/mariadb_server/vars/main.yml +++ b/roles/mariadb_server/vars/main.yml @@ -19,5 +19,5 @@ __mariadb_server__python__modules__dependent_var: - name: 'python3-PyMySQL' mariadb_server__python__modules__dependent_var: '{{ __mariadb_server__python__modules__dependent_var - | linuxfabrik.lfops.platform_select(ansible_facts) + | linuxfabrik.lfops.platform_select(ansible_facts, default=[]) }}' diff --git a/roles/monitoring_plugins/README.md b/roles/monitoring_plugins/README.md index c202738aa..5bcfa8f76 100644 --- a/roles/monitoring_plugins/README.md +++ b/roles/monitoring_plugins/README.md @@ -13,7 +13,7 @@ Notes: ## How the Role Behaves -* **Source install builds a virtual environment.** With `monitoring_plugins__install_method: 'source'`, the role deploys the plugins into a self-contained Python virtual environment under `/usr/lib64/linuxfabrik-monitoring-plugins/venv` and rewrites the plugin shebangs to that interpreter, mirroring the layout of the rpm/deb package. Check, event and notification plugins all land flat in `/usr/lib64/nagios/plugins`, and their assets in `/usr/lib64/nagios/plugins/assets`, which is where the Icinga command definitions expect them. +* **Source install builds a virtual environment.** With `monitoring_plugins__install_method: 'source'`, the role deploys the plugins into a self-contained Python virtual environment under `/usr/lib64/linuxfabrik-monitoring-plugins/venv` and rewrites the plugin shebangs to that interpreter, mirroring the layout of the rpm/deb package. Check, event and notification plugins all land flat in `/usr/lib64/nagios/plugins`, and their assets in `/usr/lib64/nagios/plugins/assets`, which is where the Icinga command definitions expect them. The OID lists and MIBs of the `snmp` plugin go into `device-oids/` and `device-mibs/` next to it. Device definitions of your own in these directories are left alone, by updates and by `monitoring_plugins:remove`. * **The role provisions a suitable Python itself.** On RHEL 8 the system Python is 3.6, which is too old. The role installs Python 3.9 (package `python39`) and builds the virtual environment with it, so a source install works on RHEL 8 without any manual Python setup. Every other supported platform already ships Python 3.9 or newer and is used as-is. A virtual environment left behind by an earlier run with a different Python is rebuilt. * **The source install needs no Internet access on the target.** The Ansible controller clones the monitoring-plugins and Linuxfabrik library repositories, downloads every Python dependency as a wheel for the interpreter, architecture and glibc of each target, and copies everything over. pip on the target installs from those files only, so air-gapped hosts are provisioned like any other. * **Dependencies are pinned and verified.** The source install uses the lockfiles of the monitoring-plugins checkout, which pin every package to an exact version and to the checksums of its files, the same set CI tests the plugins against and the rpm/deb packages ship. pip refuses any file whose checksum does not match. A later run moves an existing venv to the versions of the current lockfile. diff --git a/roles/monitoring_plugins/tasks/linux-remove.yml b/roles/monitoring_plugins/tasks/linux-remove.yml index 787ade6e2..557a05db9 100644 --- a/roles/monitoring_plugins/tasks/linux-remove.yml +++ b/roles/monitoring_plugins/tasks/linux-remove.yml @@ -126,6 +126,7 @@ vars: __monitoring_plugins__manifest_allowed: '{{ "^(/usr/lib64/nagios/plugins/[^/]+" + ~ "|/usr/lib64/nagios/plugins/device-(mibs|oids)/.+" ~ "|/usr/lib64/linuxfabrik-monitoring-plugins" ~ "|/etc/sudoers[.]d/linuxfabrik-monitoring-plugins(-logging)?" ~ "|/etc/bash_completion[.]d/linuxfabrik-monitoring-plugins" @@ -134,10 +135,29 @@ loop: '{{ (__monitoring_plugins__manifest["content"] | d("") | b64decode).splitlines() | select("match", __monitoring_plugins__manifest_allowed) - | reject("search", "/[.][.]?$") + | reject("search", "/[.][.]?(/|$)") | list }}' + # The data directories of the snmp plugin go once they are empty, as with the one-line + # installer. Device definitions an admin put there keep them in place. + - name: 'find /usr/lib64/nagios/plugins/{{ item }} -type d -empty -delete -print' + ansible.builtin.command: + argv: + - 'find' + - '/usr/lib64/nagios/plugins/{{ item }}' + - '-type' + - 'd' + - '-empty' + - '-delete' + - '-print' + removes: '/usr/lib64/nagios/plugins/{{ item }}' + register: '__monitoring_plugins__remove_snmp_dirs_result' + changed_when: '__monitoring_plugins__remove_snmp_dirs_result["stdout"] | length > 0' + loop: + - 'device-mibs' + - 'device-oids' + # Hosts set up before the manifest existed, by hand, from a zip or by an earlier version of # this role, carry whatever that release placed. The list holds every name any release ever # installed into the plugin directory. The directory is shared with the plugins of the diff --git a/roles/monitoring_plugins/tasks/linux-source.yml b/roles/monitoring_plugins/tasks/linux-source.yml index d0025eb87..2c2245a0d 100644 --- a/roles/monitoring_plugins/tasks/linux-source.yml +++ b/roles/monitoring_plugins/tasks/linux-source.yml @@ -84,6 +84,13 @@ -not -path '*/example/assets/*' \ -exec cp -- {} /tmp/ansible.monitoring-plugins-repo-flattened/assets/ \; done + # The snmp plugin reads its OID lists and MIBs from directories next to itself, so they + # keep their structure. + for dir in device-mibs device-oids; do + mkdir --parents "/tmp/ansible.monitoring-plugins-repo-flattened/$dir" + [ -d "check-plugins/snmp/$dir" ] || continue + cp --recursive -- "check-plugins/snmp/$dir/." "/tmp/ansible.monitoring-plugins-repo-flattened/$dir/" + done # A checkout always carries plugins. None of them means the clone produced something # unusable, and going on would report a successful run on a host that has no plugins. if [ "$(ls -1 /tmp/ansible.monitoring-plugins-repo-flattened/plugins | wc -l)" -eq 0 ]; then @@ -103,6 +110,20 @@ check_mode: false # run task even if `--check` is specified changed_when: false # no change on the remote host + # Each data file of the snmp plugin goes into the install manifest on its own rather than its + # directory, as the one-line installer does it: an admin puts their own device definitions + # there, and a removal must leave those alone. + - name: 'find device-mibs device-oids -type f' + ansible.builtin.command: + cmd: 'find device-mibs device-oids -type f' + chdir: '/tmp/ansible.monitoring-plugins-repo-flattened' + delegate_to: 'localhost' + become: false + throttle: 1 # serialize: shared dir on the controller, avoid races between hosts + register: '__monitoring_plugins__snmp_device_files' + check_mode: false # run task even if `--check` is specified + changed_when: false # read-only + - name: 'Make sure rsync is installed' ansible.builtin.package: name: 'rsync' @@ -439,7 +460,7 @@ group: 'root' mode: 0o755 - # The rsync options are the same for all three transfers below: + # The rsync options are the same for all transfers below: # # --chmod forces the modes instead of inheriting them from the checkout on the # controller. A hardened controller umask (027 or 077) leaves the git working @@ -480,6 +501,23 @@ - '--no-owner' - '--no-times' + # Without `delete`: these directories also hold the device definitions of the admin. A data + # file the checkout no longer carries is removed below by way of the install manifest. + - name: 'rsync the data files of the snmp plugin to remote host' + ansible.posix.synchronize: + src: '/tmp/ansible.monitoring-plugins-repo-flattened/{{ item }}/' + dest: '/usr/lib64/nagios/plugins/{{ item }}/' + mode: 'push' + rsync_opts: + - '--checksum' + - '--chmod=D755,F644' + - '--no-group' + - '--no-owner' + - '--no-times' + loop: + - 'device-mibs' + - 'device-oids' + # The Linuxfabrik library is deployed straight from its GitHub repo (not via pip), # into the plugin directory. The plugins import it as `lib.*`, which resolves because a # script's own directory is first on sys.path, so `lib/` next to the plugins wins over @@ -557,6 +595,24 @@ | list }}' + - name: 'Remove the data files of the snmp plugin the checkout no longer carries' + ansible.builtin.file: + path: '{{ item }}' + state: 'absent' + vars: + __monitoring_plugins__deployed_snmp_device_paths: '{{ + __monitoring_plugins__snmp_device_files["stdout_lines"] + | map("regex_replace", "^", "/usr/lib64/nagios/plugins/") + | list + }}' + loop: '{{ + (__monitoring_plugins__previous_manifest["content"] | d("") | b64decode).splitlines() + | select("match", "^/usr/lib64/nagios/plugins/device-(mibs|oids)/.+$") + | reject("search", "/[.][.]?(/|$)") + | reject("in", __monitoring_plugins__deployed_snmp_device_paths) + | list + }}' + tags: - 'monitoring_plugins' @@ -627,7 +683,8 @@ "/usr/lib64/linuxfabrik-monitoring-plugins", "/usr/lib64/nagios/plugins/assets", "/usr/lib64/nagios/plugins/lib"] - + (__monitoring_plugins__flattened_plugins["stdout_lines"] + + ((__monitoring_plugins__flattened_plugins["stdout_lines"] + + __monitoring_plugins__snmp_device_files["stdout_lines"]) | map("regex_replace", "^", "/usr/lib64/nagios/plugins/") | list) }}' @@ -829,6 +886,9 @@ {% for name in __monitoring_plugins__flattened_plugins["stdout_lines"] %} /usr/lib64/nagios/plugins/{{ name }} {% endfor %} + {% for name in __monitoring_plugins__snmp_device_files["stdout_lines"] %} + /usr/lib64/nagios/plugins/{{ name }} + {% endfor %} /usr/lib64/nagios/plugins/assets /usr/lib64/nagios/plugins/lib /etc/sudoers.d/linuxfabrik-monitoring-plugins diff --git a/roles/repo_baseos/templates/Rocky8/etc/yum.repos.d/Rocky-Security.repo.j2 b/roles/repo_baseos/templates/Rocky8/etc/yum.repos.d/Rocky-Security.repo.j2 index 0d1102b2a..c4704cc5c 100644 --- a/roles/repo_baseos/templates/Rocky8/etc/yum.repos.d/Rocky-Security.repo.j2 +++ b/roles/repo_baseos/templates/Rocky8/etc/yum.repos.d/Rocky-Security.repo.j2 @@ -1,5 +1,5 @@ # {{ ansible_managed }} -# 2026061601 +# 2026100101 # Rocky-Security.repo # @@ -10,13 +10,17 @@ # # If the mirrorlist does not work for you, you can try the commented out # baseurl line instead. +# +# Unlike the upstream file, the mirrorlist URLs carry no `$rltype`: it is empty on Rocky 8, +# and Rocky 8 releases before 8.5 do not define it, so dnf would pass it on literally and the +# mirrorlist answers 404. Verified with rocky-repos 8.10-1.14 and on Rocky 8.3. [security] name=Rocky Linux $releasever - Security {% if repo_baseos__mirror_url is defined and repo_baseos__mirror_url | length and not (repo_baseos__security_repo_use_upstream | bool) %} baseurl={{ repo_baseos__mirror_url }}/rocky/8/security/x86_64/os/ {% else %} -mirrorlist=https://mirrors.rockylinux.org/mirrorlist?arch=$basearch&repo=security-$releasever$rltype +mirrorlist=https://mirrors.rockylinux.org/mirrorlist?arch=$basearch&repo=security-$releasever {% endif %} #baseurl=http://dl.rockylinux.org/$contentdir/$releasever/security/$basearch/os/ gpgcheck=1 @@ -34,7 +38,7 @@ name=Rocky Linux $releasever - Security Debug {% if repo_baseos__mirror_url is defined and repo_baseos__mirror_url | length and not (repo_baseos__security_repo_use_upstream | bool) %} baseurl={{ repo_baseos__mirror_url }}/rocky/8/security/x86_64/debug/tree/ {% else %} -mirrorlist=https://mirrors.rockylinux.org/mirrorlist?arch=$basearch&repo=security-$releasever-debug$rltype +mirrorlist=https://mirrors.rockylinux.org/mirrorlist?arch=$basearch&repo=security-$releasever-debug {% endif %} #baseurl=http://dl.rockylinux.org/$contentdir/$releasever/security/$basearch/debug/tree/ gpgcheck=1 @@ -51,7 +55,7 @@ name=Rocky Linux $releasever - Security Source {% if repo_baseos__mirror_url is defined and repo_baseos__mirror_url | length and not (repo_baseos__security_repo_use_upstream | bool) %} baseurl={{ repo_baseos__mirror_url }}/rocky/8/security/source/tree/ {% else %} -mirrorlist=https://mirrors.rockylinux.org/mirrorlist?arch=$basearch&repo=security-$releasever-source$rltype +mirrorlist=https://mirrors.rockylinux.org/mirrorlist?arch=$basearch&repo=security-$releasever-source {% endif %} #baseurl=http://dl.rockylinux.org/$contentdir/$releasever/security/source/tree/ gpgcheck=1 diff --git a/tests/unit/plugins/lookup/test_bitwarden_item.py b/tests/unit/plugins/lookup/test_bitwarden_item.py index d8b1ccb27..bd0d57394 100644 --- a/tests/unit/plugins/lookup/test_bitwarden_item.py +++ b/tests/unit/plugins/lookup/test_bitwarden_item.py @@ -23,6 +23,7 @@ __metaclass__ = type import contextlib +import multiprocessing import os import unittest from typing import ClassVar @@ -45,8 +46,13 @@ class _FakeBitwarden: """Minimal stand-in for the Bitwarden client used by the lookup.""" items_by_search: ClassVar[list] = [] + # what a forced sync brings into the cache, None for no change + items_after_forced_sync = None item_by_id = None created_items: ClassVar[list] = [] + sync_calls: ClassVar[list] = [] + # what sync() reports, True for a sync that actually ran + sync_result = False vault_status = 'unlocked' mutex_held = False @@ -69,8 +75,11 @@ def status(self): def get_not_unlocked_message(self, status): return f'vault reports status "{status}"' - def sync(self, *args, **kwargs): - pass + def sync(self, force=False, run_id=None, **kwargs): + type(self).sync_calls.append({'force': force, 'run_id': run_id}) + if force and type(self).items_after_forced_sync is not None: + type(self).items_by_search = type(self).items_after_forced_sync + return force or type(self).sync_result def get_items( self, @@ -115,8 +124,11 @@ def setUp(self): self._orig = lookup_mod.Bitwarden lookup_mod.Bitwarden = _FakeBitwarden _FakeBitwarden.items_by_search = [] + _FakeBitwarden.items_after_forced_sync = None _FakeBitwarden.item_by_id = None _FakeBitwarden.created_items = [] + _FakeBitwarden.sync_calls = [] + _FakeBitwarden.sync_result = False _FakeBitwarden.mutex_held = False _FakeBitwarden.vault_status = 'unlocked' # a value leaking in from the caller's environment would flip the @@ -209,6 +221,59 @@ def test_lookup_by_id_lifts_credentials(self): self.assertEqual(result[0]['username'], 'dba') self.assertEqual(result[0]['password'], 'linuxfabrik') + def test_missing_item_is_searched_again_after_a_sync_before_it_is_created(self): + # the cache predates an item created elsewhere during this run + _FakeBitwarden.items_after_forced_sync = [ + { + 'name': 'host - db', + 'login': {'username': 'dba', 'password': 'linuxfabrik'}, + }, + ] + result = self.lookup.run([{'name': 'host - db', 'username': 'dba'}]) + self.assertEqual(_FakeBitwarden.created_items, []) + self.assertEqual(result[0]['password'], 'linuxfabrik') + self.assertEqual([c['force'] for c in _FakeBitwarden.sync_calls], [False, True]) + + def test_no_second_sync_if_the_lookup_has_just_synced(self): + _FakeBitwarden.sync_result = True + self.lookup.run([{'name': 'host - db', 'username': 'dba'}]) + self.assertEqual(len(_FakeBitwarden.created_items), 1) + self.assertEqual(len(_FakeBitwarden.sync_calls), 1) + + def test_sync_gets_the_run_id(self): + self.lookup.run([{'name': 'host - db', 'username': 'dba'}]) + self.assertEqual( + _FakeBitwarden.sync_calls[0]['run_id'], lookup_mod.get_run_id() + ) + + +def _put_run_id(queue): + queue.put(lookup_mod.get_run_id()) + + +class TestGetRunId(unittest.TestCase): + def test_forked_worker_gets_the_run_id_of_its_parent(self): + # Ansible forks its workers without exec, like multiprocessing's fork context + run_id = lookup_mod.get_run_id() + self.assertIsNotNone(run_id) + ctx = multiprocessing.get_context('fork') + queue = ctx.Queue() + worker = ctx.Process(target=_put_run_id, args=(queue,)) + worker.start() + worker.join() + self.assertEqual(queue.get(timeout=5), run_id) + + def test_no_proc_gives_no_run_id(self): + def _missing(pid, name): + raise FileNotFoundError(name) + + orig = lookup_mod._read_proc + lookup_mod._read_proc = _missing + try: + self.assertIsNone(lookup_mod.get_run_id()) + finally: + lookup_mod._read_proc = orig + if __name__ == '__main__': unittest.main() diff --git a/tests/unit/plugins/module_utils/test_bitwarden.py b/tests/unit/plugins/module_utils/test_bitwarden.py index ec18e68b0..1fa175d3f 100644 --- a/tests/unit/plugins/module_utils/test_bitwarden.py +++ b/tests/unit/plugins/module_utils/test_bitwarden.py @@ -360,6 +360,98 @@ def test_unknown_item_by_id_still_raises(self): self.assertEqual(len(calls), 3) +_SYNC_BAD_REQUEST = HTTPError( + 'http://127.0.0.1:8087/sync', 400, 'Bad Request', {}, None +) + + +class TestSyncIsRetried(unittest.TestCase): + """bw serve occasionally fails a sync with HTTP 400 or a timeout, then recovers.""" + + def setUp(self): + self._tmp = tempfile.TemporaryDirectory() + self.bw = _make_bitwarden(os.path.join(self._tmp.name, 'cache.json')) + self._orig_open_url = bitwarden.open_url + self._sleep = unittest.mock.patch.object(bitwarden.time, 'sleep') + self.sleep = self._sleep.start() + + def tearDown(self): + self._sleep.stop() + bitwarden.open_url = self._orig_open_url + self._tmp.cleanup() + + def test_bad_request_is_retried(self): + bitwarden.open_url, calls = _serve_in_order( + _SYNC_BAD_REQUEST, _SYNCED, _items(_LOGIN_ITEM) + ) + self.bw.sync(force=True) + self.assertEqual(len(calls), 3) + self.assertEqual(self.sleep.call_args_list, [unittest.mock.call(10)]) + self.assertEqual(self.bw.get_items('host - db', username='dba'), [_LOGIN_ITEM]) + + def test_timeout_is_retried(self): + bitwarden.open_url, calls = _serve_in_order( + _SYNCED, TimeoutError('timed out'), _SYNCED, _items(_LOGIN_ITEM) + ) + self.bw.sync(force=True) + self.assertEqual(len(calls), 4) + self.assertEqual(self.sleep.call_args_list, [unittest.mock.call(10)]) + + def test_last_error_aborts(self): + bitwarden.open_url, calls = _serve_in_order(*[_SYNC_BAD_REQUEST] * 4) + with self.assertRaises(bitwarden.BitwardenException) as ctx: + self.bw.sync(force=True) + self.assertEqual(len(calls), 4) + self.assertEqual([c[0][0] for c in self.sleep.call_args_list], [10, 30, 60]) + self.assertIn('400', str(ctx.exception)) + # a failed sync leaves the cache as it was + self.assertEqual(self.bw._cache['sync_timestamp'], 0) + + +class TestSyncOncePerRun(unittest.TestCase): + """Listing all items is slow, so the lookup syncs once per Ansible run.""" + + def setUp(self): + self._tmp = tempfile.TemporaryDirectory() + self.bw = _make_bitwarden(os.path.join(self._tmp.name, 'cache.json')) + self._orig_open_url = bitwarden.open_url + + def tearDown(self): + bitwarden.open_url = self._orig_open_url + self._tmp.cleanup() + + def test_second_sync_of_a_run_is_skipped(self): + bitwarden.open_url, calls = _serve_in_order(_SYNCED, _items(_LOGIN_ITEM)) + self.assertTrue(self.bw.sync(run_id='4242:1')) + # well past the interval, which does not apply with a run ID + self.bw._cache['sync_timestamp'] -= 3600 + self.assertFalse(self.bw.sync(run_id='4242:1')) + self.assertEqual(len(calls), 2) + + def test_next_run_syncs_again(self): + bitwarden.open_url, calls = _serve_in_order( + _SYNCED, _items(_LOGIN_ITEM), _SYNCED, _items(_LOGIN_ITEM) + ) + self.bw.sync(run_id='4242:1') + self.assertTrue(self.bw.sync(run_id='4343:2')) + self.assertEqual(len(calls), 4) + self.assertEqual(self.bw._cache['sync_run_id'], '4343:2') + + def test_without_run_id_the_interval_applies(self): + bitwarden.open_url, calls = _serve_in_order(_SYNCED, _items(_LOGIN_ITEM)) + self.bw.sync() + self.assertFalse(self.bw.sync()) + self.assertEqual(len(calls), 2) + + def test_force_syncs_within_a_run(self): + bitwarden.open_url, calls = _serve_in_order( + _SYNCED, _items(_LOGIN_ITEM), _SYNCED, _items(_LOGIN_ITEM) + ) + self.bw.sync(run_id='4242:1') + self.assertTrue(self.bw.sync(force=True, run_id='4242:1')) + self.assertEqual(len(calls), 4) + + def _create_if_missing(cache_file, created): """Worker for TestMutex: the lookup's search-then-create, run in its own process.""" bitwarden.CACHE_FILE = cache_file