Skip to content

fix: findings from the setup_basic lab rounds - #412

Merged
NavidSassan merged 16 commits into
mainfrom
fix/setup-basic-lab-findings
Oct 2, 2026
Merged

NavidSassan merged 16 commits into
mainfrom
fix/setup-basic-lab-findings

Conversation

@markuslf

@markuslf markuslf commented Oct 2, 2026

Copy link
Copy Markdown
Member

Fixes found while deploying setup_basic and setup_icinga2_master to lab VMs and to icinga-demo.

Changes

  • role:monitoring_plugins: the source install deploys the OID lists and MIBs of the snmp plugin.
  • role:chrony: abort if neither chrony__ntp_pools nor chrony__ntp_servers is set.
  • role:repo_baseos: the Rocky security repository works on Rocky 8 before 8.5.
  • role:kernel_settings: modprobe sunrpc is persistent, so sunrpc.* sysctls survive a reboot and tuned-adm verify no longer fails after one.
  • playbook:setup_basic: the venvs of skipped duplicity and glances are neither built nor templated.
  • role:aide: before aide --init, also wait for dnf-automatic jobs and the system_update jobs.
  • plugin:bitwarden_item: a failed vault sync is retried after 10, 30 and 60 s; the lookup syncs once per Ansible run instead of every 60 s, and syncs again before creating a missing item.

Tests

  • Unit tests (tests/unit): 357 passed, 12 of them new for the Bitwarden changes. pre-commit (yamllint, ruff, bandit, vulture, pytest) green.
  • Bitwarden run detection verified with ansible-core 2.16, strategies linear and mitogen_linear, 4 hosts: one ID for all workers of a run, a new one per run.
  • dnf-automatic unit names verified in Rocky 8, 9 and 10 containers.
  • setup_basic lazy venv injection verified with ansible-core 2.16 and 2.18 against Fedora 44.
  • The Molecule scenarios kernel_settings/sysctl and setup_basic were adapted but not run.

…nmp plugin

The source install copied only the plugins and their assets, so `snmp`
failed on every such host with "No such file or directory" for
device-oids/any-any-any.csv. Deploy device-oids/ and device-mibs/ like
build/install-plugins.sh and the one-line installer do, list each file in
the install manifest, and let monitoring_plugins:remove delete only those
files and the directories once empty, so device definitions of the admin
survive updates and removal.
The deployed chrony.conf carries none of the distribution's sources, so
without chrony__ntp_pools or chrony__ntp_servers chronyd ran without any
time source and the clock drifted unnoticed.
Rocky 8 releases before 8.5 do not define the $rltype dnf variable, so the
mirrorlist URL carried it literally and answered 404. $rltype is empty on
every Rocky 8 release, so drop it. Verified on Rocky 8.3 and with
rocky-repos 8.10-1.14.
…e a reboot

The role loaded the sunrpc module only at runtime. On a host without NFS
nothing loads it again after a reboot, so the sunrpc.* sysctls do not
exist, TuneD cannot apply sunrpc.tcp_slot_table_entries (set by the
mariadb_server role) and the next run fails in `tuned-adm verify`.
`persistent: present` writes /etc/modules-load.d/sunrpc.conf
(community.general >= 7.0.0, verified against 8.6.11).

The molecule scenario now sets a sunrpc sysctl and checks it after a
reboot.
…ty and glances

python_venv received the venvs of duplicity and glances regardless of
setup_basic__skip_duplicity and setup_basic__skip_glances. On a host
with an old duplicity venv, the pip install then failed in dependency
backtracking and aborted the run although duplicity is skipped there.
Gate both injections like the other dependent vars in the playbook.
… aide --init

The wait before `aide --init` only covered apt-daily(-upgrade).service,
which do not exist on the Red Hat family. Wait for the package jobs of
the platform (dnf-automatic.service and dnf-automatic-install.service,
verified against dnf-automatic 4.7, 4.14 and 4.20 on Rocky 8, 9 and 10)
and for security-update.service and update-and-reboot.service of the
system_update role on every platform.
…once per run

`bw serve` occasionally answers POST /sync with HTTP 400 or runs into
the 60s timeout, which aborted the whole run. sync() now tries again
after 10, 30 and 60 seconds; MUTEX_TIMEOUT grows to 600s so waiting
workers outlast the retries.

Each sync also lists the whole vault, which takes `bw serve` about 20s,
and the lookup synced every 60s. It now syncs once per Ansible run,
identified by the PID and start time of the forking ansible-playbook
process (verified with ansible-core 2.16, linear and mitogen_linear).
Without /proc and in the module, the 60s interval still applies. Before
creating a missing item, the lookup syncs again, so an item created
elsewhere during a long run is not created twice.
ternary() evaluates both of its arguments, so the venv definition of a
skipped duplicity or glances was still templated, and its platform_select
aborted on a platform the role does not support (found on Fedora 44). An
inline if/else does not help either, since Jinja resolves every name of
the expression up front. lookup('ansible.builtin.vars') only runs in the
branch taken. Verified with ansible-core 2.16 and 2.18 on Fedora 44.
mock's call.args only exists since Python 3.8. On the RHEL 8
platform-python it resolves to an attribute lookup, so the sleep
durations compared as ['args', 'args', 'args']. Index the call tuple
instead.
@markuslf
markuslf force-pushed the fix/setup-basic-lab-findings branch from e7ba87d to bdff065 Compare October 2, 2026 08:09
@NavidSassan
NavidSassan added this pull request to stack #415 October 2, 2026 08:20
@NavidSassan
NavidSassan merged commit 35aa6f7 into main Oct 2, 2026
13 checks passed
@NavidSassan
NavidSassan deleted the fix/setup-basic-lab-findings branch October 2, 2026 12:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants