actions/cache -- and actions/setup-go, which builds on it -- packs its archive
with zstdmt when zstd is present and falls back to single-threaded gzip when
it is not. The archive name says which happened: cache.tzst against cache.tgz.
Measured in l.kirchner/patchmgr, CI run 1957. The cache is GOMODCACHE plus
GOCACHE and runs to 2-5 GB. runner-gpu has zstd and writes cache.tzst;
runner-01, -02 and -03 write cache.tgz, and go test -race (stable) spent 605
seconds packing it for a job with 41 seconds of work. Across all eight jobs of
that run, 2312 of 3146 seconds went into this step.
This does not settle whether the cache is wanted at all -- in host mode with a
persistent home both directories survive between jobs anyway, and patchmgr is
switching it off for that reason. But as long as any repository on the
instance uses it, it should not be compressed on one core.
zstd is a package for the post step, not for jobs, so it sits with the others
rather than in a workflow: this runner installs nothing at job time, by
design.
The script that sets the SSH, PAM, pwquality and auditd baseline on twelve
containers existed only as a root-owned copy on the machines it hardens.
Bring it into the repo so a change reaches one place instead of twelve.
Two substantive changes over the copy that shipped on 2026-07-24:
- Deny forwarding per option instead of via DisableForwarding. Same effect,
but DisableForwarding overrides every other forwarding option and is
invisible in sshd -T, which makes a rejected port-forward read as a
configuration that should work.
- Quote the command substitution in MODDIR (SC2046).
The header and README now record the rollout command, the containers left
unhardened as break-glass foundation, and why host-specific exceptions must
live in a drop-in that sorts after 99-cis-hardening.conf.
Both from the cross-review, both fair.
"Go setzt CGO_ENABLED=0" is right about the effect and wrong about the
mechanism: Go does not set the variable, cgo simply stays off when no C
compiler is found, and go env then reports 0. Reworded.
And protoc on an instance-wide runner ties every repository to the
distribution's version -- 3.21.x on Debian 12. Unlike make and gcc that is a
code generator, so a distro upgrade changes generated code for all users at
once. The comment says so now, and says where a project that needs its own
version should pin it instead of raising it here for everybody.
The runner is host-mode, so there is no image bringing tools along: what is
not on this LXC, no job has. Measured on l.kirchner/patchmgr, a Go project,
where all six CI jobs were assigned and every one of them died in the first
seconds:
make all make: command not found
go test -race go: -race requires cgo; enable cgo by setting CGO_ENABLED=1
make proto sudo: command not found
Four packages, each for a reason:
make the gate commands are make targets
gcc Go turns CGO_ENABLED off when it finds no C compiler,
and the race detector cannot be built without cgo
protobuf-compiler protoc itself
libprotobuf-dev the well-known .proto includes under
/usr/include/google/protobuf; without them protoc fails
even though the binary is there
sudo stays absent on purpose. A workflow must not be able to install anything
on this runner -- what is needed is declared here, in the script, and not in
somebody's pipeline. That also keeps the security note at the top of this file
honest: the LXC owns nothing, and it gains nothing at a workflow's request.
Go is not in the list. Projects fetch it through actions/setup-go, because CI
matrices run more than one version.
Applied to the running LXC (301 on pve-gamer) while the runner was idle, then
verified: make 4.3, gcc 12.2.0, libprotoc 3.21.12, 11 .proto includes present.
The service PATH already contains /usr/bin, so no restart was needed.
- assert encoding BEFORE role/password mutation on re-run, so an old
SQL_ASCII DB aborts with no side effects (codex finding 1)
- ensure_utf8_locale_active honours its locale argument consistently in
match, locale.gen line and export (codex finding 2)
- assert_db_encoding_utf8 uses argv-clean runuser psql with :'db' literal
binding instead of nested su -c shell; docs keep su - postgres -c
(codex finding 3)
A PostgreSQL cluster/database freezes its encoding at initdb / CREATE
DATABASE time; a C (non-UTF-8) locale yields a SQL_ASCII cluster, which
makes psycopg3 return bytes and crashes SQLAlchemy. Harden the installer
and add a reusable pattern for future DB installers:
- ensure_utf8_locale_active: generate AND activate en_US.UTF-8 for the
install process before the server package runs initdb; abort if the
locale is not actually available
- create the database explicitly with TEMPLATE template0 ENCODING 'UTF8'
LC_COLLATE/LC_CTYPE 'en_US.UTF-8' instead of inheriting the cluster
default
- assert_db_encoding_utf8: post-install guard, abort with an actionable
message if pg_encoding_to_char is not UTF8 (catches old SQL_ASCII DBs
on re-run too)
- credentials/README docs use su - postgres -c (minimal LXCs have no sudo)
- merge the two contributing sections into one (PR + cross-review rule,
both real incidents, CI enforcement in present tense - the suite is
live on the homelab runner since K-114/PR #5)
- script catalog: runner is in production (PR #6 merged)
- usage: document input validation behavior (re-prompt on junk bytes,
env values abort when malformed)
- pattern: SSH root login prompt (sshd drop-in) and the locale fix in
setup_base_apt
- repo layout: tests/ added
LXC templates ship without a configured locale, so every apt/perl run
warned 'Setting locale failed'. setup_base_apt now exports C.UTF-8 for
the install run itself, installs the locales package, generates
en_US.UTF-8 and sets it as the system default via update-locale.
prompt_lxc_config asks 'SSH-Root-Login erlauben? [Y/n]' (env-presettable
via SSH_ROOT_LOGIN, validated, normalized to yes|no). The bootstrap passes
the value into the container; configure_ssh_root_login writes
/etc/ssh/sshd_config.d/zz-root-login.conf (yes -> PermitRootLogin yes,
no -> prohibit-password) and reloads sshd.
- CTID prompt re-prompts on invalid interactive input (was: abort)
- env-provided NAMESERVER is validated when non-empty ('' stays inherit)
- prompt_validated handles EOF (no infinite loop, clean abort under -e)
- 10# base forcing in vlan/cidr/ipv4 arithmetic (leading zeros are not
octal errors); is_clean_ascii rejects embedded newline/tab explicitly
(command substitution strips trailing newlines); is_ipv4_list checks
the whole string before word splitting
- nameref guard against reserved variable names in prompt_validated/
require_valid; source-check pattern documented as the repo contract
- 8 new test cases (41 total)
- validation helpers: sanitize_input trims CR/edge whitespace only;
embedded control/non-ASCII bytes FAIL validation and re-prompt with a
hint (2026-06-11 incident: invisible byte in a pasted VLAN tag broke
pct create mid-run) - never silently stripped
- prompt_lxc_config: every prompt validated (uint for CTID/disk/cores/
RAM, VLAN 1-4094, hostname/token formats, IP/CIDR/gateway, DNS list);
env-provided values are sanitized + validated too (abort, no re-prompt
loop in non-interactive use); helpers reusable for app prompts
- tests/test_validation.sh: 34 cases incl. the 2<0x80>0 repro, re-prompt
simulation, BASH_REMATCH clobbering regression (is_cidr), env dry-run
of prompt_lxc_config without PVE/TTY
- tests/check_ct_source.sh: every ct/*.sh must source build.func (bug
shipped twice); negative proof via prepared fixture in the test suite
- .gitea/workflows/ci.yml: bash -n over all scripts, source-check,
validation tests, shellcheck if present (documented skip otherwise)
- README: contributions via PR with cross-review (binding)
- source build.func (script was unrunnable without it)
- validate DB_NAME/DB_USER/DB_PORT/NEXUS_APP_IP before SQL/pg_hba use
- rotate password when role exists but credentials file is missing
webapp-pattern ct/install pair: PGDG repo, database nexus with
least-privilege owner role, pg_hba allowlist restricted to the nexus
app LXC (explicit reject for everything else), pgvector created by the
installer, credentials/DSN summary in /root/nexus-db.credentials.
Idempotent re-runs keep role/db and do not rotate the password.
A reboot before the first deploy would leave the enabled unit in failed
state; the condition keeps it inert until start.sh exists (same guard as
nexus-worker.service).
- new install/nexus-runtime.sh (idempotent, re-runnable on an existing
LXC): uv for the nexus user (manages Python 3.12), tesseract deu+eng,
nexus-worker.service unit (ConditionPathExists guards the skeleton
phase), sudoers extended to cover the worker service
- nexus-install.sh: RUNTIME section now invokes nexus-runtime.sh at the
end of the install (after base sudoers/units, which it extends)
Adds apply_network_profile(), which looks up DNS servers for the entered
VLAN tag in lib/networks.conf and sets --nameserver accordingly — even
when IP is DHCP. Precedence: explicit env NAMESERVER > profile > static
prompt > inherit. Comma-separated DNS is normalised to spaces for pct.
Data file mapping a VLAN tag to its DNS servers and subnet, so build.func
can set the right resolvers from the tag entered at install time — even
with DHCP. New networks are a one-line addition here.
Installs Node.js and act_runner in host mode, registers the runner as
the unprivileged webapp user, and wires up systemd units for the runner
and the next-start service. A narrow sudoers rule lets the runner restart
only webapp.service; the build/deploy itself is driven by the repo workflow.
Host-side script that creates an unprivileged Debian 12 LXC, installs
Node.js + act_runner in host mode, and registers it against the Gitea
instance. Deploy logic lives in the repo's .gitea/workflows/deploy.yml
(deploy-as-code); the runner polls outbound, so no inbound port.
For DHCP setups DNS is delivered with the lease, but with a static IP the
container inherits /etc/resolv.conf from the PVE host - which is often
unreachable from the container's VLAN.
Changes:
- New NAMESERVER prompt (only shown for static IPs, defaults to gateway)
- pct create now passes --nameserver when set
- Network wait loop tests L3 and DNS separately so failures point at the
actual cause (no route to gateway vs. bad DNS server)
- Refactored pct create args into an array for cleaner conditional flags
HOSTNAME is a bash built-in always containing the host's name, so the
"-z HOSTNAME" check never fired and the prompt was silently skipped —
containers ended up named after the Proxmox host.
Also added an optional VLAN tag prompt (empty = no tag), and the network
wait loop now exits with an error if the network never comes up instead
of silently proceeding to a guaranteed-broken apt-get update.