Commit Graph
150 Commits
Author SHA1 Message Date
Joachim Wiberg a53b7dd1ea board/common: basic support for smbios/dmi based systems in probe
This allows us to gather better product information in /run/system.json
also for x86 based systems.  For Qemu x86_64 systems the 'product-name'
changes from 'VM' to 'Standard PC (i440FX + PIIX, 1996)'.

Also, add new 'product-version' field for, e.g., board revisions, and add
missing 'serial-number' value for systems like Rasperry Pi that inject it
in the device tree.

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-11-20 09:34:00 +01:00
Joachim Wiberg 3aa22a7c63 board/common: only apply explicitly requested DHCP options in client
Validate that DHCP options were requested in the parameter request
list before applying them. This prevents malicious or misconfigured
DHCP servers from forcing unwanted configuration changes.

Validates: hostname (12), DNS (6), domain (15), search (119),
router (3), static routes (121), and NTP (42).

Fail-safe behavior: rejects options if config file unavailable.

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-11-20 09:34:00 +01:00
Joachim Wiberg 90f619bfa6 confd: add deterministic hostname management via /etc/hostname.d/
Resolve race condition between DHCP and configured hostname by introducing
priority-based hostname management using /etc/hostname.d/ directory pattern.

Priority order (highest wins):
  90-dhcp-<iface>  - DHCP assigned hostname
  50-configured    - YANG /system/hostname config
  10-default       - Bootstrap/factory default

The new /usr/libexec/infix/hostname helper reads all sources and applies the
highest priority hostname.  It exits early if hostname unchanged, preventing
unnecessary service restarts.

Fixes #1112

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-11-18 19:07:25 +01:00
Joachim Wiberg 548e0c4881 board/common: don't change link status of interface
After some discussion we agreed that this operation is the domain of
confd (or dagger) and any DHCP client should not mess with the admin
state of the interface it runs on.

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-11-18 18:11:35 +01:00
Joachim Wiberg a568f31309 confd: add support for dhcpv6 client
Basic DHCPv6 support as a per-interface ipv6 setting, not in line with
the IETF YANG model, which has this as a root container.

Note, this also fixes DHCPv4 option inference after the relocation of
/dhcp-client to /interfaces, in 764bd8e.

Fixes #1110

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-11-18 13:28:15 +01:00
Tobias WaldekranzandGitHub 20673c3a84 Merge pull request #1236 from kernelkit/recopy
bin: copy: Fix various resource leaks introduced by refactor
2025-11-06 09:26:00 +01:00
Tobias Waldekranz e718644112 board/common: mksshkey: Clean up stray temporary file 2025-11-06 08:03:55 +01:00
Joachim Wiberg b38e13a08a board/common: check if partition/LABEL is available
Check if partition/LABEL is available before trying to run tune2fs and
mount commands on it.  This change masks the kernel output:

    LABEL=var: Can't lookup blockdev

commne when running `make run` on x86_64 builds.

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-11-05 20:41:57 +01:00
Joachim Wiberg 04c86478c3 board/common: fix boot warnings, follow-up to 45efa94
In 45efa94 we introduced new mount options for the: 'aux', 'cfg', and
'var' partitions.  To save boot time, we also disabled the periodic fsck
check, which does not affect journal replay or other Ext4 safety or
integrity checks.  The latter involved calling tune2fs at runtime, but
this caused the following to appear for the 'aux' partition, which is a
reduced ext4 for systems that need to change U-Boot settings, e.g., to
set MAC address on systems without a VPD.

   EXT4-fs: Mount option(s) incompatible with ext2

Note, this is a "bogus" warning, setting the fstypt to ext2 in fstab is
not the way around this one.  That's been tried.

To make matters worse, the new mount option(s), e.g., 'commit=30', gave
us another set of warnings from the kernel.  This was due to the fstype
being 'auto', so the kernel objected to options not being applicable to
every filesystem it tried before ending up with extfs:

    squashfs: Unknown parameter 'commit'
    vfat: Unknown parameter 'commit'
    exfat: Unknown parameter 'commit'
    fuseblk: Unknown parameter 'commit'
    btrfs: Unknown parameter 'errors'

This commit explicitly sets cfg and var to ext4, leaving mnt as auto, we
also check before calling tune2fs that the partition/disk image is ext4.

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-11-05 20:41:57 +01:00
Joachim Wiberg 9c220d8a51 board/common: utilize new finit-style logging for factory reset
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-11-05 20:41:56 +01:00
Joachim Wiberg 45efa9484e board/common: speed up mounting and boot
Reduce /var image size to 128M (resized at first boot anyway) and tune
mke2fs options for faster mounting.  Set ext4 'uninit_bg' feature and
use '-m 0 -i 4096' extraargs in genimage to optimize filesystem creation
and reduce overhead.  The genimage tool has several tweaks enabled by
default already, the 'uninit_bg' feature speeds up the time to check
the file system.

Disable periodic fsck with tune2fs at boot while keeping safety checks
intact.  Adjust mount options in fstab to reduce journal syncs and
improve boot time.

Fix severe performance regression in find_partition_by_label() that was
calling sgdisk on every block device including virtual devices (ram, loop,
dm-mapper). This caused boot delays of up to 24 seconds on systems with
many block devices. Now skip virtual devices that don't have GPT tables,
reducing the delay to ~2 seconds.

Also clean up is_mmc() function to use cached result.

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-11-01 23:05:19 +01:00
Joachim Wiberg b9e53e5721 board/common: minor fixes to new boot status print
Also, these comments got lost, helps a bit understanding what's going on.

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-11-01 23:05:18 +01:00
Joachim Wiberg 4b55e38741 board/aarch64: use %m modifier in default xPi hostnames
The xPi's usually don't have a VPD so the chassis mac-address probed at
boot is usually null in /run/system.json.  This commit adds a fallbkack
mechanism to populate this field so it can be used for unique hostnames
even on these boards.

Ths ietf-hardware.yang model does not have a notion of physical address,
so we augment one tht is generic enought to be used for other hardware
components than Ethernet, similar to what ietf-interfaces.yang use.

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-10-31 13:26:08 +01:00
Joachim Wiberg e39b9cfee4 board/common: allow resizing /var on any mmc block device
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-10-31 13:26:01 +01:00
Joachim Wiberg db06ddb93b board/common: move resize2fs of /var to after reboot
On some boards, and in particular with a hybrid mbr/gpt partition table,
like on the RPi64, we must resize ext *after* reboot.

Also, do some cleanup and consolidation of error handling to prevent us
from entering an endless boot loop.

Follow-up to 391e9715

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-10-31 13:26:01 +01:00
Joachim Wiberg 6aaf612095 board/common: simplify USB port discovery and fix duplicate entries
This commit refactors USB port probing to:

- Eliminate duplicates: previously, 'authorized' and 'authorized_default'
  were listed as separate USB port entries (confusing).  Now each USB
  port is represented once, with the path pointing to the USB device
  directory, confd appends the appropriate attribute file as needed

- Add support for Raspberry Pi 4B and CM4 USB port(s) using a generic
  discovery function that scans /sys/bus/usb/devices for USB root hubs.
  This should work seamlessly across all platforms

- For backwards compatibility and better UX:
   - Single USB port systems: Named "USB" (no number)
   - Multi-port systems: Named "USB1", "USB2", etc.

- Device tree-based discovery is tried first (for boards like Alder with
  explicit DT USB port definitions), with fallback to generic discovery
  for boards without DT

Fixes: #315

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-10-31 13:26:00 +01:00
Joachim Wiberg ea2c4be237 container: refactor and cleanup per review comments
Shell script:

 - Factor out big portions of code into more logical helper functions
 - Simplify calling setup script by checking for remote image first
 - Simplify meta/sha up-to-date handling and clarify terminology
 - Consistent use of -f instead of -e in file-exists checks
 - Fix unsafe use of 'mktemp -u'

C code:
 - Clarify meta/sha terminology: rename meta-sha256 -> meta-image-sha256
 - Refactor weird archive_offset() function to local_path() helper
 - Factor out helper function calc_sha()
 - Check len of sha256 >= 64

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-10-23 15:23:55 +02:00
Joachim Wiberg da5b4bcfff container: remove old image when upgrading at startup
One unique feature of Infix OS is that you can embed your OCI archive in
the rootfs.squashfs.  This way your container instance does not need to
download anything from the network.  To upgrade you drop in a new image
and rebuild Infix, when the new Infix boots the container script 'setup'
command will recognize that the OCI archive has changed and will reload
it into the container store.

This patch is an improvement of the way too generic 'podman image prune'
command used previously.  Instead of looking for any dangling image, we
now surgically remove the old image.

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-10-23 15:23:54 +02:00
Joachim Wiberg f53d26bf34 container: upgrade fixes for mutable images
Container instances that run with mutable images, e.g., tagged with `:latest`
or similar non-versioned tags, can be upgraded without changing the config.

This commit fixes two issues found with this support:

- force container image re-fetch on upgrade, even if the file exists locally
- surgically remove old image from container store after upgrade

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-10-23 15:23:54 +02:00
Joachim Wiberg 7fbc2eba4f container: make 'container remove' cli command slightly more useful
Usually, when your system is up and running properly, you want to clean
up anything unused from your previous experiments.  This change alllows
that by calling the interactive 'podman image prune -a -f' command from
the CLI command 'container remove all'

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-10-23 15:23:54 +02:00
Joachim Wiberg 5b7325261a container: use 'nice' for podman load and create
When loading big OCI images at boot the `podman load` process completely
monopolizes all cores of an Arm Cortex-A72.  It blocks on I/O, sure, but
with 'nice' we can get some attention at least to more critical services
at boot.

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-10-23 15:23:53 +02:00
Joachim Wiberg 0eddd1ba64 container: refactor cleanup on instance removal
This commit reverts 477f7ae and bb19d06, which intended to fix an issue
with lingering old images, see #1098.  However, as detailed in #1147,
this caused severe side effects while working with multiple larger
containers.  Basically, the prune operation of one container removed
images of other containers that are just being created in parallel.

Instead of using the podman prune command we can use the meta datain the
start script to pinpoint exactly which image(s) to remove, including any
downloaded OCI archives when the container instance is removed.

Fixes #1147

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-10-23 15:23:53 +02:00
Joachim Wiberg 68bb01545c container: optimize startup of preexisting containers
This commit adds metadata to track loaded OCI archives to allow skipping
'delete + load' of OCI images when restarting either the container or the
system as a whole.  The sha256 of all loaded OCI archives is stored in a
sidecar file in our downloads directory.  Then we verify the checksum of
the OCI archives against their same-named sidecar to determine if the OCI
archive is already loaded or not.

Additionally, the instance using the image is labled with metadata to detect
changes in the container configuration.  This in turn allow skipping the
delete + create phase also of the instance.

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-10-23 15:23:52 +02:00
Joachim Wiberg c6f3e75aa4 container: only retry remote images on network changes
Fixes #1148

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-10-23 15:23:52 +02:00
Joachim Wiberg af7c2d7801 container: increase podman stop timeout
The default timeout for 'podman stop foo' is 10 seconds, which for
heavily loaded systems with intricate shutdown process is *waaaay*
too short.  Increase it to the container script default 30s, which
coincidentally is also the container@.conf template's kill delay.

Fixes #1149

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-10-23 15:23:51 +02:00
Joachim Wiberg 391e971573 board/common: expand var partition on sdcard
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-10-13 21:24:14 +02:00
Joachim Wiberg 0957fc3c11 board/common: wait for mmc probe on rpi
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-10-13 21:24:14 +02:00
Joachim Wiberg 3224f49b65 confd: initial zone-based firewall support, based on firewalld
Add supoprt for infix-firewall.yang, modeled on the zone-based firewalld
The terminology is a mix of firewalld, classic netfilter and inspired by
Ubiquity.  E.g., zone 'policy' -> 'action', and the zone matrix overview.

 - Port forwarding allows forwarding a range of ports
 - Operational data comes from firewalld active rules
 - Firewall logging goes to /var/log/firewall.log
 - Show implicit/built-in rules and zones (HOST) in firewall matrix,
   includes "locked" policy for the default-drop behavior
 - The zone services field in admin-exec 'show firewall' shows ANY when
   the zone default action is set to 'accept'
 - Zone 'forwarding' and 'masquerade' settings live in Infix in the
   policys instead, meaning users need to explicitly add a policy
   to allow both intra-zone and inter-zone forwarding
 - Support for emergency lockdown (kill switch)
 - Pre-defined services (xml+enums) are filtered and included as a
   separate YANG model, extensions added for netconf and restconf
 - Includes initial support for firewalld rich rules

firewalld policy rules, including rich rules, have an obnoxious priority
field which is extremely hard to get right, so in Infix we use the far
superior YANG construct 'ordered-by user;'.  This ensure all rules are
generated in that order by setting the priority field, on read-back from
firewalld (operational) the priority field is used to sort the output
of rules in the CLI.

Fixes #448

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-10-10 15:14:12 +02:00
Tobias WaldekranzandGitHub ebb37732f4 Merge pull request #1177 from kernelkit/multi-dsa-tree-fixes
Multi DSA tree fixes
2025-10-02 16:09:19 +02:00
Tobias Waldekranz 94f8d1a700 common: nameif: Handle nested dsa ports
In setups like this...

    CPU
    eth0
     |
.----0----.
| dst0sw0 |
'-1-2-3-4-'
        |
   .----0----.
   | dst1sw0 |
   '-1-2-3-4-'

...both eth0 and dst0sw0p4 are DSA ports. But the latter is _also_ a
physical port as far as devlink is concerned. As a result, it would
first be marked as "internal" and is then instantly reclassified as a
"port".

Catch this condition and stick with the initial "internal"
classification.
2025-10-02 14:51:34 +02:00
Tobias Waldekranz 39b4101d19 common: has-quirk: Add support for matching based on "ethtool -i"
In addition to matching on interface names, add support for matching
on ethtool information.

Example:

    {
        "@ethtool:driver=st_gmac": {
	    "broken-mqprio": true
	}
    }

This would mark any interface using the "st_gmac" driver as having a
broken mqprio implementation. Whereas this:

    {
        "@ethtool:driver=st_gmac;bus-info:30bf0000.ethernet": {
	    "broken-mqprio": true
	}
    }

Only matches an st_gmac-backed interface at the specified location.

As matching becomes more complicated, use the shell implementation
from confd as well, to make sure that they are always in agreement.
2025-10-02 14:51:29 +02:00
Joachim Wiberg 69056bfb2c board/common: fixes for unicode translation in log/pager commands
- cli: add '-r' to log/pager/follow for "raw" control chars, including unicode
 - sysklogd: backport unicode fix for em-dash character (—) in slogan
 - sysklogd: drop old patches
 - sysklogd: run in 8-bit safe mode

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-09-28 03:50:23 +02:00
Joachim Wiberg 6a745f50f8 board/common: adjust log level of udhcpc script, fix missing iface
Issue #1100

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-09-15 18:08:40 +02:00
Joachim Wiberg 77b364d682 board/common: set show-legacy wrapper as executable
Fix #1150

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-09-15 16:22:10 +02:00
Joachim Wiberg 66a5e5304b board/common: fix container network change detection on pull
When a container's image is on an inaccessible remote server, the
container wrapper script waits in the background for any netowrk
changes to retry download of the image.

This change avoids the dangerous previous construct, and is also
easier to read: timeuot after 60 seconds unless ip monitor reads
at least one event before that.

Fixes #1124

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-09-01 14:03:50 +02:00
Joachim Wiberg 6b0145f92e board/common: minor fix to usage text, and silence logs
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-09-01 14:03:50 +02:00
Joachim Wiberg b21043c7e5 package/podman: bump 4.5.0 -> 4.9.5
This major upgrade, along with the upgrade to Finit v4.14, is what is
needed to fix #1123, which was caused by some odd futex locking bug in
Podman that left lingering issues in /var/lib/containers state files.
The root cause as fixed already in v4.7.x, but since CNI is supported
up to and including 4.9.5, going with a later release seemd prudent.

Full changelogs at:
 - <https://github.com/containers/podman/releases/tag/v4.5.1>
 - <https://github.com/containers/podman/releases/tag/v4.6.0>
 - <https://github.com/containers/podman/releases/tag/v4.6.1>
 - <https://github.com/containers/podman/releases/tag/v4.6.2>
 - <https://github.com/containers/podman/releases/tag/v4.7.0>
 - <https://github.com/containers/podman/releases/tag/v4.7.1>
 - <https://github.com/containers/podman/releases/tag/v4.7.2>
 - <https://github.com/containers/podman/releases/tag/v4.8.0>
 - <https://github.com/containers/podman/releases/tag/v4.8.1>
 - <https://github.com/containers/podman/releases/tag/v4.8.2>
 - <https://github.com/containers/podman/releases/tag/v4.8.3>
 - <https://github.com/containers/podman/releases/tag/v4.9.0>
 - <https://github.com/containers/podman/releases/tag/v4.9.1>
 - <https://github.com/containers/podman/releases/tag/v4.9.2>
 - <https://github.com/containers/podman/releases/tag/v4.9.3>
 - <https://github.com/containers/podman/releases/tag/v4.9.4>
 - <https://github.com/containers/podman/releases/tag/v4.9.5>

Fixes #1123

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-09-01 12:25:55 +02:00
Joachim Wiberg 477f7aed72 Follow-up to bb19d064, prune *all* unused images
From the documentation:

> 'podman image prune' removes all dangling images from local storage.
> With the all option, all unused images are deleted (i.e., images not
> in use by any container).
>
> The image prune command does not prune cache images that only use
> layers that are necessary for other images.

So, when the container script is called in the cleanup phase of the
lifetime of a container, we can use the '--all' option to ensure we also
remove this container's loaded image.  In the case this happens before
a reboot of the system, there will be no old version of the image loaded
to /var/lib/containers after boot.

Issue #1098

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-09-01 11:05:35 +02:00
Richard Alpe bb19d064d5 containers: prune dangling images
Previously, upgrading Podman containers with the same tag left
behind dangling images, causing overlay storage to grow and fill
disk space.

This change ensures dangling images are cleaned up using
podman image prune. The command is run without -a, so only
unreferenced images are removed. This provides safe cleanup while
preventing unnecessary overlay growth.

Fixes #1098

Signed-off-by: Richard Alpe <richard@bit42.se>
2025-08-19 12:57:35 +02:00
Mattias Walström 7b3cd7d663 probe: Add USB support for RPI4
Need some special handling for RPI, since RPI does not present a phandle
for the USB in sysfs (PCI device)
2025-07-11 17:06:37 +02:00
Mattias Walström f25f0ab050 Wi-Fi: Add template finit service for wpa_supplicant 2025-06-19 15:23:29 +02:00
Mattias Walström 5560266e19 containers: Allow to have multiple mounts
Only first mount was sent to podman.
2025-05-07 13:24:07 +02:00
Mattias Walström 3375555911 imx8mp-evk: Disable mqprio for eth1
The kernel just hang, adding a quirk for it.
2025-04-28 13:44:19 +02:00
Joachim WibergandGitHub 7b20665cf9 Merge pull request #1004 from kernelkit/new-show-command
Add new show command

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-04-07 14:34:35 +02:00
Richard Alpe c584af356b show: add new show tool written in python3
The reason for this is:
1) We want to use the same fetch tool for all operational data, to
   avoid having duplicated xpaths scattered across the repo.
2) We want to call this directly from the CLI, which means it need to
   be escape safe. A unprivileged user shall be able to use it but not
   escape into a shell or do other malicious stuff.

Signed-off-by: Richard Alpe <richard@bit42.se>
2025-04-07 09:06:26 +02:00
Richard Alpe 827dc098e4 board/common: deprecate existing show command
There's several reasons for this. The main reason is that it was
developed as an intermediate way of getting system information when
the operational datastore was almost empty. There's not a lot of data
available in the operational datastore and we shall rely on that data
primarily. This command isn't safe to run from a jailed CLI as i
doesn't validate input and run commands directly in the shell, making
it hard to protect against a CLI jail break.

In upcoming patches we will introduce a new show tool that relies
solely on operational data and which sanitize the user input.

Signed-off-by: Richard Alpe <richard@bit42.se>
2025-04-07 09:06:25 +02:00
Mattias Walström b4f3e4eab8 migrate: Fix typo resulting in not logging correctly 2025-04-07 08:16:32 +02:00
Joachim Wiberg 8c115e17c4 board/common: set wheel group on /cfg/backup dir
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-04-07 05:32:05 +02:00
Joachim Wiberg 2602ddc0b8 board/common: skip product init if no product specific dir
This silences a bogus warning in the log when runparts is called with no
argument.  I.e., when there is no product specific init files available.

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-03-31 13:35:44 +02:00
Joachim Wiberg 910749bab1 board/common: send SIGTERM when container is in its setup phase
When stopping a container that is (stuck) in setup state, e.g., fetching
its container image, we need to send SIGTERM to the container wrapper
script rather than forwarding the 'stop' command to podman.

With this fix, and the Finit upgrade, there is no longer a 10 second
timeout before Finit sends SIGKILL to the PID.  Instead, the script now
immediately react to the initial SIGTERM.

Related to issue #980

Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
2025-03-31 09:54:13 +02:00