This allows us to gather better product information in /run/system.json
also for x86 based systems. For Qemu x86_64 systems the 'product-name'
changes from 'VM' to 'Standard PC (i440FX + PIIX, 1996)'.
Also, add new 'product-version' field for, e.g., board revisions, and add
missing 'serial-number' value for systems like Rasperry Pi that inject it
in the device tree.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Validate that DHCP options were requested in the parameter request
list before applying them. This prevents malicious or misconfigured
DHCP servers from forcing unwanted configuration changes.
Validates: hostname (12), DNS (6), domain (15), search (119),
router (3), static routes (121), and NTP (42).
Fail-safe behavior: rejects options if config file unavailable.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Resolve race condition between DHCP and configured hostname by introducing
priority-based hostname management using /etc/hostname.d/ directory pattern.
Priority order (highest wins):
90-dhcp-<iface> - DHCP assigned hostname
50-configured - YANG /system/hostname config
10-default - Bootstrap/factory default
The new /usr/libexec/infix/hostname helper reads all sources and applies the
highest priority hostname. It exits early if hostname unchanged, preventing
unnecessary service restarts.
Fixes#1112
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
After some discussion we agreed that this operation is the domain of
confd (or dagger) and any DHCP client should not mess with the admin
state of the interface it runs on.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Basic DHCPv6 support as a per-interface ipv6 setting, not in line with
the IETF YANG model, which has this as a root container.
Note, this also fixes DHCPv4 option inference after the relocation of
/dhcp-client to /interfaces, in 764bd8e.
Fixes#1110
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Check if partition/LABEL is available before trying to run tune2fs and
mount commands on it. This change masks the kernel output:
LABEL=var: Can't lookup blockdev
commne when running `make run` on x86_64 builds.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
In 45efa94 we introduced new mount options for the: 'aux', 'cfg', and
'var' partitions. To save boot time, we also disabled the periodic fsck
check, which does not affect journal replay or other Ext4 safety or
integrity checks. The latter involved calling tune2fs at runtime, but
this caused the following to appear for the 'aux' partition, which is a
reduced ext4 for systems that need to change U-Boot settings, e.g., to
set MAC address on systems without a VPD.
EXT4-fs: Mount option(s) incompatible with ext2
Note, this is a "bogus" warning, setting the fstypt to ext2 in fstab is
not the way around this one. That's been tried.
To make matters worse, the new mount option(s), e.g., 'commit=30', gave
us another set of warnings from the kernel. This was due to the fstype
being 'auto', so the kernel objected to options not being applicable to
every filesystem it tried before ending up with extfs:
squashfs: Unknown parameter 'commit'
vfat: Unknown parameter 'commit'
exfat: Unknown parameter 'commit'
fuseblk: Unknown parameter 'commit'
btrfs: Unknown parameter 'errors'
This commit explicitly sets cfg and var to ext4, leaving mnt as auto, we
also check before calling tune2fs that the partition/disk image is ext4.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Reduce /var image size to 128M (resized at first boot anyway) and tune
mke2fs options for faster mounting. Set ext4 'uninit_bg' feature and
use '-m 0 -i 4096' extraargs in genimage to optimize filesystem creation
and reduce overhead. The genimage tool has several tweaks enabled by
default already, the 'uninit_bg' feature speeds up the time to check
the file system.
Disable periodic fsck with tune2fs at boot while keeping safety checks
intact. Adjust mount options in fstab to reduce journal syncs and
improve boot time.
Fix severe performance regression in find_partition_by_label() that was
calling sgdisk on every block device including virtual devices (ram, loop,
dm-mapper). This caused boot delays of up to 24 seconds on systems with
many block devices. Now skip virtual devices that don't have GPT tables,
reducing the delay to ~2 seconds.
Also clean up is_mmc() function to use cached result.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The xPi's usually don't have a VPD so the chassis mac-address probed at
boot is usually null in /run/system.json. This commit adds a fallbkack
mechanism to populate this field so it can be used for unique hostnames
even on these boards.
Ths ietf-hardware.yang model does not have a notion of physical address,
so we augment one tht is generic enought to be used for other hardware
components than Ethernet, similar to what ietf-interfaces.yang use.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
On some boards, and in particular with a hybrid mbr/gpt partition table,
like on the RPi64, we must resize ext *after* reboot.
Also, do some cleanup and consolidation of error handling to prevent us
from entering an endless boot loop.
Follow-up to 391e9715
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
This commit refactors USB port probing to:
- Eliminate duplicates: previously, 'authorized' and 'authorized_default'
were listed as separate USB port entries (confusing). Now each USB
port is represented once, with the path pointing to the USB device
directory, confd appends the appropriate attribute file as needed
- Add support for Raspberry Pi 4B and CM4 USB port(s) using a generic
discovery function that scans /sys/bus/usb/devices for USB root hubs.
This should work seamlessly across all platforms
- For backwards compatibility and better UX:
- Single USB port systems: Named "USB" (no number)
- Multi-port systems: Named "USB1", "USB2", etc.
- Device tree-based discovery is tried first (for boards like Alder with
explicit DT USB port definitions), with fallback to generic discovery
for boards without DT
Fixes: #315
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Shell script:
- Factor out big portions of code into more logical helper functions
- Simplify calling setup script by checking for remote image first
- Simplify meta/sha up-to-date handling and clarify terminology
- Consistent use of -f instead of -e in file-exists checks
- Fix unsafe use of 'mktemp -u'
C code:
- Clarify meta/sha terminology: rename meta-sha256 -> meta-image-sha256
- Refactor weird archive_offset() function to local_path() helper
- Factor out helper function calc_sha()
- Check len of sha256 >= 64
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
One unique feature of Infix OS is that you can embed your OCI archive in
the rootfs.squashfs. This way your container instance does not need to
download anything from the network. To upgrade you drop in a new image
and rebuild Infix, when the new Infix boots the container script 'setup'
command will recognize that the OCI archive has changed and will reload
it into the container store.
This patch is an improvement of the way too generic 'podman image prune'
command used previously. Instead of looking for any dangling image, we
now surgically remove the old image.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Container instances that run with mutable images, e.g., tagged with `:latest`
or similar non-versioned tags, can be upgraded without changing the config.
This commit fixes two issues found with this support:
- force container image re-fetch on upgrade, even if the file exists locally
- surgically remove old image from container store after upgrade
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Usually, when your system is up and running properly, you want to clean
up anything unused from your previous experiments. This change alllows
that by calling the interactive 'podman image prune -a -f' command from
the CLI command 'container remove all'
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
When loading big OCI images at boot the `podman load` process completely
monopolizes all cores of an Arm Cortex-A72. It blocks on I/O, sure, but
with 'nice' we can get some attention at least to more critical services
at boot.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
This commit reverts 477f7ae and bb19d06, which intended to fix an issue
with lingering old images, see #1098. However, as detailed in #1147,
this caused severe side effects while working with multiple larger
containers. Basically, the prune operation of one container removed
images of other containers that are just being created in parallel.
Instead of using the podman prune command we can use the meta datain the
start script to pinpoint exactly which image(s) to remove, including any
downloaded OCI archives when the container instance is removed.
Fixes#1147
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
This commit adds metadata to track loaded OCI archives to allow skipping
'delete + load' of OCI images when restarting either the container or the
system as a whole. The sha256 of all loaded OCI archives is stored in a
sidecar file in our downloads directory. Then we verify the checksum of
the OCI archives against their same-named sidecar to determine if the OCI
archive is already loaded or not.
Additionally, the instance using the image is labled with metadata to detect
changes in the container configuration. This in turn allow skipping the
delete + create phase also of the instance.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The default timeout for 'podman stop foo' is 10 seconds, which for
heavily loaded systems with intricate shutdown process is *waaaay*
too short. Increase it to the container script default 30s, which
coincidentally is also the container@.conf template's kill delay.
Fixes#1149
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Add supoprt for infix-firewall.yang, modeled on the zone-based firewalld
The terminology is a mix of firewalld, classic netfilter and inspired by
Ubiquity. E.g., zone 'policy' -> 'action', and the zone matrix overview.
- Port forwarding allows forwarding a range of ports
- Operational data comes from firewalld active rules
- Firewall logging goes to /var/log/firewall.log
- Show implicit/built-in rules and zones (HOST) in firewall matrix,
includes "locked" policy for the default-drop behavior
- The zone services field in admin-exec 'show firewall' shows ANY when
the zone default action is set to 'accept'
- Zone 'forwarding' and 'masquerade' settings live in Infix in the
policys instead, meaning users need to explicitly add a policy
to allow both intra-zone and inter-zone forwarding
- Support for emergency lockdown (kill switch)
- Pre-defined services (xml+enums) are filtered and included as a
separate YANG model, extensions added for netconf and restconf
- Includes initial support for firewalld rich rules
firewalld policy rules, including rich rules, have an obnoxious priority
field which is extremely hard to get right, so in Infix we use the far
superior YANG construct 'ordered-by user;'. This ensure all rules are
generated in that order by setting the priority field, on read-back from
firewalld (operational) the priority field is used to sort the output
of rules in the CLI.
Fixes#448
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
In setups like this...
CPU
eth0
|
.----0----.
| dst0sw0 |
'-1-2-3-4-'
|
.----0----.
| dst1sw0 |
'-1-2-3-4-'
...both eth0 and dst0sw0p4 are DSA ports. But the latter is _also_ a
physical port as far as devlink is concerned. As a result, it would
first be marked as "internal" and is then instantly reclassified as a
"port".
Catch this condition and stick with the initial "internal"
classification.
In addition to matching on interface names, add support for matching
on ethtool information.
Example:
{
"@ethtool:driver=st_gmac": {
"broken-mqprio": true
}
}
This would mark any interface using the "st_gmac" driver as having a
broken mqprio implementation. Whereas this:
{
"@ethtool:driver=st_gmac;bus-info:30bf0000.ethernet": {
"broken-mqprio": true
}
}
Only matches an st_gmac-backed interface at the specified location.
As matching becomes more complicated, use the shell implementation
from confd as well, to make sure that they are always in agreement.
- cli: add '-r' to log/pager/follow for "raw" control chars, including unicode
- sysklogd: backport unicode fix for em-dash character (—) in slogan
- sysklogd: drop old patches
- sysklogd: run in 8-bit safe mode
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
When a container's image is on an inaccessible remote server, the
container wrapper script waits in the background for any netowrk
changes to retry download of the image.
This change avoids the dangerous previous construct, and is also
easier to read: timeuot after 60 seconds unless ip monitor reads
at least one event before that.
Fixes#1124
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
From the documentation:
> 'podman image prune' removes all dangling images from local storage.
> With the all option, all unused images are deleted (i.e., images not
> in use by any container).
>
> The image prune command does not prune cache images that only use
> layers that are necessary for other images.
So, when the container script is called in the cleanup phase of the
lifetime of a container, we can use the '--all' option to ensure we also
remove this container's loaded image. In the case this happens before
a reboot of the system, there will be no old version of the image loaded
to /var/lib/containers after boot.
Issue #1098
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Previously, upgrading Podman containers with the same tag left
behind dangling images, causing overlay storage to grow and fill
disk space.
This change ensures dangling images are cleaned up using
podman image prune. The command is run without -a, so only
unreferenced images are removed. This provides safe cleanup while
preventing unnecessary overlay growth.
Fixes#1098
Signed-off-by: Richard Alpe <richard@bit42.se>
The reason for this is:
1) We want to use the same fetch tool for all operational data, to
avoid having duplicated xpaths scattered across the repo.
2) We want to call this directly from the CLI, which means it need to
be escape safe. A unprivileged user shall be able to use it but not
escape into a shell or do other malicious stuff.
Signed-off-by: Richard Alpe <richard@bit42.se>
There's several reasons for this. The main reason is that it was
developed as an intermediate way of getting system information when
the operational datastore was almost empty. There's not a lot of data
available in the operational datastore and we shall rely on that data
primarily. This command isn't safe to run from a jailed CLI as i
doesn't validate input and run commands directly in the shell, making
it hard to protect against a CLI jail break.
In upcoming patches we will introduce a new show tool that relies
solely on operational data and which sanitize the user input.
Signed-off-by: Richard Alpe <richard@bit42.se>
This silences a bogus warning in the log when runparts is called with no
argument. I.e., when there is no product specific init files available.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
When stopping a container that is (stuck) in setup state, e.g., fetching
its container image, we need to send SIGTERM to the container wrapper
script rather than forwarding the 'stop' command to podman.
With this fix, and the Finit upgrade, there is no longer a 10 second
timeout before Finit sends SIGKILL to the PID. Instead, the script now
immediately react to the initial SIGTERM.
Related to issue #980
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>