The \r-overwrite of an in-progress line breaks when something else
writes to stderr between conout(3) and the final call, leaving blank
"[ OK ]" lines. Cache the description and reprint it whole, so the
final status line is robust to intervening output.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
When an interface has been handed off to a container it lives in
another netns, so `ip neigh/addr flush dev FOO` fails on the host.
The failure aborts dagger, and since interfaces_change() runs before
containers_change() in change_cb(), the container delete path is
never reached -- the stale container keeps the interface trapped in
its netns, breaking the next reconfiguration.
Guard both the neighbor and address flush exit scripts with
`if_nametoindex(ifname)` -- true exactly when the interface is in
the host netns, false for both "in a container" and "already gone".
This replaces the pre-existing `!cni_find(ifname) && if_nametoindex(
ifname)` guard at the addr site: cni_find() added no information
for this check and would popen(container find) for nothing when the
interface had been deleted entirely.
Also harden wrap() in /usr/sbin/container so a stale setup pidfile
doesn't short-circuit Finit's stop attempt -- kill the setup PID and
still ask podman to stop the container.
Fixes#1493
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Add support for static ARP (IPv4) and neighbor cache (IPv6) entries per
interface. Static entries are installed as permanent kernel neighbor
table entries that are never evicted by normal ARP/NDP aging.
Fixes#819
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
On multi-chip DSA hardware (e.g., boards with multiple mv88e6xxx
chips), each switch chip has its own independent PHC device. With
boundary_clock_jbod enabled, ptp4l starts but only disciplines the
active slave port's PHC — the others drift.
Automatically start phc2sys -a alongside any BC or TC instance using
hardware timestamping. It subscribes to ptp4l's UDS socket, tracks
BMCA, and disciplines all non-active PHCs to match the active one.
On single-chip hardware it is a harmless no-op.
CLOCK_REALTIME is intentionally left untouched. Syncing the system
clock to PTP (phc2sys -rr), feeding the PHC from GPS/NTP (ts2phc,
phc2sys reverse), and full multi-source coordination (timemaster) are
planned as follow-on phases; see the issue tracker for the roadmap.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
A container's command line, as reported by podman, may comtain environment
variables, which the current regexp (added in ed4fe58) does not support.
Extend the regexp to allow environment variables and add an operational,
config false, 'cmdline' leaf node to allow any characters to be reported
for the full command line in the operational output.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Unfortunatly the fix was entered in 25.10, but confd was not stepped
up until 25.11, so this should be for 25.10 but now it is for confd
version 1.6 (infix 25.11), as close as we can get.
this is a follwup for commit 7e37fc49a3
confd: prevent IP addresses on bridge ports
Bridge ports should not have IP addresses configured. The IP address
should be configured on the bridge interface itself, not its member ports.
Add YANG must expression to enforce this rule at configuration time.
Fixes#1122
Add a second discovery backend that reads mDNS neighbor data directly
from the sysrepo operational datastore via `copy operational-state -x
/mdns`, selectable with --backend operational.
Add an auto backend (the new default) that tries the operational data
first and falls back to avahi-browse if the `copy` command is
unavailable or mDNS is disabled (visible via "enabled": false in the
JSON). This makes the service work correctly in all configurations
without manual tuning.
Refactor TXT record parsing into parseTxt() and service/host assembly
into buildHosts(), shared by both backends. Extract fetchOpRoot() and
parseOperational() so scanOperational() and scanAuto() share the same
JSON fetch/decode and host-building logic without duplication.
Fix a bug where TXT values containing semicolons were truncated
(parts[9] → strings.Join(parts[9:], ";")), and apply decode() to TXT
values so avahi \DDD escapes are resolved. Add Apple device model
support via the am= (RAOP/AirPlay 1) and model= (AirPlay 2) TXT keys.
Display the last-updated timestamp in 24-hour format regardless of the
browser's locale.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Add reliable avahi client reconnect via explicit free+recreate, reduced
log verbosity (single 10s warning instead of 3×2s loop). Also, binary
TXT record filtering (UTF-8 + XML validity), and finally add a SIGHUP
handler for on-demand reconnect.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The setting 'stratumweight 0.0' disables stratum in Chrony source
selection (pure distance-based), making client_stratum_selection
non-deterministic on a LAN. Setting it to 1.0 gives srv1 a 1-second
effective advantage per stratum level, which no realistic distance
fluctuation can overcome.
Also correct the YANG descriptions in infix-system and infix-ntp which
had the semantics backwards — claiming 0.0 "ensures lower stratum is
always preferred" when in fact higher values do that.
Fixes#1361
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
After a successful bootstrap confd writes a sentinel to /run/confd.boot
If Finit restarts confd — whether after a crash or a clean exit — the
sentinel is found and the destructive bootstrap phases are skipped:
- gen-config fork (factory/failure configs already exist)
- wipe_sysrepo_shm() (other daemons, e.g. statd, are live)
- sr_install_factory_config() (datastores are already initialised)
- sr_replace_config(NULL, NULL) (running datastore is consistent)
- bootstrap_config() / load startup-config (not needed; sysrepo has
the right state; plugins resync via SR_EV_ENABLED on re-subscribe)
On restart confd reconnects to sysrepo, re-initialises plugins (which
re-subscribe and receive SR_EV_ENABLED to resync with the live running
datastore), then enters the steady-state event loop.
The sentinel lives on tmpfs so a real reboot always produces a clean
slate. Crash-loop protection is delegated to Finit's max-restarts (10).
As a side-effect this also enables a future "run-once" mode for resource
constrained systems: confd can exit after bootstrap and the sentinel
ensures any later restart just re-attaches without re-bootstrapping.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Add a finit_enable/disable/reload() family in core.c that directly
manipulates Finit's service state without fork+exec overhead:
finit_enable(svc) -- create symlink in /etc/finit.d/enabled/
finit_disable(svc) -- remove symlink from /etc/finit.d/enabled/
finit_delete(svc) -- remove both symlink and service entirely
finit_reload(svc) -- utimensat() on .conf to schedule reload
Printf-style variants (finit_enablef/disablef/reloadf) handle
template instance names such as container@foo and hostapd@wlan0.
All systemf("initctl ... enable/disable/touch ...") call sites across
containers, dhcp-server, firewall, hardware, ntp, routing, services,
syslog, and system are converted to the new API.
As a related cleanup in services.c, drop the remaining srx_enabled()
calls in favour of reading the already-fetched config tree directly
via lydx_is_enabled(lydx_get_xpathf(config, ...)), eliminating the
last sysrepo round-trips from that module.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Previously svc_enadis() would call 'initctl enable + touch' for every
config change event, even when the service's enabled state was unchanged.
This caused rousette to be unnecessarily restarted on every test_reset(),
racing with the active RESTCONF connection on slow hardware.
Replace svc_enadis() and svc_change() with svc_enable() which only
manages nginx symlinks and calls 'initctl enable/disable' -- never touch.
Each handler now checks the diff for the specific leaf that changed:
- If /enabled appears in diff: call svc_enable() to start or stop it
- If other config leaves changed with service already enabled: touch only
This ensures rousette, ttyd, netbrowse, avahi, sshd, and lldpd are only
restarted when their configuration actually requires it.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Save a few CPU cycles by skipping a new dagger generation when no
interfaces have been modified/added/deleted.
Uses d->next_fp as the sentinel: NULL means no claim was made for this
transaction. dagger_evolve() and dagger_abandon() now NULL it after
fclose, so subsequent unclaimed transactions also get the clean early
return.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
The operator sees Infix through YANG models and should not need to know
which library implements a given feature. Rename the public-facing
parts of the avahi module to use the mdns vocabulary:
- Log strings: "avahi: ..." → "mdns: ..."
- Public API: avahi_ctx_init/exit → mdns_ctx_init/exit
- Main type: struct avahi_ctx → struct mdns_ctx
- statd field: statd.avahi → statd.mdns
Internal types (struct avahi_neighbor, avahi_service, …) and file names
(avahi.c, avahi.h) are kept as-is — developers debugging at the C level
benefit from knowing the underlying implementation.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
When the mdns service is stop-started (e.g. after a config change),
statd's avahi client fires AVAHI_CLIENT_FAILURE momentarily. With
AVAHI_CLIENT_NO_FAIL the client reconnects automatically, but the
immediate ERROR log is misleading:
ERROR: avahi: client failure: Daemon connection failed
New behavior:
- On AVAHI_CLIENT_FAILURE: start a 2 s deferred timer (no immediate log)
- Timer fires up to 3 times (~6 s total); on the 3rd attempt, check if
mDNS is enabled in the running config via a temporary sysrepo session
- Log ERROR only if the daemon is still down AND mDNS is enabled
- On AVAHI_CLIENT_S_RUNNING: cancel the timer, reset the counter, and
log NOTE "mDNS daemon reconnected" if a failure had been seen
This silences the error entirely when the operator has disabled mDNS
(expected), and defers it by ~6 s for a brief restart (self-heals
before the timer fires).
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
- Use same log frameworks as reset of confd
- Use existing primitives from libite + libsrx
- Drop remaining pthreads
- Coding style fixes
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Replace logging + logging.handlers with a lightweight syslog wrapper,
and argparse with manual argv parsing. On a sama7g54, this cuts yanger
startup from ~770ms to ~470ms by eliminating ~300ms of stdlib imports.
Also batch external command invocations:
- ietf_routing: two sysctl calls instead of two per interface
- ietf_hardware: one ls per hwmon device instead of six
- bridge: fetch mctl querier data once instead of once per VLAN
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Only install the keys on CHANGE event, fixes this annoying issue:
Nov 5 01:32:10 ix confd[2011]: Installing HTTPS gencert certificate "self-signed"
Nov 5 01:32:10 ix confd[2011]: Installing SSH host key "genkey".
Nov 5 01:32:11 ix confd[2011]: Installing HTTPS gencert certificate "self-signed
Nov 5 01:32:11 ix confd[2011]: Installing SSH host key "genkey".
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Use SR_SUBSCR_NO_THREAD for all subscriptions and integrate sysrepo
event pipes into a libev event loop. This eliminates approximately 30
per-subscription threads, reducing overhead on embedded ARM hardware.
A temporary poll-based "event pump" thread handles callback dispatch
during bootstrap (where sr_replace_config blocks waiting for callbacks),
then exits. After bootstrap, the single-threaded libev loop takes over
for steady-state event processing.
Note, the confd-test-mode plugin still requires use of threads so we do
not create deadlocks when calling sr_replace_config().
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
- Add x509-public-key-format identity to crypto-types
- Add certificate node to web services container
- Use certificate from ietf-keystore as web cert
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Relocate probe of wifi radios from gen-hardware to 00-probe. This saves
one python invocation and some precious CPU cycles at boot.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Allow confd to start even earlier at boot and call 'gen-config' as a
background task, pending on it to complete before we load the sysrepo
factory datastore.
Also, add Finit style progress to console so users can see the phases
of the bootstrap process.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
On single-core Cortex-A7, the YANG bootstrap and config loading is the
dominant boot bottleneck. The current sequence spawns three serial
phases (bootstrap, sysrepo-plugind, load), each performing independent
sr_connect()/sr_disconnect() cycles, and every sysrepoctl/sysrepocfg
invocation is a fork+exec that rebuilds SHM from scratch.
Replace all three with a single confd binary that does one sr_connect()
and performs all datastore operations in-process:
- Wipe stale /dev/shm/sr_* for a clean slate
- sr_install_factory_config() from the generated JSON
- Smart migration: compare config version via libjansson, only
fork+exec the migrate script when versions actually differ
- Load startup-config (or test-config) via lyd_parse_data() +
sr_replace_config(), mirroring what sysrepocfg -I does internally
- On failure: revert to factory-default, load failure-config, set
login banners (Fail Secure mode)
- On first boot: copy factory-default to running, export to file
- dlopen plugins and enter event loop
The bootstrap shell script is split: config generation (gen-hostname,
gen-interfaces, etc.) stays in the new gen-config script, while all
sysrepo operations move into the C daemon. The finit boot sequence
collapses from 5 stanzas to 2 (gen-config -> confd).
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Replace ERROR() with ERRNO() in all error paths following POSIX API
calls (fopen, rename, realloc, fmkpath, etc.), removing redundant
manual strerror(errno) formatting where present.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Ensure we translate the hostname before we generate the udhcpc command
line options, otherwise we'll send '-x hostname:rpi-%m' to udhcpc.
Also update all board specific factory configs with the missing setting
'"value": "auto"' for the DHCP hostname option. This to match what we
already do when we infer options when enabling a DHCP client in the CLI.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Two races could prevent DHCP-learned default routes from being installed
at boot:
1. The signal from the DHCP client script could be lost leaving conf.d
updated but frr.conf stale.
2. Even when the signal was received, 'vtysh -b' could fail because the
FRR daemon chain (mgmtd→zebra→staticd) writes PID files before being
fully operational, causing netd to give up with no retry.
Fix both by refactoring netd to use libev:
- Use an inotify watch of CONF_DIR, so netd reacts directly to file
changes without depending on signal delivery.
- On backend_apply() failure, schedule a retry in 5s and call pidfile()
unconditionally so dependent services are not blocked waiting for the
Finit condition 'pid/netd' to be satisfied.
We also take this opportunity to rename /etc/netd/conf.d/ to /etc/net.d/
The extra /etc/netd/ directory level served no purpose — nothing else
lives there. Flatten to /etc/net.d/ which reads more naturally as the
drop-in directory for network configuration snippets.
Also, reduce logging a bit since each netd backend already logs success
or fail which is sufficient to know that a configuration change has been
applied or not.
Finally, with the new inotify processing in netd it's redundant to call
'initctl reload netd' from the udhcpc script, in fact it will only cause
unnecessary overhead, so we drop it.
Fixes#1438
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Add a must expression to ensure users do not set static neighbors when
the interface type is not point-to-multipoint or non-broadcast.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Two related problems prevented devices behind an intermediate hub (e.g.
the VIA Labs hub on the RPi 400's VL805 xHCI) from being authorized
when confd unlocks a USB bus:
1. Cascade ordering: usb_authorize() ran nftw() with FTW_DEPTH, which
visits children before the root-hub entry. The hub was authorized
while authorized_default was still 2, so when the kernel probed the
hub's children it immediately denied them. Fix: write
authorized_default=1 on the bus before entering nftw.
2. Slow/async probe: hub port enumeration is asynchronous — devices
behind a hub may appear in sysfs after nftw has already finished.
Fix: add a udev rule that catches usb_device add events on buses
where authorized_default=1 is already set on any ancestor root hub
and authorizes them immediately.
The ATTRS{} matcher in udev walks the full sysfs parent chain, so
authorized_default=1 on the root hub is visible even for devices several
hubs deep. Buses still locked at authorized_default=0 are left alone.
Also expand the docstring in generic_usb_ports() to document the
design: each root hub gets its own uniquely-named entry (USB, or USB1/
USB2/... when multiple) so buses can be individually controlled, and
boards needing finer control should use DT usb-ports annotations.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
Add three SSH-related commands to the operational CLI:
ssh [user <name>] [port <num>] <host>
Connect to a remote device over SSH, running as the
CLI user (not root) by dropping privileges before exec.
set ssh known-hosts <host> <keytype> <pubkey>
Pre-enroll a host public key received out-of-band (e.g.
via email after a factory reset) into ~/.ssh/known_hosts,
avoiding a TOFU prompt on first connect.
no ssh known-hosts <host>
Remove a stale host key entry using ssh-keygen -R, e.g.
after a device factory reset causes a key mismatch.
Tab completion is provided for key types (ssh-ed25519,
ecdsa-sha2-nistp256, etc.) and for known host names/IPs.
A new run_as_user() helper is introduced alongside the existing
run(), factoring out the fork+setuid+execvp pattern used by
infix_shell() so it can be shared across the SSH functions.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
avahi-browse returns 127.0.0.1 (or ::1) when resolving services on the
same machine netbrowse is running on. These addresses are meaningless
for display and misleading as click targets. Skip the entire 127.x/::1
range and link-local (fe80:) when choosing the preferred address for a
host card; the card addr field is simply omitted if no routable address
is seen.
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>
- IP address shown in soft parens on the card header line, e.g.
"infix.local (192.168.1.10)"; IPv4 preferred over IPv6, link-local
skipped
- Product name and OS version (from mDNS TXT records "product=" and
"ov=") shown as a secondary line below the header when available
- Theme toggle icons changed from ◐/●/○ to ◐/☽/☀ (system/dark/light)
- browse.go: introduce Host struct (addr, product, version, other, svcs)
replacing the flat []Service map; scan() now returns map[string]Host
- Search also matches against IP, product, and version fields
Signed-off-by: Joachim Wiberg <troglobit@gmail.com>