bajia/README.md

302 lines
11 KiB
Markdown

# bajia
an init system (PID 1) for embedded / VM targets, written in C++20 and
configured with a small declarative language inspired by Android's init `.rc`
format.
## Status
- `.rc` config parser (services + on-trigger action blocks)
- Supervisor event loop built on `signalfd` + `epoll`
- service spawn / reap / respawn (per-service restart policy)
- crash-window rate limiting (`crash-threshold`/`crash-window` cap crash
restarts within a rolling window; a throttled service needs `bctl start`)
- action commands: `start`, `stop`, `restart`, `exec`, `mkdir`, `chmod`,
`chown`, `setenv`, `write`, `symlink`, `mount`, `switch_root`, `log`
- ordered `reboot`/`poweroff` shutdown (graceful stop, unmount, reboot)
- logger with a ring buffer that flushes to the console once available
- first/second stage boot via initramfs + `switch_root`
- dependency ordering (`depends = NAME`) with cycle detection
- property triggers (`on property:K=V`) + `setprop`/`getprop`
- per-service logs (`logfile = PATH`) capturing stdout+stderr
roadmap:
- readiness/socket activation
## building
requires a C++20 compiler and [Ninja](https://ninja-build.org/)
```sh
python3 configure.py # generates build/build.ninja
ninja -C build # produces build/bajia
```
`configure.py` also writes a thin `Makefile` convenience wrapper
(`make`, `make clean`, `make format`, `--asan`, `--debug`).
## running
As a real init, the kernel must launch it as PID 1:
```
init=/path/to/bajia
```
Or by hand against a config (useful for development, may not behave like a
real boot). bajia normally refuses to start unless it is PID 1; pass
`--run-as-user` to override:
```sh
./build/bajia --run-as-user etc/init.rc
```
If no files are given it looks for `/etc/bajia/init.rc`.
## configuration language
See [`etc/init.rc`](etc/init.rc) for a complete example.
### services
```rc
service NAME /path/to/exe [args...]
user = root|other # uid after the privilege drop (name or number)
group = GROUP [GROUP...] # primary gid + supplementary groups (names/numbers)
oneshot # run once and exit, never respawn
disabled # not started by the boot sequence
console # bind stdio to /dev/console
logfile = PATH # redirect stdout+stderr to PATH (append)
class = NAME # grouping (default "default")
respawn = never|on-failure|always # restart policy (default always)
crash-threshold = N # restarts allowed per window
crash-window = SECS
depends = NAME [NAME...] # start these (transitively) first; cycle-checked
seclabel = CONTEXT # SELinux exec context (--selinux build)
setenv = K=V # extra environment (repeatable)
cwd = /path
```
`depends = NAME [NAME...]` (repeatable, like `group`) gives dependency
ordering: `start app` (from a trigger or `bctl start`) starts every transitive
dependency first — in declared order — before the service itself. Shared and
already-running dependencies are started once; a self- or mutual-cycle is
reported and the start refused. This is *ordering* only (dependencies are
spawned before dependents); waiting for a dependency to signal readiness is a
separate, later feature.
`logfile = PATH` captures the service's `stdout`+`stderr` to `PATH`, appending
across restarts, instead of the console. The parent directory must already
exist (make it with `mkdir` in `early-init`). The file is opened as root
*before* the privilege drop, so a service running as an unprivileged user can
still append to a root-owned log. `logfile` and `console` are alternatives;
`logfile` wins when both are set.
### actions
```rc
on TRIGGER
start NAME | stop NAME | restart NAME
exec /cmd args...
mkdir PATH [mode]
chmod PATH mode
chown PATH uid gid
setenv K V
write PATH CONTENT
symlink TARGET LINK
mount SOURCE TARGET FSTYPE
switch_root NEW_ROOT INIT [ARGS...]
setprop KEY VALUE | getprop KEY
log message
```
boot triggers fire in order: `early-init`, `init`, `boot`. `shutdown` triggers
fire when the system is winding down. `service-*` triggers are on the roadmap.
Services run as `root` by default; `user`/`group` trigger a full privilege
drop (supplementary groups, then gid, then uid) before exec.
### property triggers
An `on property:KEY=VALUE` action fires whenever `setprop KEY VALUE` is run
(by an action, or via the control socket). It is the Android-style way to wake
a later step once an earlier one signals a condition:
```rc
on boot
setprop net.up 1
on property:net.up=1
start webserver
```
Properties live in a small supervisor store and are addressable from actions
(`setprop`/`getprop`) and from `bctl setprop KEY VALUE` / `bctl getprop KEY`.
A `setprop` fires *all* matching `property:KEY=VALUE` actions; a nested
`setprop`-in-trigger storm is capped to avoid infinite recursion.
### imports
Configs can be split across files with `@import PATH` (column 0, before any
section in that file):
```rc
@import extra-services.rc
```
The path resolves relative to the importing file's directory (absolute paths
pass through). A file reached by several imports is only parsed once;
self/cyclic imports are reported as errors. Imported files may import other
files and define services and actions like any other rc. `reload` re-parses
the whole import tree, so imported changes take effect on `bctl reload`.
## first/second stage boot (initramfs)
bajia supports the classic two-stage boot: a minimal initramfs runs it as
PID 1, and once the real root is mounted it `switch_root`es onto that root and
hands off to the real init (usually this same binary). The two stages are just
two different `.rc` configs; the kernel boots into the first-stage config, and
its `switch_root` line re-execs the second-stage binary with the full config.
```
# first stage - etc/initramfs.rc (loaded by the kernel into the initramfs)
on early-init
mount proc /proc proc
mount sysfs /sys sysfs
mount devtmpfs /dev devtmpfs
on init
mkdir /mnt/root 0755
mount /dev/sda1 /mnt/root ext4 # real root filesystem
mount /mnt/root/boot /mnt/root/boot # etc., as needed
# hand control to the real init on the new root (still PID 1)
on boot
switch_root /mnt/root /sbin/init /etc/bajia/init.rc
```
The `switch_root NEW_ROOT INIT [ARGS...]` command:
1. bind-mounts `NEW_ROOT` onto itself so it is a proper mount point,
2. moves `/dev`, `/proc`, `/sys` into the new root,
3. `chdir`s there, calls `pivot_root` (stashing the initramfs root at
`/initrd`) and detaches that root to reclaim its backing RAM,
4. re-execs `INIT` with `ARGS` as PID 1 (usually bajia again, i.e. a second
invocation that reads the real config and runs the full `boot` services).
`switch_root` only succeeds on a real mount point and refuses to pivot onto
`/`; it never returns on success. It is meant to be the last command of the
first-stage `boot` trigger.
### testing it in a VM
`tools/run_vm.py` has a `--two-stage` mode that boots the whole chain in QEMU:
a minimal stage-1 initramfs (bajia as `/init` + a generated first-stage rc that
mounts a real root and `switch_root`s onto it) and a stage-2 writable ext4 root
disk (bajia at `/sbin/init` + the full config). It needs a static bajia, a
static busybox, and a kernel with virtio-blk built in (stock distro kernels
qualify):
```sh
python3 tools/run_vm.py --two-stage --busybox /path/to/busybox-static \
--root-dev /dev/vda --nographic
```
You should see two `Welcome to bajia 0.1` banners (one per stage) and end up at
a `bajia login:` prompt on the second-stage root. `--config my-init.rc` selects
the second-stage config; `--root-dev`/`--root-fstype` (default `/dev/vda`/ext4)
and `--root-size` (default 128 MiB) tune the real root disk. Without
`--two-stage`, the tool keeps its original single-root initramfs behaviour.
## control
A running init listens on an abstract unix socket (`@bajia`). The bundled
`bctl` client drives it:
```sh
bctl status # list services + state
bctl start NAME # start a service
bctl stop NAME # graceful stop (SIGTERM)
bctl restart NAME # restart a service
bctl trigger EVENT # fire an action trigger
bctl setprop KEY VALUE # set a property; fires matching on property: triggers
bctl getprop KEY # print a property's value
bctl reload # re-parse init.rc and reconcile services
bctl shutdown [poweroff|reboot]
```
`reload` (also `kill -HUP 1`) re-parses the rc files: removed services are
stopped, added services registered, and running services whose definition
changed are restarted with the new definition. A parse error rejects the
reload and keeps the live config.
## SELinux (experimental)
SELinux support is opt-in (`configure.py --selinux`, adds `-DBAJIA_SELINUX`
+ `-lselinux`). With it enabled, PID 1 mounts selinuxfs, loads the policy
from `/etc/selinux/config`, calls `selinux_restorecon` on the core tree, and
applies a per-service exec label via the `seclabel = CONTEXT` service option.
To bring up a policy without hand-writing one, reuse the host's installed
policy (Fedora/SELinux hosts have one at `/etc/selinux/<type>/`) in
permissive mode:
```sh
python3 tools/run_vm.py --selinux
```
This bundles the host `policy.policy.<vers>` and `file_contexts` into the
initramfs, writes `SELINUX=permissive`, and boots `selinux=1 enforcing=0`.
Watch the serial console for `selinux: policy loaded, enforcing=0`; a
`selinux-probe` service prints the runtime exec contexts:
```
probe-ctx=system_u:system_r:init_t:s0 init-ctx=system_u:system_r:kernel_t:s0
```
The guest kernel must support SELinux and the policy version must match the
kernel's (`cat /sys/fs/selinux/policyvers`). Once permissive is stable,
read `avc: denied` lines from dmesg and iterate toward a minimal custom
policy with `checkpolicy`/`audit2allow`, then flip to `enforcing=1`.
Limitations: uses the dynamic libselinux (no static build on Fedora), so the
`--selinux` init is dynamically linked and the loader + libs (`libselinux`,
`libpcre2-8`, glibc) are bundled into the initramfs.
## development & testing
Host-side unit tests (no framework, no dependencies) cover the rc parser, the
`@import` machinery, and the pure supervisor helpers:
```sh
make test # builds build/unit_tests and runs it
```
The parser is also fuzz-tested with libFuzzer (needs clang):
```sh
python3 tools/fuzz.py --seconds 300
```
This drives random bytes through the same `parse_rc_stream` path the real init
uses, with `@import` rejected so fuzz input can never open real files (e.g.
`/dev/zero`). Crashes are saved under `build/fuzz-`; seeds accumulate in
`build/fuzz-corpus` and grow between runs. A grammar dictionary (auto-seeded
at `build/fuzz.dict`, overridable via `--dict`, disabled with `--no-dict`)
guides coverage toward real rc keywords. Leak detection is on by default:
`tools/lsan.supp` silences the spurious `strdup` that a torsocks `LD_PRELOAD`
on the dev host allocates at startup, so any real leak in `parse_rc_stream` is
saved as a `leak-*` artifact; pass `--no-detect-leaks` to disable it on a
clean host. Peak fuzz RSS is driven mostly by ASan's freed-memory quarantine
(256MiB default); fuzz.py pins it to 64MiB (`--quarantine-mb N`, 0 to
disable), which roughly halves peak RSS.
To leak-check the host-side unit tests under ASan/LSan instead:
```sh
make test-asan
```
(clang++ and a `leak:tsocks_once` suppression are used automatically; the
default `make test` runs the same assertions without the sanitizer).