302 lines
11 KiB
Markdown
302 lines
11 KiB
Markdown
# bajia
|
|
|
|
an init system (PID 1) for embedded / VM targets, written in C++20 and
|
|
configured with a small declarative language inspired by Android's init `.rc`
|
|
format.
|
|
|
|
## Status
|
|
|
|
- `.rc` config parser (services + on-trigger action blocks)
|
|
- Supervisor event loop built on `signalfd` + `epoll`
|
|
- service spawn / reap / respawn (per-service restart policy)
|
|
- crash-window rate limiting (`crash-threshold`/`crash-window` cap crash
|
|
restarts within a rolling window; a throttled service needs `bctl start`)
|
|
- action commands: `start`, `stop`, `restart`, `exec`, `mkdir`, `chmod`,
|
|
`chown`, `setenv`, `write`, `symlink`, `mount`, `switch_root`, `log`
|
|
- ordered `reboot`/`poweroff` shutdown (graceful stop, unmount, reboot)
|
|
- logger with a ring buffer that flushes to the console once available
|
|
- first/second stage boot via initramfs + `switch_root`
|
|
- dependency ordering (`depends = NAME`) with cycle detection
|
|
- property triggers (`on property:K=V`) + `setprop`/`getprop`
|
|
- per-service logs (`logfile = PATH`) capturing stdout+stderr
|
|
|
|
roadmap:
|
|
|
|
- readiness/socket activation
|
|
|
|
## building
|
|
|
|
requires a C++20 compiler and [Ninja](https://ninja-build.org/)
|
|
|
|
```sh
|
|
python3 configure.py # generates build/build.ninja
|
|
ninja -C build # produces build/bajia
|
|
```
|
|
|
|
`configure.py` also writes a thin `Makefile` convenience wrapper
|
|
(`make`, `make clean`, `make format`, `--asan`, `--debug`).
|
|
|
|
## running
|
|
|
|
As a real init, the kernel must launch it as PID 1:
|
|
|
|
```
|
|
init=/path/to/bajia
|
|
```
|
|
|
|
Or by hand against a config (useful for development, may not behave like a
|
|
real boot). bajia normally refuses to start unless it is PID 1; pass
|
|
`--run-as-user` to override:
|
|
|
|
```sh
|
|
./build/bajia --run-as-user etc/init.rc
|
|
```
|
|
|
|
If no files are given it looks for `/etc/bajia/init.rc`.
|
|
|
|
## configuration language
|
|
|
|
See [`etc/init.rc`](etc/init.rc) for a complete example.
|
|
|
|
### services
|
|
|
|
```rc
|
|
service NAME /path/to/exe [args...]
|
|
user = root|other # uid after the privilege drop (name or number)
|
|
group = GROUP [GROUP...] # primary gid + supplementary groups (names/numbers)
|
|
oneshot # run once and exit, never respawn
|
|
disabled # not started by the boot sequence
|
|
console # bind stdio to /dev/console
|
|
logfile = PATH # redirect stdout+stderr to PATH (append)
|
|
class = NAME # grouping (default "default")
|
|
respawn = never|on-failure|always # restart policy (default always)
|
|
crash-threshold = N # restarts allowed per window
|
|
crash-window = SECS
|
|
depends = NAME [NAME...] # start these (transitively) first; cycle-checked
|
|
seclabel = CONTEXT # SELinux exec context (--selinux build)
|
|
setenv = K=V # extra environment (repeatable)
|
|
cwd = /path
|
|
```
|
|
|
|
`depends = NAME [NAME...]` (repeatable, like `group`) gives dependency
|
|
ordering: `start app` (from a trigger or `bctl start`) starts every transitive
|
|
dependency first — in declared order — before the service itself. Shared and
|
|
already-running dependencies are started once; a self- or mutual-cycle is
|
|
reported and the start refused. This is *ordering* only (dependencies are
|
|
spawned before dependents); waiting for a dependency to signal readiness is a
|
|
separate, later feature.
|
|
|
|
`logfile = PATH` captures the service's `stdout`+`stderr` to `PATH`, appending
|
|
across restarts, instead of the console. The parent directory must already
|
|
exist (make it with `mkdir` in `early-init`). The file is opened as root
|
|
*before* the privilege drop, so a service running as an unprivileged user can
|
|
still append to a root-owned log. `logfile` and `console` are alternatives;
|
|
`logfile` wins when both are set.
|
|
|
|
### actions
|
|
|
|
```rc
|
|
on TRIGGER
|
|
start NAME | stop NAME | restart NAME
|
|
exec /cmd args...
|
|
mkdir PATH [mode]
|
|
chmod PATH mode
|
|
chown PATH uid gid
|
|
setenv K V
|
|
write PATH CONTENT
|
|
symlink TARGET LINK
|
|
mount SOURCE TARGET FSTYPE
|
|
switch_root NEW_ROOT INIT [ARGS...]
|
|
setprop KEY VALUE | getprop KEY
|
|
log message
|
|
```
|
|
|
|
boot triggers fire in order: `early-init`, `init`, `boot`. `shutdown` triggers
|
|
fire when the system is winding down. `service-*` triggers are on the roadmap.
|
|
|
|
Services run as `root` by default; `user`/`group` trigger a full privilege
|
|
drop (supplementary groups, then gid, then uid) before exec.
|
|
|
|
### property triggers
|
|
|
|
An `on property:KEY=VALUE` action fires whenever `setprop KEY VALUE` is run
|
|
(by an action, or via the control socket). It is the Android-style way to wake
|
|
a later step once an earlier one signals a condition:
|
|
|
|
```rc
|
|
on boot
|
|
setprop net.up 1
|
|
|
|
on property:net.up=1
|
|
start webserver
|
|
```
|
|
|
|
Properties live in a small supervisor store and are addressable from actions
|
|
(`setprop`/`getprop`) and from `bctl setprop KEY VALUE` / `bctl getprop KEY`.
|
|
A `setprop` fires *all* matching `property:KEY=VALUE` actions; a nested
|
|
`setprop`-in-trigger storm is capped to avoid infinite recursion.
|
|
|
|
### imports
|
|
|
|
Configs can be split across files with `@import PATH` (column 0, before any
|
|
section in that file):
|
|
|
|
```rc
|
|
@import extra-services.rc
|
|
```
|
|
|
|
The path resolves relative to the importing file's directory (absolute paths
|
|
pass through). A file reached by several imports is only parsed once;
|
|
self/cyclic imports are reported as errors. Imported files may import other
|
|
files and define services and actions like any other rc. `reload` re-parses
|
|
the whole import tree, so imported changes take effect on `bctl reload`.
|
|
|
|
## first/second stage boot (initramfs)
|
|
|
|
bajia supports the classic two-stage boot: a minimal initramfs runs it as
|
|
PID 1, and once the real root is mounted it `switch_root`es onto that root and
|
|
hands off to the real init (usually this same binary). The two stages are just
|
|
two different `.rc` configs; the kernel boots into the first-stage config, and
|
|
its `switch_root` line re-execs the second-stage binary with the full config.
|
|
|
|
```
|
|
# first stage - etc/initramfs.rc (loaded by the kernel into the initramfs)
|
|
on early-init
|
|
mount proc /proc proc
|
|
mount sysfs /sys sysfs
|
|
mount devtmpfs /dev devtmpfs
|
|
|
|
on init
|
|
mkdir /mnt/root 0755
|
|
mount /dev/sda1 /mnt/root ext4 # real root filesystem
|
|
mount /mnt/root/boot /mnt/root/boot # etc., as needed
|
|
|
|
# hand control to the real init on the new root (still PID 1)
|
|
on boot
|
|
switch_root /mnt/root /sbin/init /etc/bajia/init.rc
|
|
```
|
|
|
|
The `switch_root NEW_ROOT INIT [ARGS...]` command:
|
|
|
|
1. bind-mounts `NEW_ROOT` onto itself so it is a proper mount point,
|
|
2. moves `/dev`, `/proc`, `/sys` into the new root,
|
|
3. `chdir`s there, calls `pivot_root` (stashing the initramfs root at
|
|
`/initrd`) and detaches that root to reclaim its backing RAM,
|
|
4. re-execs `INIT` with `ARGS` as PID 1 (usually bajia again, i.e. a second
|
|
invocation that reads the real config and runs the full `boot` services).
|
|
|
|
`switch_root` only succeeds on a real mount point and refuses to pivot onto
|
|
`/`; it never returns on success. It is meant to be the last command of the
|
|
first-stage `boot` trigger.
|
|
|
|
### testing it in a VM
|
|
|
|
`tools/run_vm.py` has a `--two-stage` mode that boots the whole chain in QEMU:
|
|
a minimal stage-1 initramfs (bajia as `/init` + a generated first-stage rc that
|
|
mounts a real root and `switch_root`s onto it) and a stage-2 writable ext4 root
|
|
disk (bajia at `/sbin/init` + the full config). It needs a static bajia, a
|
|
static busybox, and a kernel with virtio-blk built in (stock distro kernels
|
|
qualify):
|
|
|
|
```sh
|
|
python3 tools/run_vm.py --two-stage --busybox /path/to/busybox-static \
|
|
--root-dev /dev/vda --nographic
|
|
```
|
|
|
|
You should see two `Welcome to bajia 0.1` banners (one per stage) and end up at
|
|
a `bajia login:` prompt on the second-stage root. `--config my-init.rc` selects
|
|
the second-stage config; `--root-dev`/`--root-fstype` (default `/dev/vda`/ext4)
|
|
and `--root-size` (default 128 MiB) tune the real root disk. Without
|
|
`--two-stage`, the tool keeps its original single-root initramfs behaviour.
|
|
|
|
## control
|
|
|
|
A running init listens on an abstract unix socket (`@bajia`). The bundled
|
|
`bctl` client drives it:
|
|
|
|
```sh
|
|
bctl status # list services + state
|
|
bctl start NAME # start a service
|
|
bctl stop NAME # graceful stop (SIGTERM)
|
|
bctl restart NAME # restart a service
|
|
bctl trigger EVENT # fire an action trigger
|
|
bctl setprop KEY VALUE # set a property; fires matching on property: triggers
|
|
bctl getprop KEY # print a property's value
|
|
bctl reload # re-parse init.rc and reconcile services
|
|
bctl shutdown [poweroff|reboot]
|
|
```
|
|
|
|
`reload` (also `kill -HUP 1`) re-parses the rc files: removed services are
|
|
stopped, added services registered, and running services whose definition
|
|
changed are restarted with the new definition. A parse error rejects the
|
|
reload and keeps the live config.
|
|
|
|
## SELinux (experimental)
|
|
|
|
SELinux support is opt-in (`configure.py --selinux`, adds `-DBAJIA_SELINUX`
|
|
+ `-lselinux`). With it enabled, PID 1 mounts selinuxfs, loads the policy
|
|
from `/etc/selinux/config`, calls `selinux_restorecon` on the core tree, and
|
|
applies a per-service exec label via the `seclabel = CONTEXT` service option.
|
|
|
|
To bring up a policy without hand-writing one, reuse the host's installed
|
|
policy (Fedora/SELinux hosts have one at `/etc/selinux/<type>/`) in
|
|
permissive mode:
|
|
|
|
```sh
|
|
python3 tools/run_vm.py --selinux
|
|
```
|
|
|
|
This bundles the host `policy.policy.<vers>` and `file_contexts` into the
|
|
initramfs, writes `SELINUX=permissive`, and boots `selinux=1 enforcing=0`.
|
|
Watch the serial console for `selinux: policy loaded, enforcing=0`; a
|
|
`selinux-probe` service prints the runtime exec contexts:
|
|
|
|
```
|
|
probe-ctx=system_u:system_r:init_t:s0 init-ctx=system_u:system_r:kernel_t:s0
|
|
```
|
|
|
|
The guest kernel must support SELinux and the policy version must match the
|
|
kernel's (`cat /sys/fs/selinux/policyvers`). Once permissive is stable,
|
|
read `avc: denied` lines from dmesg and iterate toward a minimal custom
|
|
policy with `checkpolicy`/`audit2allow`, then flip to `enforcing=1`.
|
|
|
|
Limitations: uses the dynamic libselinux (no static build on Fedora), so the
|
|
`--selinux` init is dynamically linked and the loader + libs (`libselinux`,
|
|
`libpcre2-8`, glibc) are bundled into the initramfs.
|
|
|
|
## development & testing
|
|
|
|
Host-side unit tests (no framework, no dependencies) cover the rc parser, the
|
|
`@import` machinery, and the pure supervisor helpers:
|
|
|
|
```sh
|
|
make test # builds build/unit_tests and runs it
|
|
```
|
|
|
|
The parser is also fuzz-tested with libFuzzer (needs clang):
|
|
|
|
```sh
|
|
python3 tools/fuzz.py --seconds 300
|
|
```
|
|
|
|
This drives random bytes through the same `parse_rc_stream` path the real init
|
|
uses, with `@import` rejected so fuzz input can never open real files (e.g.
|
|
`/dev/zero`). Crashes are saved under `build/fuzz-`; seeds accumulate in
|
|
`build/fuzz-corpus` and grow between runs. A grammar dictionary (auto-seeded
|
|
at `build/fuzz.dict`, overridable via `--dict`, disabled with `--no-dict`)
|
|
guides coverage toward real rc keywords. Leak detection is on by default:
|
|
`tools/lsan.supp` silences the spurious `strdup` that a torsocks `LD_PRELOAD`
|
|
on the dev host allocates at startup, so any real leak in `parse_rc_stream` is
|
|
saved as a `leak-*` artifact; pass `--no-detect-leaks` to disable it on a
|
|
clean host. Peak fuzz RSS is driven mostly by ASan's freed-memory quarantine
|
|
(256MiB default); fuzz.py pins it to 64MiB (`--quarantine-mb N`, 0 to
|
|
disable), which roughly halves peak RSS.
|
|
|
|
To leak-check the host-side unit tests under ASan/LSan instead:
|
|
|
|
```sh
|
|
make test-asan
|
|
```
|
|
|
|
(clang++ and a `leak:tsocks_once` suppression are used automatically; the
|
|
default `make test` runs the same assertions without the sanitizer).
|