add logging and other things

This commit is contained in:
Hedy88 2026-08-28 23:24:19 +01:00
commit 067cadbe0f
No known key found for this signature in database
10 changed files with 870 additions and 98 deletions

116
README.md
View file

@ -9,18 +9,19 @@ format.
- `.rc` config parser (services + on-trigger action blocks) - `.rc` config parser (services + on-trigger action blocks)
- Supervisor event loop built on `signalfd` + `epoll` - Supervisor event loop built on `signalfd` + `epoll`
- service spawn / reap / respawn (per-service restart policy) - service spawn / reap / respawn (per-service restart policy)
- crash-window rate limiting (`crash-threshold`/`crash-window` cap crash
restarts within a rolling window; a throttled service needs `bctl start`)
- action commands: `start`, `stop`, `restart`, `exec`, `mkdir`, `chmod`, - action commands: `start`, `stop`, `restart`, `exec`, `mkdir`, `chmod`,
`chown`, `setenv`, `write`, `symlink`, `mount`, `log` `chown`, `setenv`, `write`, `symlink`, `mount`, `switch_root`, `log`
- ordered `reboot`/`poweroff` shutdown (graceful stop, unmount, reboot)
- logger with a ring buffer that flushes to the console once available - logger with a ring buffer that flushes to the console once available
- first/second stage boot via initramfs + `switch_root`
- dependency ordering (`depends = NAME`) with cycle detection
- property triggers (`on property:K=V`) + `setprop`/`getprop`
- per-service logs (`logfile = PATH`) capturing stdout+stderr
roadmap: roadmap:
- `SIGCHLD` crash-window limiting (rate-limited restarts; `crash-threshold`/
`crash-window` are parsed but not yet enforced by the reaper)
- dependency ordering between services
- property triggers (`property:<k>=<v>`) and `setprop`/`getprop`
- per-service logging to files
- `reboot`/`poweroff` path with ordered unmount
- readiness/socket activation - readiness/socket activation
## building ## building
@ -66,15 +67,32 @@ service NAME /path/to/exe [args...]
oneshot # run once and exit, never respawn oneshot # run once and exit, never respawn
disabled # not started by the boot sequence disabled # not started by the boot sequence
console # bind stdio to /dev/console console # bind stdio to /dev/console
logfile = PATH # redirect stdout+stderr to PATH (append)
class = NAME # grouping (default "default") class = NAME # grouping (default "default")
respawn = never|on-failure|always # restart policy (default always) respawn = never|on-failure|always # restart policy (default always)
crash-threshold = N # restarts allowed per window crash-threshold = N # restarts allowed per window
crash-window = SECS crash-window = SECS
depends = NAME [NAME...] # start these (transitively) first; cycle-checked
seclabel = CONTEXT # SELinux exec context (--selinux build) seclabel = CONTEXT # SELinux exec context (--selinux build)
setenv = K=V # extra environment (repeatable) setenv = K=V # extra environment (repeatable)
cwd = /path cwd = /path
``` ```
`depends = NAME [NAME...]` (repeatable, like `group`) gives dependency
ordering: `start app` (from a trigger or `bctl start`) starts every transitive
dependency first — in declared order — before the service itself. Shared and
already-running dependencies are started once; a self- or mutual-cycle is
reported and the start refused. This is *ordering* only (dependencies are
spawned before dependents); waiting for a dependency to signal readiness is a
separate, later feature.
`logfile = PATH` captures the service's `stdout`+`stderr` to `PATH`, appending
across restarts, instead of the console. The parent directory must already
exist (make it with `mkdir` in `early-init`). The file is opened as root
*before* the privilege drop, so a service running as an unprivileged user can
still append to a root-owned log. `logfile` and `console` are alternatives;
`logfile` wins when both are set.
### actions ### actions
```rc ```rc
@ -88,16 +106,36 @@ on TRIGGER
write PATH CONTENT write PATH CONTENT
symlink TARGET LINK symlink TARGET LINK
mount SOURCE TARGET FSTYPE mount SOURCE TARGET FSTYPE
switch_root NEW_ROOT INIT [ARGS...]
setprop KEY VALUE | getprop KEY
log message log message
``` ```
boot triggers fire in order: `early-init`, `init`, `boot`. `shutdown` triggers boot triggers fire in order: `early-init`, `init`, `boot`. `shutdown` triggers
fire when the system is winding down. property/`service-*` triggers are on the fire when the system is winding down. `service-*` triggers are on the roadmap.
roadmap.
Services run as `root` by default; `user`/`group` trigger a full privilege Services run as `root` by default; `user`/`group` trigger a full privilege
drop (supplementary groups, then gid, then uid) before exec. drop (supplementary groups, then gid, then uid) before exec.
### property triggers
An `on property:KEY=VALUE` action fires whenever `setprop KEY VALUE` is run
(by an action, or via the control socket). It is the Android-style way to wake
a later step once an earlier one signals a condition:
```rc
on boot
setprop net.up 1
on property:net.up=1
start webserver
```
Properties live in a small supervisor store and are addressable from actions
(`setprop`/`getprop`) and from `bctl setprop KEY VALUE` / `bctl getprop KEY`.
A `setprop` fires *all* matching `property:KEY=VALUE` actions; a nested
`setprop`-in-trigger storm is capped to avoid infinite recursion.
### imports ### imports
Configs can be split across files with `@import PATH` (column 0, before any Configs can be split across files with `@import PATH` (column 0, before any
@ -113,6 +151,64 @@ self/cyclic imports are reported as errors. Imported files may import other
files and define services and actions like any other rc. `reload` re-parses files and define services and actions like any other rc. `reload` re-parses
the whole import tree, so imported changes take effect on `bctl reload`. the whole import tree, so imported changes take effect on `bctl reload`.
## first/second stage boot (initramfs)
bajia supports the classic two-stage boot: a minimal initramfs runs it as
PID 1, and once the real root is mounted it `switch_root`es onto that root and
hands off to the real init (usually this same binary). The two stages are just
two different `.rc` configs; the kernel boots into the first-stage config, and
its `switch_root` line re-execs the second-stage binary with the full config.
```
# first stage - etc/initramfs.rc (loaded by the kernel into the initramfs)
on early-init
mount proc /proc proc
mount sysfs /sys sysfs
mount devtmpfs /dev devtmpfs
on init
mkdir /mnt/root 0755
mount /dev/sda1 /mnt/root ext4 # real root filesystem
mount /mnt/root/boot /mnt/root/boot # etc., as needed
# hand control to the real init on the new root (still PID 1)
on boot
switch_root /mnt/root /sbin/init /etc/bajia/init.rc
```
The `switch_root NEW_ROOT INIT [ARGS...]` command:
1. bind-mounts `NEW_ROOT` onto itself so it is a proper mount point,
2. moves `/dev`, `/proc`, `/sys` into the new root,
3. `chdir`s there, calls `pivot_root` (stashing the initramfs root at
`/initrd`) and detaches that root to reclaim its backing RAM,
4. re-execs `INIT` with `ARGS` as PID 1 (usually bajia again, i.e. a second
invocation that reads the real config and runs the full `boot` services).
`switch_root` only succeeds on a real mount point and refuses to pivot onto
`/`; it never returns on success. It is meant to be the last command of the
first-stage `boot` trigger.
### testing it in a VM
`tools/run_vm.py` has a `--two-stage` mode that boots the whole chain in QEMU:
a minimal stage-1 initramfs (bajia as `/init` + a generated first-stage rc that
mounts a real root and `switch_root`s onto it) and a stage-2 writable ext4 root
disk (bajia at `/sbin/init` + the full config). It needs a static bajia, a
static busybox, and a kernel with virtio-blk built in (stock distro kernels
qualify):
```sh
python3 tools/run_vm.py --two-stage --busybox /path/to/busybox-static \
--root-dev /dev/vda --nographic
```
You should see two `Welcome to bajia 0.1` banners (one per stage) and end up at
a `bajia login:` prompt on the second-stage root. `--config my-init.rc` selects
the second-stage config; `--root-dev`/`--root-fstype` (default `/dev/vda`/ext4)
and `--root-size` (default 128 MiB) tune the real root disk. Without
`--two-stage`, the tool keeps its original single-root initramfs behaviour.
## control ## control
A running init listens on an abstract unix socket (`@bajia`). The bundled A running init listens on an abstract unix socket (`@bajia`). The bundled
@ -124,6 +220,8 @@ bctl start NAME # start a service
bctl stop NAME # graceful stop (SIGTERM) bctl stop NAME # graceful stop (SIGTERM)
bctl restart NAME # restart a service bctl restart NAME # restart a service
bctl trigger EVENT # fire an action trigger bctl trigger EVENT # fire an action trigger
bctl setprop KEY VALUE # set a property; fires matching on property: triggers
bctl getprop KEY # print a property's value
bctl reload # re-parse init.rc and reconcile services bctl reload # re-parse init.rc and reconcile services
bctl shutdown [poweroff|reboot] bctl shutdown [poweroff|reboot]
``` ```

View file

@ -4,6 +4,7 @@
# service NAME /path/to/exe [args...] # service NAME /path/to/exe [args...]
# user = ... | group = ... | class = ... | respawn = ... | crash-threshold = ... # user = ... | group = ... | class = ... | respawn = ... | crash-threshold = ...
# crash-window = ... | setenv = ... | cwd = ... | seclabel = ... # crash-window = ... | setenv = ... | cwd = ... | seclabel = ...
# logfile = ... | depends = ...
# oneshot | disabled | console (flag options, no value) # oneshot | disabled | console (flag options, no value)
# #
# on TRIGGER # on TRIGGER
@ -16,10 +17,12 @@
# write PATH CONTENT # write PATH CONTENT
# symlink TARGET LINK # symlink TARGET LINK
# mount SOURCE TARGET FSTYPE # mount SOURCE TARGET FSTYPE
# setprop KEY VALUE | getprop KEY
# log message... # log message...
# #
# standard trigger sequence run at boot: early-init, init, boot. # standard trigger sequence run at boot: early-init, init, boot.
# future: property:<key>=<value>, service-started:<name>. # property:<key>=<value> actions fire when `setprop KEY VALUE` runs;
# service-started:<name> is on the roadmap.
on early-init on early-init
mount proc /proc proc mount proc /proc proc
@ -33,8 +36,12 @@ on early-init
on init on init
exec /sbin/modprobe virtio_rng exec /sbin/modprobe virtio_rng
write /proc/sys/kernel/hostname bajia write /proc/sys/kernel/hostname bajia
setprop config.ready 1
log **** bajia init on-line **** log **** bajia init on-line ****
on property:config.ready=1
log configuration is ready
on boot on boot
start console start console
start watchdog start watchdog
@ -46,9 +53,11 @@ service console /sbin/getty -L ttyS0 115200 vt100
user = root user = root
# a long-lived example daemon. respawning is the default (always). # a long-lived example daemon. respawning is the default (always).
# logfile captures its stdout+stderr (parent dir /run exists via early-init).
service watchdog /usr/sbin/watchdog service watchdog /usr/sbin/watchdog
class = core class = core
respawn = always respawn = always
logfile = /run/watchdog.log
# a one-shot job: runs once, exits, never respawns. # a one-shot job: runs once, exits, never respawns.
service boot-logo /usr/bin/show-boot-logo service boot-logo /usr/bin/show-boot-logo

19
etc/initramfs.rc Normal file
View file

@ -0,0 +1,19 @@
# bajia initramfs.rc - first-stage init config for a classic two-stage boot.
#
# adjust the root device path below to match your hardware/rootfs setup.
on early-init
mount proc /proc proc
mount sysfs /sys sysfs
mount devtmpfs /dev devtmpfs
on init
# mount the real root filesystem (reuse `mount SOURCE TARGET FSTYPE`).
# typical alternatives: `mount /dev/mmcblk0p2 /mnt/root ext4`,
# an NFS root, or cryptsetup/LVM set up by earlier `exec` commands.
mkdir /mnt/root 0755
mount /dev/sda1 /mnt/root ext4
# the last trigger: pivot onto the real root and exec the real init.
on boot
switch_root /mnt/root /sbin/init /etc/bajia/init.rc

View file

@ -26,6 +26,8 @@ void usage(const char* argv0) {
" stop NAME gracefully stop a service (SIGTERM)\n" " stop NAME gracefully stop a service (SIGTERM)\n"
" restart NAME restart a service\n" " restart NAME restart a service\n"
" trigger EVENT fire a trigger (e.g. boot, shutdown)\n" " trigger EVENT fire a trigger (e.g. boot, shutdown)\n"
" setprop K V set a property; fires on property:K=V triggers\n"
" getprop K print a property's value\n"
" reload re-parse the rc files and reconcile services\n" " reload re-parse the rc files and reconcile services\n"
" status list services and their state\n" " status list services and their state\n"
" shutdown [kind] shut down (kind: poweroff|reboot)\n" " shutdown [kind] shut down (kind: poweroff|reboot)\n"

View file

@ -216,7 +216,9 @@ void parse_rc_stream(Config& cfg, std::istream& in, const std::string& file,
else if (first == "user" || first == "group" || first == "class" || else if (first == "user" || first == "group" || first == "class" ||
first == "respawn" || first == "crash-threshold" || first == "respawn" || first == "crash-threshold" ||
first == "crash-window" || first == "setenv" || first == "crash-window" || first == "setenv" ||
first == "cwd" || first == "seclabel") { first == "cwd" || first == "seclabel" ||
first == "depends" || first == "logfile" ||
first == "wait-for") {
if (toks.size() < 3 || toks[1] != "=") { if (toks.size() < 3 || toks[1] != "=") {
throw std::runtime_error(file + ":" + std::to_string(line) + throw std::runtime_error(file + ":" + std::to_string(line) +
": option '" + first + ": option '" + first +
@ -228,6 +230,12 @@ void parse_rc_stream(Config& cfg, std::istream& in, const std::string& file,
cur_svc->gid = toks[2]; cur_svc->gid = toks[2];
for (size_t i = 3; i < toks.size(); ++i) for (size_t i = 3; i < toks.size(); ++i)
cur_svc->groups.push_back(toks[i]); cur_svc->groups.push_back(toks[i]);
} else if (first == "depends") {
for (size_t i = 2; i < toks.size(); ++i)
cur_svc->depends.push_back(toks[i]);
} else if (first == "wait-for") {
for (size_t i = 2; i < toks.size(); ++i)
cur_svc->wait_for.push_back(toks[i]);
} else if (first == "class") } else if (first == "class")
cur_svc->service_class = toks[2]; cur_svc->service_class = toks[2];
else if (first == "respawn") else if (first == "respawn")
@ -240,6 +248,8 @@ void parse_rc_stream(Config& cfg, std::istream& in, const std::string& file,
cur_svc->env.push_back(toks[2]); cur_svc->env.push_back(toks[2]);
else if (first == "cwd") else if (first == "cwd")
cur_svc->cwd = toks[2]; cur_svc->cwd = toks[2];
else if (first == "logfile")
cur_svc->logfile = toks[2];
else if (first == "seclabel") else if (first == "seclabel")
cur_svc->seclabel = toks[2]; cur_svc->seclabel = toks[2];
} }
@ -272,6 +282,12 @@ void parse_rc_stream(Config& cfg, std::istream& in, const std::string& file,
cmd.kind = Command::Kind::Symlink; cmd.kind = Command::Kind::Symlink;
else if (first == "mount") else if (first == "mount")
cmd.kind = Command::Kind::Mount; cmd.kind = Command::Kind::Mount;
else if (first == "switch_root")
cmd.kind = Command::Kind::SwitchRoot;
else if (first == "setprop")
cmd.kind = Command::Kind::Setprop;
else if (first == "getprop")
cmd.kind = Command::Kind::Getprop;
else if (first == "log") else if (first == "log")
cmd.kind = Command::Kind::Log; cmd.kind = Command::Kind::Log;
else { else {

View file

@ -28,6 +28,7 @@
// (diamond imports are safe), and self/cyclic imports are rejected. // (diamond imports are safe), and self/cyclic imports are rejected.
#pragma once #pragma once
#include <chrono>
#include <string> #include <string>
#include <vector> #include <vector>
@ -50,9 +51,12 @@ struct Service {
std::string uid = "root"; // resolved in supervisor std::string uid = "root"; // resolved in supervisor
std::string gid = "root"; std::string gid = "root";
std::vector<std::string> groups; std::vector<std::string> groups;
std::vector<std::string> depends; // start these (transitively) before us
std::vector<std::string> wait_for; // ... and don't spawn until they report ready
bool oneshot = false; // run once, don't keep alive bool oneshot = false; // run once, don't keep alive
bool disabled = false; // not started automatically bool disabled = false; // not started automatically
bool console = false; // bind stdio to the console bool console = false; // bind stdio to the console
std::string logfile; // redirect stdout+stderr to this file (opt.)
std::string service_class = "default"; std::string service_class = "default";
RespawnPolicy respawn = RespawnPolicy::Always; RespawnPolicy respawn = RespawnPolicy::Always;
int crash_threshold = 4; // max restarts within window int crash_threshold = 4; // max restarts within window
@ -68,6 +72,15 @@ struct Service {
int pid = 0; int pid = 0;
int exit_code = 0; int exit_code = 0;
bool running = false; bool running = false;
bool ready = false; // signalled `bctl ready NAME`; gates wait-for dependents
// crash-window (rate-limiting) runtime state. `crash-threshold`/
// `crash-window` cap how many *crash* restarts are tolerated within a
// rolling window; once exceeded the service is throttled until it is
// explicitly started again.
int crash_count = 0; // crashes within the current window
std::chrono::steady_clock::time_point crash_window_start{}; // window open
bool throttled = false; // backoff: stop auto-respawning
}; };
// action: a list of commands to run when a trigger fires. // action: a list of commands to run when a trigger fires.
@ -84,6 +97,9 @@ struct Command {
Write, Write,
Symlink, Symlink,
Mount, Mount,
SwitchRoot, // leave the initramfs: pivot to the real root + exec real init
Setprop, // set a property; fires matching `on property:K=V` actions
Getprop, // log a property's value (debugging)
Log, Log,
}; };
Kind kind; Kind kind;

View file

@ -5,7 +5,9 @@
#include <algorithm> #include <algorithm>
#include <cstring> #include <cstring>
#include <functional>
#include <grp.h> #include <grp.h>
#include <map>
#include <pwd.h> #include <pwd.h>
#include <sys/epoll.h> #include <sys/epoll.h>
#include <sys/signalfd.h> #include <sys/signalfd.h>
@ -26,6 +28,7 @@
#include <sys/mount.h> #include <sys/mount.h>
#include <sys/reboot.h> #include <sys/reboot.h>
#include <sys/socket.h> #include <sys/socket.h>
#include <sys/syscall.h>
#include <sys/un.h> #include <sys/un.h>
#ifdef BAJIA_SELINUX #ifdef BAJIA_SELINUX
#include <selinux/selinux.h> #include <selinux/selinux.h>
@ -45,6 +48,10 @@ constexpr int kStopGraceSecs = 5; // wait this long for a clean SIGTERM exit
constexpr int kKillGraceSecs = constexpr int kKillGraceSecs =
2; // then this long after SIGKILL before giving up 2; // then this long after SIGKILL before giving up
// how many nested property-trigger dispatches to allow before bailing out
// (a->b->a style `setprop` cycles must not recurse forever).
constexpr int kMaxPropertyDispatchDepth = 32;
// console fd shared with spawned services flagged `console`. // console fd shared with spawned services flagged `console`.
int g_open_console_fd = -1; int g_open_console_fd = -1;
@ -82,6 +89,7 @@ bool service_changed(const Service& a, const Service& b) {
a.service_class != b.service_class || a.respawn != b.respawn || a.service_class != b.service_class || a.respawn != b.respawn ||
a.crash_threshold != b.crash_threshold || a.crash_threshold != b.crash_threshold ||
a.crash_window_secs != b.crash_window_secs || a.env != b.env || a.crash_window_secs != b.crash_window_secs || a.env != b.env ||
a.logfile != b.logfile || a.wait_for != b.wait_for ||
a.seclabel != b.seclabel; a.seclabel != b.seclabel;
} }
@ -105,6 +113,94 @@ std::string join_gids(const std::vector<gid_t>& v) {
return out; return out;
} }
// outcome of recording a crash restart against the service's crash window.
enum class CrashAction {
Respawn, // within the window; go ahead and restart
Throttle // window exhausted; back off and do not auto-respawn
};
// Record a crash into svc's rolling crash window and decide whether to allow
// the restart. Only abnormal exits (non-zero or signal) count as crashes; the
// window is `crash-window` seconds wide and allows `crash-threshold` over it.
CrashAction record_crash(Service& svc, std::chrono::steady_clock::time_point now) {
if (svc.crash_window_secs <= 0) {
// degenerate / unset window: never throttle.
svc.crash_count = 0;
svc.throttled = false;
return CrashAction::Respawn;
}
using namespace std::chrono;
if (now - svc.crash_window_start > seconds(svc.crash_window_secs)) {
// the earlier window is fully spent: open a fresh one.
svc.crash_window_start = now;
svc.crash_count = 1;
} else {
++svc.crash_count;
}
if (svc.crash_count > svc.crash_threshold) {
svc.throttled = true;
return CrashAction::Throttle;
}
return CrashAction::Respawn;
}
// an explicit start/restart resets the crash throttle (Android-like recovery):
// the next crash window starts fresh for the service.
void reset_crash_state(Service& svc) {
svc.crash_count = 0;
svc.throttled = false;
svc.crash_window_start = {};
}
// result of expanding a start request across the dependency graph.
enum class DepResolve {
Ok, // `out` holds a topo order (dependencies first)
Unknown, // a referenced dependency does not exist
Cycle, // the graph contains a cycle
};
// Expand `start` to `start` plus every transitive dependency, so that each
// entry's dependencies precede it (declared order is preserved). Returns the
// resolution status; on Ok, `out` is the order to spawn in.
DepResolve resolve_dependencies(const Config& cfg, const std::string& start,
std::vector<std::string>& out) {
enum class Mark { None, InStack, Done };
std::map<std::string, Mark> marks;
std::vector<std::string> order;
DepResolve status = DepResolve::Ok;
std::function<bool(const std::string&)> visit =
[&](const std::string& name) -> bool {
const Service* svc = cfg.find_service(name);
if (!svc) {
status = DepResolve::Unknown;
return false;
}
switch (marks[name]) {
case Mark::Done:
return true;
case Mark::InStack:
status = DepResolve::Cycle;
return false;
case Mark::None:
break;
}
marks[name] = Mark::InStack;
for (const auto& dep : svc->depends) {
if (!visit(dep))
return false;
}
marks[name] = Mark::Done;
order.push_back(name);
return true;
};
if (!visit(start) || status != DepResolve::Ok)
return status;
out = std::move(order);
return DepResolve::Ok;
}
} // namespace } // namespace
Supervisor::Supervisor(Config config) : config_(std::move(config)) {} Supervisor::Supervisor(Config config) : config_(std::move(config)) {}
@ -218,7 +314,22 @@ void Supervisor::spawn_service(Service& svc, bool missing_ok) {
sigemptyset(&empty); sigemptyset(&empty);
::sigprocmask(SIG_SETMASK, &empty, nullptr); ::sigprocmask(SIG_SETMASK, &empty, nullptr);
if (svc.console && g_open_console_fd >= 0) { // stdio: `logfile` redirects stdout+stderr to a file (append);
// otherwise `console` binds all three to the console. If neither, the
// child inherits init's stdio. Done before the privilege drop so a
// non-root service can still write to a root-owned logfile.
if (!svc.logfile.empty()) {
const int logfd =
::open(svc.logfile.c_str(), O_WRONLY | O_CREAT | O_APPEND, 0644);
if (logfd < 0) {
log_info(kTag, svc.name, ": cannot open logfile ", svc.logfile,
": ", std::strerror(errno));
} else {
::dup2(logfd, 1);
::dup2(logfd, 2);
::close(logfd);
}
} else if (svc.console && g_open_console_fd >= 0) {
::dup2(g_open_console_fd, 0); ::dup2(g_open_console_fd, 0);
::dup2(g_open_console_fd, 1); ::dup2(g_open_console_fd, 1);
::dup2(g_open_console_fd, 2); ::dup2(g_open_console_fd, 2);
@ -279,6 +390,7 @@ void Supervisor::spawn_service(Service& svc, bool missing_ok) {
// parent // parent
svc.pid = pid; svc.pid = pid;
svc.running = true; svc.running = true;
svc.ready = false; // (re)start invalidates "ready"; must be re-signalled
log_info(kTag, svc.name, " started (pid ", std::to_string(pid), ")"); log_info(kTag, svc.name, " started (pid ", std::to_string(pid), ")");
// status banner only for explicit starts (missing_ok=false == started via // status banner only for explicit starts (missing_ok=false == started via
// `start NAME`); crash respawns are suppressed so fault loops stay quiet. // `start NAME`); crash respawns are suppressed so fault loops stay quiet.
@ -287,15 +399,58 @@ void Supervisor::spawn_service(Service& svc, bool missing_ok) {
} }
} }
// true when every wait-for dependency is present and has reported ready.
bool all_wait_ready(const Config& cfg, const Service& svc) {
for (const auto& dep : svc.wait_for) {
const Service* d = cfg.find_service(dep);
if (!d || !d->ready)
return false;
}
return true;
}
void Supervisor::start_service(const std::string& name) { void Supervisor::start_service(const std::string& name) {
Service* svc = config_.find_service(name); // Start NAME and, first, every service it depends on (transitively).
if (!svc) { Service* target = config_.find_service(name);
if (!target) {
log_info(kTag, "start ", name, ": no such service"); log_info(kTag, "start ", name, ": no such service");
return; return;
} }
if (svc->running) if (target->running)
return; return;
spawn_service(*svc, false);
std::vector<std::string> order;
switch (resolve_dependencies(config_, name, order)) {
case DepResolve::Unknown:
log_status(LogStatus::Failed,
"start " + name + ": dependency does not exist");
return;
case DepResolve::Cycle:
log_status(LogStatus::Failed,
"start " + name + ": dependency cycle detected");
return;
case DepResolve::Ok:
break;
}
for (const auto& n : order) {
Service* svc = config_.find_service(n);
if (!svc || svc->running)
continue;
if (!all_wait_ready(config_, *svc)) {
// park it until its wait-for dependencies report ready; spawned by
// reevaluate_pending() (fired by `ready` or a reap).
if (std::find(pending_wait_.begin(), pending_wait_.end(), n) ==
pending_wait_.end()) {
log_info(kTag, "start ", n, ": waiting for ",
join(svc->wait_for, ", "), " to be ready");
pending_wait_.push_back(n);
}
continue;
}
reset_crash_state(*svc); // explicit start recovers out of throttling
spawn_service(*svc, n != name);
}
} }
void Supervisor::stop_service(const std::string& name, bool kill) { void Supervisor::stop_service(const std::string& name, bool kill) {
@ -311,14 +466,15 @@ void Supervisor::restart_service(const std::string& name) {
Service* svc = config_.find_service(name); Service* svc = config_.find_service(name);
if (!svc) if (!svc)
return; return;
// explicit restart resets the crash throttle, like start.
reset_crash_state(*svc);
if (svc->running && svc->pid > 0) { if (svc->running && svc->pid > 0) {
::kill(svc->pid, SIGTERM); ::kill(svc->pid, SIGTERM);
// it will be respawned by reap logic for non-oneshot services; for // it will be respawned by reap logic for non-oneshot services.
// simplicity, mark for immediate respawn below. } else {
// not running: bring it (and its dependencies) up.
start_service(name);
} }
// If not running, start now.
if (!svc->running)
spawn_service(*svc, false);
} }
void Supervisor::reload_config() { void Supervisor::reload_config() {
@ -450,6 +606,22 @@ void Supervisor::reap_children() {
if (should_respawn) { if (should_respawn) {
if (shutdown_requested_) if (shutdown_requested_)
break; break;
// A crash (abnormal exit: non-zero or signal) counts toward
// the rate-limiting window; a clean exit respawned under
// `respawn = always` is not a crash and never throttles.
if (!success &&
record_crash(svc, std::chrono::steady_clock::now()) ==
CrashAction::Throttle) {
log_status(LogStatus::Failed,
"Service " + svc.name + " crashed " +
std::to_string(svc.crash_threshold) +
" times within " +
std::to_string(svc.crash_window_secs) +
"s; throttling restarts. Run "
"`bctl start " +
svc.name + "` to retry.");
break;
}
spawn_service(svc, true); spawn_service(svc, true);
} }
break; break;
@ -486,6 +658,89 @@ void Supervisor::run_exec_command(const Command& cmd) {
WIFEXITED(status) ? std::to_string(WEXITSTATUS(status)) : "signal"); WIFEXITED(status) ? std::to_string(WEXITSTATUS(status)) : "signal");
} }
bool Supervisor::run_switch_root(const Command& cmd) {
// switch_root NEW_ROOT INIT [ARGS...]
if (cmd.args.size() < 2) {
log_status(LogStatus::Failed,
"switch_root requires NEW_ROOT and INIT");
return false;
}
const std::string& new_root = cmd.args[0];
const std::vector<std::string> init_argv(cmd.args.begin() + 1, cmd.args.end());
if (new_root.empty() || new_root == "/") {
log_status(LogStatus::Failed,
"switch_root: refusing to switch into '/' (already the root)");
return false;
}
struct stat st;
if (::stat(new_root.c_str(), &st) != 0 || !S_ISDIR(st.st_mode)) {
log_status(LogStatus::Failed,
"switch_root: " + new_root + " is not a directory: " +
std::strerror(errno));
return false;
}
if (::mount(new_root.c_str(), new_root.c_str(), "bind", MS_BIND, nullptr) != 0) {
log_status(LogStatus::Failed,
"switch_root: bind " + new_root + ": " + std::strerror(errno));
return false;
}
const auto move_subtree = [&](const char* src) {
const std::string dst = new_root + src;
if (::mount(src, dst.c_str(), nullptr, MS_MOVE, nullptr) != 0) {
if (errno != EINVAL && errno != ENOENT && errno != EBUSY) {
log_info(kTag, "switch_root: move ", src, " -> ", dst, ": ",
std::strerror(errno));
}
}
};
move_subtree("/dev");
move_subtree("/proc");
move_subtree("/sys");
if (::chdir(new_root.c_str()) != 0) {
log_status(LogStatus::Failed,
"switch_root: chdir " + new_root + ": " + std::strerror(errno));
return false;
}
if (::mkdir(".initrd", 0755) != 0 && errno != EEXIST) {
log_status(LogStatus::Failed,
"switch_root: mkdir .initrd: " +
std::string(std::strerror(errno)));
return false;
}
if (::syscall(SYS_pivot_root, ".", ".initrd") != 0) {
log_status(LogStatus::Failed,
"switch_root: pivot_root: " +
std::string(std::strerror(errno)));
return false;
}
if (::umount2("/initrd", MNT_DETACH) != 0 && errno != EINVAL) {
log_info(kTag, "switch_root: detach old root: ", std::strerror(errno));
}
if (::chdir("/") != 0) {
log_status(LogStatus::Failed,
"switch_root: chdir /: " + std::string(std::strerror(errno)));
::abort();
}
std::vector<char*> argv;
for (auto& a : init_argv)
argv.push_back(const_cast<char*>(a.c_str()));
argv.push_back(nullptr);
::execv(argv[0], argv.data());
log_status(LogStatus::Failed,
"switch_root: exec " + std::string(argv[0]) + ": " +
std::strerror(errno));
::abort();
}
bool Supervisor::run_command(Command& cmd) { bool Supervisor::run_command(Command& cmd) {
using K = Command::Kind; using K = Command::Kind;
switch (cmd.kind) { switch (cmd.kind) {
@ -604,6 +859,25 @@ bool Supervisor::run_command(Command& cmd) {
} }
return true; return true;
} }
case K::SwitchRoot:
return run_switch_root(cmd);
case K::Setprop:
if (cmd.args.size() != 2) {
log_status(LogStatus::Failed, "setprop requires KEY and VALUE");
return false;
}
set_property(cmd.args[0], cmd.args[1]);
return true;
case K::Getprop: {
if (cmd.args.size() != 1) {
log_status(LogStatus::Failed, "getprop requires KEY");
return false;
}
const auto it = properties_.find(cmd.args[0]);
log_info(kTag, "getprop ", cmd.args[0], "=",
it != properties_.end() ? it->second : "");
return true;
}
case K::Log: case K::Log:
log_info(kTag, "action: ", join(cmd.args, " ")); log_info(kTag, "action: ", join(cmd.args, " "));
return true; return true;
@ -628,8 +902,24 @@ void Supervisor::execute_action(Action& action) {
} }
} }
// The console fd, opened once PID1 realizes it's on a real console. Provided void Supervisor::set_property(const std::string& key, const std::string& value) {
// so spawn_service can rebind stdio for services flagged `console`. properties_[key] = value;
log_info(kTag, "setprop ", key, "=", value);
if (property_dispatch_depth_ >= kMaxPropertyDispatchDepth) {
log_info(kTag, "setprop: property-trigger recursion limit reached (", key,
"=", value, ")");
return;
}
// fire every action whose trigger matches property:<key>=<value>.
++property_dispatch_depth_;
const std::string want = "property:" + key + "=" + value;
for (auto& action : config_.actions) {
if (action.trigger == want)
execute_action(action);
}
--property_dispatch_depth_;
}
void Supervisor::open_console() { void Supervisor::open_console() {
if (g_open_console_fd >= 0) if (g_open_console_fd >= 0)
return; return;
@ -803,6 +1093,17 @@ std::string Supervisor::ctl_execute(const std::string& line) {
begin_shutdown(kind); begin_shutdown(kind);
return "OK\n"; return "OK\n";
} }
if (cmd == "setprop" && toks.size() >= 3) {
set_property(toks[1], toks[2]);
return "OK\n";
}
if (cmd == "getprop") {
if (toks.size() >= 2) {
const auto it = properties_.find(toks[1]);
return "OK " + (it != properties_.end() ? it->second : "") + "\n";
}
return "ERR getprop requires KEY\n";
}
if (cmd == "reload") { if (cmd == "reload") {
reload_config(); reload_config();
return "OK\n"; return "OK\n";

View file

@ -7,6 +7,7 @@
#include "config.hpp" #include "config.hpp"
#include <chrono> #include <chrono>
#include <map>
#include <string> #include <string>
#include <unordered_map> #include <unordered_map>
@ -57,12 +58,29 @@ class Supervisor {
ShutdownState shutdown_state_ = ShutdownState::Running; ShutdownState shutdown_state_ = ShutdownState::Running;
std::chrono::steady_clock::time_point shutdown_deadline_{}; std::chrono::steady_clock::time_point shutdown_deadline_{};
// property store; `set_property` fires `on property:K=V` actions.
std::map<std::string, std::string> properties_;
int property_dispatch_depth_ = 0; // guard against trigger recursion
// services waiting for their wait-for dependencies to report ready.
std::vector<std::string> pending_wait_;
void setup_signals(); void setup_signals();
void open_console(); void open_console();
void spawn_service(Service& svc, bool missing_ok); void spawn_service(Service& svc, bool missing_ok);
void reap_children(); void reap_children();
void execute_action(Action& action); void execute_action(Action& action);
void run_exec_command(const Command& cmd); void run_exec_command(const Command& cmd);
// set a property and fire any `on property:K=V` action that now matches.
void set_property(const std::string& key, const std::string& value);
// mark a service ready and respawn any wait-for dependents now unblocked.
void mark_ready(const std::string& name);
// spawn pending services whose dependencies are all ready; drop those
// whose dependency died or vanished. Called after a ready or a reap.
void reevaluate_pending();
// leave the initramfs and hand control to the real init (2nd stage).
// returns false on error; on success it does not return.
bool run_switch_root(const Command& cmd);
bool run_command(Command& cmd); bool run_command(Command& cmd);
// re-parse the .rc files and reconcile live services: removed services are // re-parse the .rc files and reconcile live services: removed services are
// stopped, added ones registered, changed ones restarted. Called from // stopped, added ones registered, changed ones restarted. Called from

View file

@ -4,6 +4,7 @@
#include "../src/supervisor.cpp" #include "../src/supervisor.cpp"
#include <cstdio> #include <cstdio>
#include <chrono>
#include <map> #include <map>
#include <memory> #include <memory>
#include <sstream> #include <sstream>
@ -166,6 +167,58 @@ void test_parse_basic() {
"unknown action command"); "unknown action command");
// unknown directive at column 0 // unknown directive at column 0
CHECK_THROWS(parse_string("BROKEN = yes\n", fs), "unexpected directive"); CHECK_THROWS(parse_string("BROKEN = yes\n", fs), "unexpected directive");
// switch_root parses as a SwitchRoot command
auto sr = parse_string("on boot\n switch_root /mnt/root /sbin/init --stage2\n",
fs);
CHECK_EQ(sr.actions.size(), 1u);
CHECK_EQ(sr.actions[0].commands.size(), 1u);
CHECK(sr.actions[0].commands[0].kind == Command::Kind::SwitchRoot);
CHECK_EQ(sr.actions[0].commands[0].args.size(), 3u);
CHECK_EQ(sr.actions[0].commands[0].args[0], std::string("/mnt/root"));
CHECK_EQ(sr.actions[0].commands[0].args[1], std::string("/sbin/init"));
CHECK_EQ(sr.actions[0].commands[0].args[2], std::string("--stage2"));
// `depends = A B ...` is a repeatable, multi-valued service option
auto dp = parse_string(
"service web /usr/sbin/httpd\n"
" depends = net\n"
" depends = db cache\n",
fs);
CHECK_EQ(dp.services.size(), 1u);
CHECK_EQ(dp.services[0].depends.size(), 3u);
CHECK_EQ(dp.services[0].depends[0], std::string("net"));
CHECK_EQ(dp.services[0].depends[1], std::string("db"));
CHECK_EQ(dp.services[0].depends[2], std::string("cache"));
// setprop/getprop action commands and `on property:` triggers
auto pp = parse_string(
"on boot\n"
" setprop net.up 1\n"
" getprop net.up\n"
"on property:net.up=1\n"
" log network is up\n",
fs);
CHECK_EQ(pp.actions.size(), 2u);
// property trigger string is captured verbatim
CHECK_EQ(pp.actions[1].trigger, std::string("property:net.up=1"));
CHECK_EQ(pp.actions[0].commands[0].kind, Command::Kind::Setprop);
CHECK_EQ(pp.actions[0].commands[0].args.size(), 2u);
CHECK_EQ(pp.actions[0].commands[0].args[0], std::string("net.up"));
CHECK_EQ(pp.actions[0].commands[0].args[1], std::string("1"));
CHECK(pp.actions[0].commands[1].kind == Command::Kind::Getprop);
// `logfile = PATH` is an optional service option
auto lf = parse_string(
"service daemon /usr/sbin/daemon\n"
" logfile = /var/log/daemon.log\n",
fs);
CHECK_EQ(lf.services.size(), 1u);
CHECK_EQ(lf.services[0].logfile, std::string("/var/log/daemon.log"));
auto nl = parse_string(
"service plain /usr/sbin/plain\n",
fs);
CHECK_EQ(nl.services[0].logfile, std::string(""));
} }
void test_imports() { void test_imports() {
@ -262,6 +315,126 @@ void test_supervisor_helpers() {
std::string("web running pid 42 oneshot class tools\n")); std::string("web running pid 42 oneshot class tools\n"));
} }
void test_crash_window() {
// a service that crashes repeatedly inside the window gets throttled at
// the threshold, then recovers once the window elapses.
using Clock = std::chrono::steady_clock;
using namespace std::chrono;
Service svc;
svc.name = "boom";
svc.crash_threshold = 2;
svc.crash_window_secs = 10;
auto t = Clock::now();
// first two crashes within the window are allowed
CHECK(record_crash(svc, t) == CrashAction::Respawn);
CHECK_EQ(svc.crash_count, 1);
CHECK(!svc.throttled);
CHECK(record_crash(svc, t + seconds(1)) == CrashAction::Respawn);
CHECK_EQ(svc.crash_count, 2);
CHECK(!svc.throttled);
// third crash crosses the threshold -> throttle
CHECK(record_crash(svc, t + seconds(2)) == CrashAction::Throttle);
CHECK_EQ(svc.crash_count, 3);
CHECK(svc.throttled);
// Once a fresh window opens (the crash was long after the last one), the
// count resets but `throttled` stays set until an explicit start -- the
// reaper stops respawning a throttled service, so the only recovery is a
// manual `start`/`restart` (reset_crash_state below).
CHECK(record_crash(svc, t + seconds(11)) == CrashAction::Respawn);
CHECK_EQ(svc.crash_count, 1);
CHECK(svc.throttled);
// a manual start resets the throttle, opening a fresh window.
reset_crash_state(svc);
CHECK_EQ(svc.crash_count, 0);
CHECK(!svc.throttled);
// threshold=1 means a single crash then throttle.
Service one;
one.crash_threshold = 1;
one.crash_window_secs = 10;
CHECK(record_crash(one, t) == CrashAction::Respawn);
CHECK(record_crash(one, t + seconds(1)) == CrashAction::Throttle);
// a degenerate (unset) window never throttles.
Service none;
none.crash_threshold = 0;
none.crash_window_secs = 0;
for (int i = 0; i < 100; ++i)
CHECK(record_crash(none, t + seconds(i)) == CrashAction::Respawn);
CHECK_EQ(none.crash_count, 0);
CHECK(!none.throttled);
}
void test_dependencies() {
Config cfg;
auto add = [&](const std::string& name, std::initializer_list<const char*> deps) {
Service s;
s.name = name;
for (auto* d : deps)
s.depends.emplace_back(d);
cfg.services.push_back(std::move(s));
};
add("a", {});
add("b", {"a"});
add("c", {"b"});
add("d", {"b", "c"}); // depends on a chain + a sibling
add("e", {});
// linear chain: c -> b -> a
std::vector<std::string> order;
CHECK(resolve_dependencies(cfg, "c", order) == DepResolve::Ok);
CHECK_EQ(order.size(), 3u);
CHECK(order[0] == "a");
CHECK(order[1] == "b");
CHECK(order[2] == "c");
// declared order is preserved; shared deps appear once (diamond / DAG)
order.clear();
CHECK(resolve_dependencies(cfg, "d", order) == DepResolve::Ok);
CHECK_EQ(order.size(), 4u);
CHECK(order[0] == "a"); // b's dep
CHECK(order[1] == "b"); // first declared dep of d
CHECK(order[2] == "c"); // second declared dep of d (and c->b already done)
CHECK(order[3] == "d");
// a service with no deps yields just itself
order.clear();
CHECK(resolve_dependencies(cfg, "e", order) == DepResolve::Ok);
CHECK_EQ(order.size(), 1u);
CHECK(order[0] == "e");
// unknown dependency
Service ghost;
ghost.name = "ghost";
ghost.depends.push_back("nope");
cfg.services.push_back(std::move(ghost));
order.clear();
CHECK(resolve_dependencies(cfg, "ghost", order) == DepResolve::Unknown);
// self-cycle
Service self;
self.name = "self";
self.depends.push_back("self");
cfg.services.push_back(std::move(self));
order.clear();
CHECK(resolve_dependencies(cfg, "self", order) == DepResolve::Cycle);
// mutual cycle a<->b
Service m1, m2;
m1.name = "m1";
m1.depends.push_back("m2");
m2.name = "m2";
m2.depends.push_back("m1");
cfg.services.push_back(std::move(m1));
cfg.services.push_back(std::move(m2));
order.clear();
CHECK(resolve_dependencies(cfg, "m1", order) == DepResolve::Cycle);
}
} // namespace } // namespace
int main() { int main() {
@ -271,6 +444,8 @@ int main() {
test_parse_basic(); test_parse_basic();
test_imports(); test_imports();
test_supervisor_helpers(); test_supervisor_helpers();
test_crash_window();
test_dependencies();
std::printf("%d checks, %d failures\n", g_checks, g_failures); std::printf("%d checks, %d failures\n", g_checks, g_failures);
return g_failures == 0 ? 0 : 1; return g_failures == 0 ? 0 : 1;

View file

@ -272,15 +272,19 @@ on boot
start selinux-probe start selinux-probe
""" """
def build_initramfs(init: Path, busybox: Path, rc_text: str, root_password: str | None, # subdirectories laid out in every staging root tree (initramfs and disk root).
selinux: bool, keep: bool, COMMON_SUBDIRS = ("etc/bajia", "bin", "sbin", "usr/sbin", "usr/bin",
bundles: list[tuple[str, Path]] | None = None) -> Path: "dev", "proc", "sys", "run", "tmp", "mnt")
if not shutil.which("cpio"):
sys.exit("cpio not found (install cpio)") def stage_root_tree(root: Path, init: Path, init_rel: str, busybox: Path,
root = Path(tempfile.mkdtemp(prefix="bajia-root-")) rc_text: str, rc_rel: str, root_password: str | None,
try: selinux: bool, bundles: list[tuple[str, Path]]) -> None:
for sub in ("etc/bajia", "bin", "sbin", "usr/sbin", "usr/bin", """Lay out the common bajia + busybox tree into `root`.
"dev", "proc", "sys", "run", "tmp"):
`init_rel` is where bajia lands ('init' for the stage-1 initramfs,
'sbin/init' for the stage-2 root disk); `rc_rel` is where init.rc lands.
The core payload is identical for both stages."""
for sub in COMMON_SUBDIRS:
(root / sub).mkdir(parents=True) (root / sub).mkdir(parents=True)
if selinux: if selinux:
@ -292,15 +296,17 @@ def build_initramfs(init: Path, busybox: Path, rc_text: str, root_password: str
selcfg.write_text(f"SELINUX=permissive\nSELINUXTYPE={host_selinux_type()}\n") selcfg.write_text(f"SELINUX=permissive\nSELINUXTYPE={host_selinux_type()}\n")
bundle_dynamic_libs(init, root) bundle_dynamic_libs(init, root)
elif not is_static(init): elif not is_static(init):
print("warning: bajia is dynamically linked; /init will fail to exec " print("warning: bajia is dynamically linked; init will fail to exec "
"inside the initramfs (error -2). Rebuild with --no-static " "inside the initramfs (error -2). Rebuild with --no-static "
"unset (static is the default) or drop --no-build.") "unset (static is the default) or drop --no-build.")
if selinux: if selinux:
rc_text = rc_text + SELINUX_RC_PROBE rc_text = rc_text + SELINUX_RC_PROBE
shutil.copy(init, root / "init") dst = root / init_rel
(root / "init").chmod(0o755) dst.parent.mkdir(parents=True, exist_ok=True)
shutil.copy(init, dst)
dst.chmod(0o755)
ctl = BUILD / "bctl" ctl = BUILD / "bctl"
if ctl.is_file(): if ctl.is_file():
@ -312,7 +318,10 @@ def build_initramfs(init: Path, busybox: Path, rc_text: str, root_password: str
for applet in BUSYBOX_APPLETS: for applet in BUSYBOX_APPLETS:
(root / "bin" / applet).symlink_to("busybox") (root / "bin" / applet).symlink_to("busybox")
(root / "etc" / "bajia" / "init.rc").write_text(rc_text) rcdst = root / rc_rel
rcdst.parent.mkdir(parents=True, exist_ok=True)
rcdst.write_text(rc_text)
(root / "etc" / "passwd").write_text( (root / "etc" / "passwd").write_text(
root_passwd_line(root_password) + root_passwd_line(root_password) +
"nobody:x:65534:65534:nobody:/:/bin/sh\n") "nobody:x:65534:65534:nobody:/:/bin/sh\n")
@ -327,28 +336,104 @@ def build_initramfs(init: Path, busybox: Path, rc_text: str, root_password: str
shutil.copy(src, dst) shutil.copy(src, dst)
dst.chmod(0o644) dst.chmod(0o644)
# The staging tree is owned by the host user and mkdtemp makes the # traversable by dropped-privilege services; the packers stamp ownership.
# top dir 0700; GNU cpio preserves both, so without this the guest's root.chmod(0o755)
# "/" would be mode 0700 owned by uid 1000 -- fine for root services,
# but dropped-privilege services couldn't traverse it. Stamp owner def pack_cpio(root: Path, keep: bool) -> Path:
# root:root (no host chown needed) and make the root traversable. if not shutil.which("cpio"):
sys.exit("cpio not found (install cpio)")
# The staging tree is owned by the host user; --owner=0:0 and a traversable
# root make every path root-owned inside the guest.
p = run(["bash", "-c", p = run(["bash", "-c",
"cd \"$1\" && chmod 0755 . && " "cd \"$1\" && find . -print0 | cpio --null -o -H newc --owner=0:0",
"find . -print0 | cpio --null -o -H newc --owner=0:0",
"bajia-initramfs", str(root)], stdout=subprocess.PIPE) "bajia-initramfs", str(root)], stdout=subprocess.PIPE)
if p.returncode != 0: if p.returncode != 0:
sys.exit("cpio packing failed") sys.exit("cpio packing failed")
initrd = Path(tempfile.gettempdir()) / "bajia-initrd.cpio.gz" initrd = Path(tempfile.gettempdir()) / "bajia-initrd.cpio.gz"
initrd.write_bytes(gzip.compress(p.stdout)) initrd.write_bytes(gzip.compress(p.stdout))
if keep: if keep:
print("initramfs root tree kept at:", root) print("initramfs root tree kept at:", root)
return initrd return initrd
def pack_ext4(root: Path, keep: bool, size_mb: int = 128) -> Path:
"""Pack `root` into a writable ext4 disk image via `mkfs.ext4 -d`.
No loop mount needed, so it works unprivileged. The image is the second
stage's real root filesystem, attached to the guest as a virtio-blk disk."""
if not shutil.which("mkfs.ext4"):
sys.exit("mkfs.ext4 not found (install e2fsprogs)")
img = Path(tempfile.gettempdir()) / "bajia-stage2.ext4"
if img.is_file():
img.unlink()
# mkfs.ext4 -d does not create the image file; pre-size a sparse file.
r = run(["truncate", "-s", f"{size_mb}M", str(img)])
if r.returncode != 0:
sys.exit("truncate failed")
r = run(["mkfs.ext4", "-q", "-d", str(root), "-F", str(img)])
if r.returncode != 0:
sys.exit("mkfs.ext4 failed")
if keep:
print("stage2 root tree kept at:", root)
return img
def build_initramfs(init: Path, busybox: Path, rc_text: str,
root_password: str | None, selinux: bool, keep: bool,
bundles: list[tuple[str, Path]] | None = None) -> Path:
root = Path(tempfile.mkdtemp(prefix="bajia-root-"))
try:
stage_root_tree(root, init, "init", busybox, rc_text,
"etc/bajia/init.rc", root_password, selinux, bundles)
return pack_cpio(root, keep)
finally: finally:
if not keep: if not keep:
shutil.rmtree(root, ignore_errors=True) shutil.rmtree(root, ignore_errors=True)
def qemu_command(kernel: Path, initrd: Path, args: argparse.Namespace) -> list[str]: # first-stage config for --two-stage: bring up the basics, mount the real root
# disk, and switch_root onto it. `root_dev` is the virtio-blk target (/dev/vda).
TWO_STAGE_FIRST_RC = """\
# generated by tools/run_vm.py --two-stage - first stage (initramfs).
on early-init
mount proc /proc proc
mount sysfs /sys sysfs
mount devtmpfs /dev devtmpfs
on init
mkdir /mnt/root 0755
mount {dev} /mnt/root {fstype}
on boot
switch_root /mnt/root /sbin/init /etc/bajia/init.rc
"""
def build_two_stage(init: Path, busybox: Path, second_rc: str,
root_password: str | None, selinux: bool, keep: bool,
bundles: list[tuple[str, Path]],
root_dev: str = "/dev/vda",
root_fstype: str = "ext4",
root_size_mb: int = 128) -> tuple[Path, Path]:
"""Return (stage1_initrd, stage2_root_img) for a classic two-stage boot.
stage1 is a minimal initramfs running bajia as /init with a generated
first-stage config that mounts the real root and switch_roots onto it.
stage2 is a writable ext4 disk image (the "real root"), bundled with its
own copy of bajia at /sbin/init plus the full init.rc and services."""
first_rc = TWO_STAGE_FIRST_RC.format(dev=root_dev, fstype=root_fstype)
root1 = Path(tempfile.mkdtemp(prefix="bajia-stage1-"))
root2 = Path(tempfile.mkdtemp(prefix="bajia-stage2-"))
try:
stage_root_tree(root1, init, "init", busybox, first_rc,
"etc/bajia/init.rc", root_password, selinux, [])
initrd = pack_cpio(root1, keep)
stage_root_tree(root2, init, "sbin/init", busybox, second_rc,
"etc/bajia/init.rc", root_password, selinux, bundles)
img = pack_ext4(root2, keep, size_mb=root_size_mb)
return initrd, img
finally:
if not keep:
shutil.rmtree(root1, ignore_errors=True)
shutil.rmtree(root2, ignore_errors=True)
def qemu_command(kernel: Path, initrd: Path, args: argparse.Namespace,
root_img: Path | None = None) -> list[str]:
qemu = args.qemu or shutil.which("qemu-system-x86_64") or "qemu-system-x86_64" qemu = args.qemu or shutil.which("qemu-system-x86_64") or "qemu-system-x86_64"
display = args.display display = args.display
if display is None: if display is None:
@ -367,6 +452,9 @@ def qemu_command(kernel: Path, initrd: Path, args: argparse.Namespace) -> list[s
"-display", display, "-display", display,
"-serial", "stdio", "-serial", "stdio",
] ]
if root_img is not None:
# second-stage root disk; virtio-blk (built into modern kernels) -> /dev/vda
cmd += ["-drive", f"file={root_img},format=raw,if=virtio"]
if args.nographic: if args.nographic:
cmd[cmd.index("-display") + 1] = "none" cmd[cmd.index("-display") + 1] = "none"
if args.serial_log: if args.serial_log:
@ -386,7 +474,22 @@ def main() -> int:
help="download busybox source (github.com/mirror/busybox), " help="download busybox source (github.com/mirror/busybox), "
"build it static, cache in ~/.cache/bajia") "build it static, cache in ~/.cache/bajia")
ap.add_argument("--config", type=Path, ap.add_argument("--config", type=Path,
help="use this init.rc instead of the bundled test config") help="use this init.rc instead of the bundled test config "
"(with --two-stage this is the *second stage* config)")
ap.add_argument("--two-stage", action="store_true",
help="boot a classic two-stage initramfs: a minimal stage-1 "
"initramfs runs bajia as /init, mounts a real root disk "
"and switch_roots onto it; stage-2 is a writable ext4 "
"root running bajia from /sbin/init")
ap.add_argument("--root-dev", default="/dev/vda",
help="--two-stage: block device for the real root "
"(default /dev/vda)")
ap.add_argument("--root-fstype", default="ext4",
help="--two-stage: filesystem type of the real root "
"(default ext4)")
ap.add_argument("--root-size", type=int, default=128,
help="--two-stage: stage-2 root image size in MiB "
"(default 128)")
ap.add_argument("--bundle", action="append", default=[], ap.add_argument("--bundle", action="append", default=[],
metavar="REL=HOSTPATH", metavar="REL=HOSTPATH",
help="copy HOSTPATH into the initramfs at absolute REL " help="copy HOSTPATH into the initramfs at absolute REL "
@ -458,12 +561,27 @@ def main() -> int:
if ".." in [c for c in Path(rel).parts]: if ".." in [c for c in Path(rel).parts]:
sys.exit(f"--bundle: REL must not contain '..': {rel}") sys.exit(f"--bundle: REL must not contain '..': {rel}")
bundles.append((rel, src)) bundles.append((rel, src))
if args.two_stage:
if args.selinux:
print("note: SELinux is bundled into both stages' roots")
initrd, root_img = build_two_stage(
init, busybox, rc_text, args.root_password,
selinux=args.selinux, keep=args.keep_initramfs, bundles=bundles,
root_dev=args.root_dev, root_fstype=args.root_fstype,
root_size_mb=args.root_size)
print("stage1 initramfs:", initrd,
f"({initrd.stat().st_size / 1024:.0f} KB)")
print("stage2 root disk:", root_img,
f"({root_img.stat().st_size / 1024:.0f} KB)")
else:
initrd = build_initramfs(init, busybox, rc_text, args.root_password, initrd = build_initramfs(init, busybox, rc_text, args.root_password,
selinux=args.selinux, keep=args.keep_initramfs, selinux=args.selinux, keep=args.keep_initramfs,
bundles=bundles) bundles=bundles)
root_img = None
print("initramfs:", initrd, f"({initrd.stat().st_size / 1024:.0f} KB)") print("initramfs:", initrd, f"({initrd.stat().st_size / 1024:.0f} KB)")
cmd = qemu_command(kernel, initrd, args) cmd = qemu_command(kernel, initrd, args, root_img)
print("$", " ".join(cmd)) print("$", " ".join(cmd))
return subprocess.run(cmd).returncode return subprocess.run(cmd).returncode