add logging and other things
This commit is contained in:
parent
a4650fd589
commit
067cadbe0f
10 changed files with 870 additions and 98 deletions
116
README.md
116
README.md
|
|
@ -9,18 +9,19 @@ format.
|
|||
- `.rc` config parser (services + on-trigger action blocks)
|
||||
- Supervisor event loop built on `signalfd` + `epoll`
|
||||
- service spawn / reap / respawn (per-service restart policy)
|
||||
- crash-window rate limiting (`crash-threshold`/`crash-window` cap crash
|
||||
restarts within a rolling window; a throttled service needs `bctl start`)
|
||||
- action commands: `start`, `stop`, `restart`, `exec`, `mkdir`, `chmod`,
|
||||
`chown`, `setenv`, `write`, `symlink`, `mount`, `log`
|
||||
`chown`, `setenv`, `write`, `symlink`, `mount`, `switch_root`, `log`
|
||||
- ordered `reboot`/`poweroff` shutdown (graceful stop, unmount, reboot)
|
||||
- logger with a ring buffer that flushes to the console once available
|
||||
- first/second stage boot via initramfs + `switch_root`
|
||||
- dependency ordering (`depends = NAME`) with cycle detection
|
||||
- property triggers (`on property:K=V`) + `setprop`/`getprop`
|
||||
- per-service logs (`logfile = PATH`) capturing stdout+stderr
|
||||
|
||||
roadmap:
|
||||
|
||||
- `SIGCHLD` crash-window limiting (rate-limited restarts; `crash-threshold`/
|
||||
`crash-window` are parsed but not yet enforced by the reaper)
|
||||
- dependency ordering between services
|
||||
- property triggers (`property:<k>=<v>`) and `setprop`/`getprop`
|
||||
- per-service logging to files
|
||||
- `reboot`/`poweroff` path with ordered unmount
|
||||
- readiness/socket activation
|
||||
|
||||
## building
|
||||
|
|
@ -66,15 +67,32 @@ service NAME /path/to/exe [args...]
|
|||
oneshot # run once and exit, never respawn
|
||||
disabled # not started by the boot sequence
|
||||
console # bind stdio to /dev/console
|
||||
logfile = PATH # redirect stdout+stderr to PATH (append)
|
||||
class = NAME # grouping (default "default")
|
||||
respawn = never|on-failure|always # restart policy (default always)
|
||||
crash-threshold = N # restarts allowed per window
|
||||
crash-window = SECS
|
||||
depends = NAME [NAME...] # start these (transitively) first; cycle-checked
|
||||
seclabel = CONTEXT # SELinux exec context (--selinux build)
|
||||
setenv = K=V # extra environment (repeatable)
|
||||
cwd = /path
|
||||
```
|
||||
|
||||
`depends = NAME [NAME...]` (repeatable, like `group`) gives dependency
|
||||
ordering: `start app` (from a trigger or `bctl start`) starts every transitive
|
||||
dependency first — in declared order — before the service itself. Shared and
|
||||
already-running dependencies are started once; a self- or mutual-cycle is
|
||||
reported and the start refused. This is *ordering* only (dependencies are
|
||||
spawned before dependents); waiting for a dependency to signal readiness is a
|
||||
separate, later feature.
|
||||
|
||||
`logfile = PATH` captures the service's `stdout`+`stderr` to `PATH`, appending
|
||||
across restarts, instead of the console. The parent directory must already
|
||||
exist (make it with `mkdir` in `early-init`). The file is opened as root
|
||||
*before* the privilege drop, so a service running as an unprivileged user can
|
||||
still append to a root-owned log. `logfile` and `console` are alternatives;
|
||||
`logfile` wins when both are set.
|
||||
|
||||
### actions
|
||||
|
||||
```rc
|
||||
|
|
@ -88,16 +106,36 @@ on TRIGGER
|
|||
write PATH CONTENT
|
||||
symlink TARGET LINK
|
||||
mount SOURCE TARGET FSTYPE
|
||||
switch_root NEW_ROOT INIT [ARGS...]
|
||||
setprop KEY VALUE | getprop KEY
|
||||
log message
|
||||
```
|
||||
|
||||
boot triggers fire in order: `early-init`, `init`, `boot`. `shutdown` triggers
|
||||
fire when the system is winding down. property/`service-*` triggers are on the
|
||||
roadmap.
|
||||
fire when the system is winding down. `service-*` triggers are on the roadmap.
|
||||
|
||||
Services run as `root` by default; `user`/`group` trigger a full privilege
|
||||
drop (supplementary groups, then gid, then uid) before exec.
|
||||
|
||||
### property triggers
|
||||
|
||||
An `on property:KEY=VALUE` action fires whenever `setprop KEY VALUE` is run
|
||||
(by an action, or via the control socket). It is the Android-style way to wake
|
||||
a later step once an earlier one signals a condition:
|
||||
|
||||
```rc
|
||||
on boot
|
||||
setprop net.up 1
|
||||
|
||||
on property:net.up=1
|
||||
start webserver
|
||||
```
|
||||
|
||||
Properties live in a small supervisor store and are addressable from actions
|
||||
(`setprop`/`getprop`) and from `bctl setprop KEY VALUE` / `bctl getprop KEY`.
|
||||
A `setprop` fires *all* matching `property:KEY=VALUE` actions; a nested
|
||||
`setprop`-in-trigger storm is capped to avoid infinite recursion.
|
||||
|
||||
### imports
|
||||
|
||||
Configs can be split across files with `@import PATH` (column 0, before any
|
||||
|
|
@ -113,6 +151,64 @@ self/cyclic imports are reported as errors. Imported files may import other
|
|||
files and define services and actions like any other rc. `reload` re-parses
|
||||
the whole import tree, so imported changes take effect on `bctl reload`.
|
||||
|
||||
## first/second stage boot (initramfs)
|
||||
|
||||
bajia supports the classic two-stage boot: a minimal initramfs runs it as
|
||||
PID 1, and once the real root is mounted it `switch_root`es onto that root and
|
||||
hands off to the real init (usually this same binary). The two stages are just
|
||||
two different `.rc` configs; the kernel boots into the first-stage config, and
|
||||
its `switch_root` line re-execs the second-stage binary with the full config.
|
||||
|
||||
```
|
||||
# first stage - etc/initramfs.rc (loaded by the kernel into the initramfs)
|
||||
on early-init
|
||||
mount proc /proc proc
|
||||
mount sysfs /sys sysfs
|
||||
mount devtmpfs /dev devtmpfs
|
||||
|
||||
on init
|
||||
mkdir /mnt/root 0755
|
||||
mount /dev/sda1 /mnt/root ext4 # real root filesystem
|
||||
mount /mnt/root/boot /mnt/root/boot # etc., as needed
|
||||
|
||||
# hand control to the real init on the new root (still PID 1)
|
||||
on boot
|
||||
switch_root /mnt/root /sbin/init /etc/bajia/init.rc
|
||||
```
|
||||
|
||||
The `switch_root NEW_ROOT INIT [ARGS...]` command:
|
||||
|
||||
1. bind-mounts `NEW_ROOT` onto itself so it is a proper mount point,
|
||||
2. moves `/dev`, `/proc`, `/sys` into the new root,
|
||||
3. `chdir`s there, calls `pivot_root` (stashing the initramfs root at
|
||||
`/initrd`) and detaches that root to reclaim its backing RAM,
|
||||
4. re-execs `INIT` with `ARGS` as PID 1 (usually bajia again, i.e. a second
|
||||
invocation that reads the real config and runs the full `boot` services).
|
||||
|
||||
`switch_root` only succeeds on a real mount point and refuses to pivot onto
|
||||
`/`; it never returns on success. It is meant to be the last command of the
|
||||
first-stage `boot` trigger.
|
||||
|
||||
### testing it in a VM
|
||||
|
||||
`tools/run_vm.py` has a `--two-stage` mode that boots the whole chain in QEMU:
|
||||
a minimal stage-1 initramfs (bajia as `/init` + a generated first-stage rc that
|
||||
mounts a real root and `switch_root`s onto it) and a stage-2 writable ext4 root
|
||||
disk (bajia at `/sbin/init` + the full config). It needs a static bajia, a
|
||||
static busybox, and a kernel with virtio-blk built in (stock distro kernels
|
||||
qualify):
|
||||
|
||||
```sh
|
||||
python3 tools/run_vm.py --two-stage --busybox /path/to/busybox-static \
|
||||
--root-dev /dev/vda --nographic
|
||||
```
|
||||
|
||||
You should see two `Welcome to bajia 0.1` banners (one per stage) and end up at
|
||||
a `bajia login:` prompt on the second-stage root. `--config my-init.rc` selects
|
||||
the second-stage config; `--root-dev`/`--root-fstype` (default `/dev/vda`/ext4)
|
||||
and `--root-size` (default 128 MiB) tune the real root disk. Without
|
||||
`--two-stage`, the tool keeps its original single-root initramfs behaviour.
|
||||
|
||||
## control
|
||||
|
||||
A running init listens on an abstract unix socket (`@bajia`). The bundled
|
||||
|
|
@ -124,6 +220,8 @@ bctl start NAME # start a service
|
|||
bctl stop NAME # graceful stop (SIGTERM)
|
||||
bctl restart NAME # restart a service
|
||||
bctl trigger EVENT # fire an action trigger
|
||||
bctl setprop KEY VALUE # set a property; fires matching on property: triggers
|
||||
bctl getprop KEY # print a property's value
|
||||
bctl reload # re-parse init.rc and reconcile services
|
||||
bctl shutdown [poweroff|reboot]
|
||||
```
|
||||
|
|
|
|||
11
etc/init.rc
11
etc/init.rc
|
|
@ -4,6 +4,7 @@
|
|||
# service NAME /path/to/exe [args...]
|
||||
# user = ... | group = ... | class = ... | respawn = ... | crash-threshold = ...
|
||||
# crash-window = ... | setenv = ... | cwd = ... | seclabel = ...
|
||||
# logfile = ... | depends = ...
|
||||
# oneshot | disabled | console (flag options, no value)
|
||||
#
|
||||
# on TRIGGER
|
||||
|
|
@ -16,10 +17,12 @@
|
|||
# write PATH CONTENT
|
||||
# symlink TARGET LINK
|
||||
# mount SOURCE TARGET FSTYPE
|
||||
# setprop KEY VALUE | getprop KEY
|
||||
# log message...
|
||||
#
|
||||
# standard trigger sequence run at boot: early-init, init, boot.
|
||||
# future: property:<key>=<value>, service-started:<name>.
|
||||
# property:<key>=<value> actions fire when `setprop KEY VALUE` runs;
|
||||
# service-started:<name> is on the roadmap.
|
||||
|
||||
on early-init
|
||||
mount proc /proc proc
|
||||
|
|
@ -33,8 +36,12 @@ on early-init
|
|||
on init
|
||||
exec /sbin/modprobe virtio_rng
|
||||
write /proc/sys/kernel/hostname bajia
|
||||
setprop config.ready 1
|
||||
log **** bajia init on-line ****
|
||||
|
||||
on property:config.ready=1
|
||||
log configuration is ready
|
||||
|
||||
on boot
|
||||
start console
|
||||
start watchdog
|
||||
|
|
@ -46,9 +53,11 @@ service console /sbin/getty -L ttyS0 115200 vt100
|
|||
user = root
|
||||
|
||||
# a long-lived example daemon. respawning is the default (always).
|
||||
# logfile captures its stdout+stderr (parent dir /run exists via early-init).
|
||||
service watchdog /usr/sbin/watchdog
|
||||
class = core
|
||||
respawn = always
|
||||
logfile = /run/watchdog.log
|
||||
|
||||
# a one-shot job: runs once, exits, never respawns.
|
||||
service boot-logo /usr/bin/show-boot-logo
|
||||
|
|
|
|||
19
etc/initramfs.rc
Normal file
19
etc/initramfs.rc
Normal file
|
|
@ -0,0 +1,19 @@
|
|||
# bajia initramfs.rc - first-stage init config for a classic two-stage boot.
|
||||
#
|
||||
# adjust the root device path below to match your hardware/rootfs setup.
|
||||
|
||||
on early-init
|
||||
mount proc /proc proc
|
||||
mount sysfs /sys sysfs
|
||||
mount devtmpfs /dev devtmpfs
|
||||
|
||||
on init
|
||||
# mount the real root filesystem (reuse `mount SOURCE TARGET FSTYPE`).
|
||||
# typical alternatives: `mount /dev/mmcblk0p2 /mnt/root ext4`,
|
||||
# an NFS root, or cryptsetup/LVM set up by earlier `exec` commands.
|
||||
mkdir /mnt/root 0755
|
||||
mount /dev/sda1 /mnt/root ext4
|
||||
|
||||
# the last trigger: pivot onto the real root and exec the real init.
|
||||
on boot
|
||||
switch_root /mnt/root /sbin/init /etc/bajia/init.rc
|
||||
|
|
@ -26,6 +26,8 @@ void usage(const char* argv0) {
|
|||
" stop NAME gracefully stop a service (SIGTERM)\n"
|
||||
" restart NAME restart a service\n"
|
||||
" trigger EVENT fire a trigger (e.g. boot, shutdown)\n"
|
||||
" setprop K V set a property; fires on property:K=V triggers\n"
|
||||
" getprop K print a property's value\n"
|
||||
" reload re-parse the rc files and reconcile services\n"
|
||||
" status list services and their state\n"
|
||||
" shutdown [kind] shut down (kind: poweroff|reboot)\n"
|
||||
|
|
|
|||
|
|
@ -216,7 +216,9 @@ void parse_rc_stream(Config& cfg, std::istream& in, const std::string& file,
|
|||
else if (first == "user" || first == "group" || first == "class" ||
|
||||
first == "respawn" || first == "crash-threshold" ||
|
||||
first == "crash-window" || first == "setenv" ||
|
||||
first == "cwd" || first == "seclabel") {
|
||||
first == "cwd" || first == "seclabel" ||
|
||||
first == "depends" || first == "logfile" ||
|
||||
first == "wait-for") {
|
||||
if (toks.size() < 3 || toks[1] != "=") {
|
||||
throw std::runtime_error(file + ":" + std::to_string(line) +
|
||||
": option '" + first +
|
||||
|
|
@ -228,6 +230,12 @@ void parse_rc_stream(Config& cfg, std::istream& in, const std::string& file,
|
|||
cur_svc->gid = toks[2];
|
||||
for (size_t i = 3; i < toks.size(); ++i)
|
||||
cur_svc->groups.push_back(toks[i]);
|
||||
} else if (first == "depends") {
|
||||
for (size_t i = 2; i < toks.size(); ++i)
|
||||
cur_svc->depends.push_back(toks[i]);
|
||||
} else if (first == "wait-for") {
|
||||
for (size_t i = 2; i < toks.size(); ++i)
|
||||
cur_svc->wait_for.push_back(toks[i]);
|
||||
} else if (first == "class")
|
||||
cur_svc->service_class = toks[2];
|
||||
else if (first == "respawn")
|
||||
|
|
@ -240,6 +248,8 @@ void parse_rc_stream(Config& cfg, std::istream& in, const std::string& file,
|
|||
cur_svc->env.push_back(toks[2]);
|
||||
else if (first == "cwd")
|
||||
cur_svc->cwd = toks[2];
|
||||
else if (first == "logfile")
|
||||
cur_svc->logfile = toks[2];
|
||||
else if (first == "seclabel")
|
||||
cur_svc->seclabel = toks[2];
|
||||
}
|
||||
|
|
@ -272,6 +282,12 @@ void parse_rc_stream(Config& cfg, std::istream& in, const std::string& file,
|
|||
cmd.kind = Command::Kind::Symlink;
|
||||
else if (first == "mount")
|
||||
cmd.kind = Command::Kind::Mount;
|
||||
else if (first == "switch_root")
|
||||
cmd.kind = Command::Kind::SwitchRoot;
|
||||
else if (first == "setprop")
|
||||
cmd.kind = Command::Kind::Setprop;
|
||||
else if (first == "getprop")
|
||||
cmd.kind = Command::Kind::Getprop;
|
||||
else if (first == "log")
|
||||
cmd.kind = Command::Kind::Log;
|
||||
else {
|
||||
|
|
|
|||
|
|
@ -28,6 +28,7 @@
|
|||
// (diamond imports are safe), and self/cyclic imports are rejected.
|
||||
#pragma once
|
||||
|
||||
#include <chrono>
|
||||
#include <string>
|
||||
#include <vector>
|
||||
|
||||
|
|
@ -50,9 +51,12 @@ struct Service {
|
|||
std::string uid = "root"; // resolved in supervisor
|
||||
std::string gid = "root";
|
||||
std::vector<std::string> groups;
|
||||
std::vector<std::string> depends; // start these (transitively) before us
|
||||
std::vector<std::string> wait_for; // ... and don't spawn until they report ready
|
||||
bool oneshot = false; // run once, don't keep alive
|
||||
bool disabled = false; // not started automatically
|
||||
bool console = false; // bind stdio to the console
|
||||
std::string logfile; // redirect stdout+stderr to this file (opt.)
|
||||
std::string service_class = "default";
|
||||
RespawnPolicy respawn = RespawnPolicy::Always;
|
||||
int crash_threshold = 4; // max restarts within window
|
||||
|
|
@ -68,6 +72,15 @@ struct Service {
|
|||
int pid = 0;
|
||||
int exit_code = 0;
|
||||
bool running = false;
|
||||
bool ready = false; // signalled `bctl ready NAME`; gates wait-for dependents
|
||||
|
||||
// crash-window (rate-limiting) runtime state. `crash-threshold`/
|
||||
// `crash-window` cap how many *crash* restarts are tolerated within a
|
||||
// rolling window; once exceeded the service is throttled until it is
|
||||
// explicitly started again.
|
||||
int crash_count = 0; // crashes within the current window
|
||||
std::chrono::steady_clock::time_point crash_window_start{}; // window open
|
||||
bool throttled = false; // backoff: stop auto-respawning
|
||||
};
|
||||
|
||||
// action: a list of commands to run when a trigger fires.
|
||||
|
|
@ -84,6 +97,9 @@ struct Command {
|
|||
Write,
|
||||
Symlink,
|
||||
Mount,
|
||||
SwitchRoot, // leave the initramfs: pivot to the real root + exec real init
|
||||
Setprop, // set a property; fires matching `on property:K=V` actions
|
||||
Getprop, // log a property's value (debugging)
|
||||
Log,
|
||||
};
|
||||
Kind kind;
|
||||
|
|
|
|||
|
|
@ -5,7 +5,9 @@
|
|||
|
||||
#include <algorithm>
|
||||
#include <cstring>
|
||||
#include <functional>
|
||||
#include <grp.h>
|
||||
#include <map>
|
||||
#include <pwd.h>
|
||||
#include <sys/epoll.h>
|
||||
#include <sys/signalfd.h>
|
||||
|
|
@ -26,6 +28,7 @@
|
|||
#include <sys/mount.h>
|
||||
#include <sys/reboot.h>
|
||||
#include <sys/socket.h>
|
||||
#include <sys/syscall.h>
|
||||
#include <sys/un.h>
|
||||
#ifdef BAJIA_SELINUX
|
||||
#include <selinux/selinux.h>
|
||||
|
|
@ -45,6 +48,10 @@ constexpr int kStopGraceSecs = 5; // wait this long for a clean SIGTERM exit
|
|||
constexpr int kKillGraceSecs =
|
||||
2; // then this long after SIGKILL before giving up
|
||||
|
||||
// how many nested property-trigger dispatches to allow before bailing out
|
||||
// (a->b->a style `setprop` cycles must not recurse forever).
|
||||
constexpr int kMaxPropertyDispatchDepth = 32;
|
||||
|
||||
// console fd shared with spawned services flagged `console`.
|
||||
int g_open_console_fd = -1;
|
||||
|
||||
|
|
@ -82,6 +89,7 @@ bool service_changed(const Service& a, const Service& b) {
|
|||
a.service_class != b.service_class || a.respawn != b.respawn ||
|
||||
a.crash_threshold != b.crash_threshold ||
|
||||
a.crash_window_secs != b.crash_window_secs || a.env != b.env ||
|
||||
a.logfile != b.logfile || a.wait_for != b.wait_for ||
|
||||
a.seclabel != b.seclabel;
|
||||
}
|
||||
|
||||
|
|
@ -105,6 +113,94 @@ std::string join_gids(const std::vector<gid_t>& v) {
|
|||
return out;
|
||||
}
|
||||
|
||||
// outcome of recording a crash restart against the service's crash window.
|
||||
enum class CrashAction {
|
||||
Respawn, // within the window; go ahead and restart
|
||||
Throttle // window exhausted; back off and do not auto-respawn
|
||||
};
|
||||
|
||||
// Record a crash into svc's rolling crash window and decide whether to allow
|
||||
// the restart. Only abnormal exits (non-zero or signal) count as crashes; the
|
||||
// window is `crash-window` seconds wide and allows `crash-threshold` over it.
|
||||
CrashAction record_crash(Service& svc, std::chrono::steady_clock::time_point now) {
|
||||
if (svc.crash_window_secs <= 0) {
|
||||
// degenerate / unset window: never throttle.
|
||||
svc.crash_count = 0;
|
||||
svc.throttled = false;
|
||||
return CrashAction::Respawn;
|
||||
}
|
||||
using namespace std::chrono;
|
||||
if (now - svc.crash_window_start > seconds(svc.crash_window_secs)) {
|
||||
// the earlier window is fully spent: open a fresh one.
|
||||
svc.crash_window_start = now;
|
||||
svc.crash_count = 1;
|
||||
} else {
|
||||
++svc.crash_count;
|
||||
}
|
||||
if (svc.crash_count > svc.crash_threshold) {
|
||||
svc.throttled = true;
|
||||
return CrashAction::Throttle;
|
||||
}
|
||||
return CrashAction::Respawn;
|
||||
}
|
||||
|
||||
// an explicit start/restart resets the crash throttle (Android-like recovery):
|
||||
// the next crash window starts fresh for the service.
|
||||
void reset_crash_state(Service& svc) {
|
||||
svc.crash_count = 0;
|
||||
svc.throttled = false;
|
||||
svc.crash_window_start = {};
|
||||
}
|
||||
|
||||
// result of expanding a start request across the dependency graph.
|
||||
enum class DepResolve {
|
||||
Ok, // `out` holds a topo order (dependencies first)
|
||||
Unknown, // a referenced dependency does not exist
|
||||
Cycle, // the graph contains a cycle
|
||||
};
|
||||
|
||||
// Expand `start` to `start` plus every transitive dependency, so that each
|
||||
// entry's dependencies precede it (declared order is preserved). Returns the
|
||||
// resolution status; on Ok, `out` is the order to spawn in.
|
||||
DepResolve resolve_dependencies(const Config& cfg, const std::string& start,
|
||||
std::vector<std::string>& out) {
|
||||
enum class Mark { None, InStack, Done };
|
||||
std::map<std::string, Mark> marks;
|
||||
std::vector<std::string> order;
|
||||
DepResolve status = DepResolve::Ok;
|
||||
|
||||
std::function<bool(const std::string&)> visit =
|
||||
[&](const std::string& name) -> bool {
|
||||
const Service* svc = cfg.find_service(name);
|
||||
if (!svc) {
|
||||
status = DepResolve::Unknown;
|
||||
return false;
|
||||
}
|
||||
switch (marks[name]) {
|
||||
case Mark::Done:
|
||||
return true;
|
||||
case Mark::InStack:
|
||||
status = DepResolve::Cycle;
|
||||
return false;
|
||||
case Mark::None:
|
||||
break;
|
||||
}
|
||||
marks[name] = Mark::InStack;
|
||||
for (const auto& dep : svc->depends) {
|
||||
if (!visit(dep))
|
||||
return false;
|
||||
}
|
||||
marks[name] = Mark::Done;
|
||||
order.push_back(name);
|
||||
return true;
|
||||
};
|
||||
|
||||
if (!visit(start) || status != DepResolve::Ok)
|
||||
return status;
|
||||
out = std::move(order);
|
||||
return DepResolve::Ok;
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
Supervisor::Supervisor(Config config) : config_(std::move(config)) {}
|
||||
|
|
@ -218,7 +314,22 @@ void Supervisor::spawn_service(Service& svc, bool missing_ok) {
|
|||
sigemptyset(&empty);
|
||||
::sigprocmask(SIG_SETMASK, &empty, nullptr);
|
||||
|
||||
if (svc.console && g_open_console_fd >= 0) {
|
||||
// stdio: `logfile` redirects stdout+stderr to a file (append);
|
||||
// otherwise `console` binds all three to the console. If neither, the
|
||||
// child inherits init's stdio. Done before the privilege drop so a
|
||||
// non-root service can still write to a root-owned logfile.
|
||||
if (!svc.logfile.empty()) {
|
||||
const int logfd =
|
||||
::open(svc.logfile.c_str(), O_WRONLY | O_CREAT | O_APPEND, 0644);
|
||||
if (logfd < 0) {
|
||||
log_info(kTag, svc.name, ": cannot open logfile ", svc.logfile,
|
||||
": ", std::strerror(errno));
|
||||
} else {
|
||||
::dup2(logfd, 1);
|
||||
::dup2(logfd, 2);
|
||||
::close(logfd);
|
||||
}
|
||||
} else if (svc.console && g_open_console_fd >= 0) {
|
||||
::dup2(g_open_console_fd, 0);
|
||||
::dup2(g_open_console_fd, 1);
|
||||
::dup2(g_open_console_fd, 2);
|
||||
|
|
@ -279,6 +390,7 @@ void Supervisor::spawn_service(Service& svc, bool missing_ok) {
|
|||
// parent
|
||||
svc.pid = pid;
|
||||
svc.running = true;
|
||||
svc.ready = false; // (re)start invalidates "ready"; must be re-signalled
|
||||
log_info(kTag, svc.name, " started (pid ", std::to_string(pid), ")");
|
||||
// status banner only for explicit starts (missing_ok=false == started via
|
||||
// `start NAME`); crash respawns are suppressed so fault loops stay quiet.
|
||||
|
|
@ -287,15 +399,58 @@ void Supervisor::spawn_service(Service& svc, bool missing_ok) {
|
|||
}
|
||||
}
|
||||
|
||||
// true when every wait-for dependency is present and has reported ready.
|
||||
bool all_wait_ready(const Config& cfg, const Service& svc) {
|
||||
for (const auto& dep : svc.wait_for) {
|
||||
const Service* d = cfg.find_service(dep);
|
||||
if (!d || !d->ready)
|
||||
return false;
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
void Supervisor::start_service(const std::string& name) {
|
||||
Service* svc = config_.find_service(name);
|
||||
if (!svc) {
|
||||
// Start NAME and, first, every service it depends on (transitively).
|
||||
Service* target = config_.find_service(name);
|
||||
if (!target) {
|
||||
log_info(kTag, "start ", name, ": no such service");
|
||||
return;
|
||||
}
|
||||
if (svc->running)
|
||||
if (target->running)
|
||||
return;
|
||||
spawn_service(*svc, false);
|
||||
|
||||
std::vector<std::string> order;
|
||||
switch (resolve_dependencies(config_, name, order)) {
|
||||
case DepResolve::Unknown:
|
||||
log_status(LogStatus::Failed,
|
||||
"start " + name + ": dependency does not exist");
|
||||
return;
|
||||
case DepResolve::Cycle:
|
||||
log_status(LogStatus::Failed,
|
||||
"start " + name + ": dependency cycle detected");
|
||||
return;
|
||||
case DepResolve::Ok:
|
||||
break;
|
||||
}
|
||||
|
||||
for (const auto& n : order) {
|
||||
Service* svc = config_.find_service(n);
|
||||
if (!svc || svc->running)
|
||||
continue;
|
||||
if (!all_wait_ready(config_, *svc)) {
|
||||
// park it until its wait-for dependencies report ready; spawned by
|
||||
// reevaluate_pending() (fired by `ready` or a reap).
|
||||
if (std::find(pending_wait_.begin(), pending_wait_.end(), n) ==
|
||||
pending_wait_.end()) {
|
||||
log_info(kTag, "start ", n, ": waiting for ",
|
||||
join(svc->wait_for, ", "), " to be ready");
|
||||
pending_wait_.push_back(n);
|
||||
}
|
||||
continue;
|
||||
}
|
||||
reset_crash_state(*svc); // explicit start recovers out of throttling
|
||||
spawn_service(*svc, n != name);
|
||||
}
|
||||
}
|
||||
|
||||
void Supervisor::stop_service(const std::string& name, bool kill) {
|
||||
|
|
@ -311,14 +466,15 @@ void Supervisor::restart_service(const std::string& name) {
|
|||
Service* svc = config_.find_service(name);
|
||||
if (!svc)
|
||||
return;
|
||||
// explicit restart resets the crash throttle, like start.
|
||||
reset_crash_state(*svc);
|
||||
if (svc->running && svc->pid > 0) {
|
||||
::kill(svc->pid, SIGTERM);
|
||||
// it will be respawned by reap logic for non-oneshot services; for
|
||||
// simplicity, mark for immediate respawn below.
|
||||
// it will be respawned by reap logic for non-oneshot services.
|
||||
} else {
|
||||
// not running: bring it (and its dependencies) up.
|
||||
start_service(name);
|
||||
}
|
||||
// If not running, start now.
|
||||
if (!svc->running)
|
||||
spawn_service(*svc, false);
|
||||
}
|
||||
|
||||
void Supervisor::reload_config() {
|
||||
|
|
@ -450,6 +606,22 @@ void Supervisor::reap_children() {
|
|||
if (should_respawn) {
|
||||
if (shutdown_requested_)
|
||||
break;
|
||||
// A crash (abnormal exit: non-zero or signal) counts toward
|
||||
// the rate-limiting window; a clean exit respawned under
|
||||
// `respawn = always` is not a crash and never throttles.
|
||||
if (!success &&
|
||||
record_crash(svc, std::chrono::steady_clock::now()) ==
|
||||
CrashAction::Throttle) {
|
||||
log_status(LogStatus::Failed,
|
||||
"Service " + svc.name + " crashed " +
|
||||
std::to_string(svc.crash_threshold) +
|
||||
" times within " +
|
||||
std::to_string(svc.crash_window_secs) +
|
||||
"s; throttling restarts. Run "
|
||||
"`bctl start " +
|
||||
svc.name + "` to retry.");
|
||||
break;
|
||||
}
|
||||
spawn_service(svc, true);
|
||||
}
|
||||
break;
|
||||
|
|
@ -486,6 +658,89 @@ void Supervisor::run_exec_command(const Command& cmd) {
|
|||
WIFEXITED(status) ? std::to_string(WEXITSTATUS(status)) : "signal");
|
||||
}
|
||||
|
||||
bool Supervisor::run_switch_root(const Command& cmd) {
|
||||
// switch_root NEW_ROOT INIT [ARGS...]
|
||||
if (cmd.args.size() < 2) {
|
||||
log_status(LogStatus::Failed,
|
||||
"switch_root requires NEW_ROOT and INIT");
|
||||
return false;
|
||||
}
|
||||
const std::string& new_root = cmd.args[0];
|
||||
const std::vector<std::string> init_argv(cmd.args.begin() + 1, cmd.args.end());
|
||||
|
||||
if (new_root.empty() || new_root == "/") {
|
||||
log_status(LogStatus::Failed,
|
||||
"switch_root: refusing to switch into '/' (already the root)");
|
||||
return false;
|
||||
}
|
||||
|
||||
struct stat st;
|
||||
if (::stat(new_root.c_str(), &st) != 0 || !S_ISDIR(st.st_mode)) {
|
||||
log_status(LogStatus::Failed,
|
||||
"switch_root: " + new_root + " is not a directory: " +
|
||||
std::strerror(errno));
|
||||
return false;
|
||||
}
|
||||
|
||||
if (::mount(new_root.c_str(), new_root.c_str(), "bind", MS_BIND, nullptr) != 0) {
|
||||
log_status(LogStatus::Failed,
|
||||
"switch_root: bind " + new_root + ": " + std::strerror(errno));
|
||||
return false;
|
||||
}
|
||||
|
||||
const auto move_subtree = [&](const char* src) {
|
||||
const std::string dst = new_root + src;
|
||||
if (::mount(src, dst.c_str(), nullptr, MS_MOVE, nullptr) != 0) {
|
||||
if (errno != EINVAL && errno != ENOENT && errno != EBUSY) {
|
||||
log_info(kTag, "switch_root: move ", src, " -> ", dst, ": ",
|
||||
std::strerror(errno));
|
||||
}
|
||||
}
|
||||
};
|
||||
move_subtree("/dev");
|
||||
move_subtree("/proc");
|
||||
move_subtree("/sys");
|
||||
|
||||
if (::chdir(new_root.c_str()) != 0) {
|
||||
log_status(LogStatus::Failed,
|
||||
"switch_root: chdir " + new_root + ": " + std::strerror(errno));
|
||||
return false;
|
||||
}
|
||||
|
||||
if (::mkdir(".initrd", 0755) != 0 && errno != EEXIST) {
|
||||
log_status(LogStatus::Failed,
|
||||
"switch_root: mkdir .initrd: " +
|
||||
std::string(std::strerror(errno)));
|
||||
return false;
|
||||
}
|
||||
if (::syscall(SYS_pivot_root, ".", ".initrd") != 0) {
|
||||
log_status(LogStatus::Failed,
|
||||
"switch_root: pivot_root: " +
|
||||
std::string(std::strerror(errno)));
|
||||
return false;
|
||||
}
|
||||
if (::umount2("/initrd", MNT_DETACH) != 0 && errno != EINVAL) {
|
||||
log_info(kTag, "switch_root: detach old root: ", std::strerror(errno));
|
||||
}
|
||||
|
||||
if (::chdir("/") != 0) {
|
||||
log_status(LogStatus::Failed,
|
||||
"switch_root: chdir /: " + std::string(std::strerror(errno)));
|
||||
::abort();
|
||||
}
|
||||
|
||||
std::vector<char*> argv;
|
||||
for (auto& a : init_argv)
|
||||
argv.push_back(const_cast<char*>(a.c_str()));
|
||||
argv.push_back(nullptr);
|
||||
::execv(argv[0], argv.data());
|
||||
|
||||
log_status(LogStatus::Failed,
|
||||
"switch_root: exec " + std::string(argv[0]) + ": " +
|
||||
std::strerror(errno));
|
||||
::abort();
|
||||
}
|
||||
|
||||
bool Supervisor::run_command(Command& cmd) {
|
||||
using K = Command::Kind;
|
||||
switch (cmd.kind) {
|
||||
|
|
@ -604,6 +859,25 @@ bool Supervisor::run_command(Command& cmd) {
|
|||
}
|
||||
return true;
|
||||
}
|
||||
case K::SwitchRoot:
|
||||
return run_switch_root(cmd);
|
||||
case K::Setprop:
|
||||
if (cmd.args.size() != 2) {
|
||||
log_status(LogStatus::Failed, "setprop requires KEY and VALUE");
|
||||
return false;
|
||||
}
|
||||
set_property(cmd.args[0], cmd.args[1]);
|
||||
return true;
|
||||
case K::Getprop: {
|
||||
if (cmd.args.size() != 1) {
|
||||
log_status(LogStatus::Failed, "getprop requires KEY");
|
||||
return false;
|
||||
}
|
||||
const auto it = properties_.find(cmd.args[0]);
|
||||
log_info(kTag, "getprop ", cmd.args[0], "=",
|
||||
it != properties_.end() ? it->second : "");
|
||||
return true;
|
||||
}
|
||||
case K::Log:
|
||||
log_info(kTag, "action: ", join(cmd.args, " "));
|
||||
return true;
|
||||
|
|
@ -628,8 +902,24 @@ void Supervisor::execute_action(Action& action) {
|
|||
}
|
||||
}
|
||||
|
||||
// The console fd, opened once PID1 realizes it's on a real console. Provided
|
||||
// so spawn_service can rebind stdio for services flagged `console`.
|
||||
void Supervisor::set_property(const std::string& key, const std::string& value) {
|
||||
properties_[key] = value;
|
||||
log_info(kTag, "setprop ", key, "=", value);
|
||||
if (property_dispatch_depth_ >= kMaxPropertyDispatchDepth) {
|
||||
log_info(kTag, "setprop: property-trigger recursion limit reached (", key,
|
||||
"=", value, ")");
|
||||
return;
|
||||
}
|
||||
// fire every action whose trigger matches property:<key>=<value>.
|
||||
++property_dispatch_depth_;
|
||||
const std::string want = "property:" + key + "=" + value;
|
||||
for (auto& action : config_.actions) {
|
||||
if (action.trigger == want)
|
||||
execute_action(action);
|
||||
}
|
||||
--property_dispatch_depth_;
|
||||
}
|
||||
|
||||
void Supervisor::open_console() {
|
||||
if (g_open_console_fd >= 0)
|
||||
return;
|
||||
|
|
@ -803,6 +1093,17 @@ std::string Supervisor::ctl_execute(const std::string& line) {
|
|||
begin_shutdown(kind);
|
||||
return "OK\n";
|
||||
}
|
||||
if (cmd == "setprop" && toks.size() >= 3) {
|
||||
set_property(toks[1], toks[2]);
|
||||
return "OK\n";
|
||||
}
|
||||
if (cmd == "getprop") {
|
||||
if (toks.size() >= 2) {
|
||||
const auto it = properties_.find(toks[1]);
|
||||
return "OK " + (it != properties_.end() ? it->second : "") + "\n";
|
||||
}
|
||||
return "ERR getprop requires KEY\n";
|
||||
}
|
||||
if (cmd == "reload") {
|
||||
reload_config();
|
||||
return "OK\n";
|
||||
|
|
|
|||
|
|
@ -7,6 +7,7 @@
|
|||
#include "config.hpp"
|
||||
|
||||
#include <chrono>
|
||||
#include <map>
|
||||
#include <string>
|
||||
#include <unordered_map>
|
||||
|
||||
|
|
@ -57,12 +58,29 @@ class Supervisor {
|
|||
ShutdownState shutdown_state_ = ShutdownState::Running;
|
||||
std::chrono::steady_clock::time_point shutdown_deadline_{};
|
||||
|
||||
// property store; `set_property` fires `on property:K=V` actions.
|
||||
std::map<std::string, std::string> properties_;
|
||||
int property_dispatch_depth_ = 0; // guard against trigger recursion
|
||||
|
||||
// services waiting for their wait-for dependencies to report ready.
|
||||
std::vector<std::string> pending_wait_;
|
||||
|
||||
void setup_signals();
|
||||
void open_console();
|
||||
void spawn_service(Service& svc, bool missing_ok);
|
||||
void reap_children();
|
||||
void execute_action(Action& action);
|
||||
void run_exec_command(const Command& cmd);
|
||||
// set a property and fire any `on property:K=V` action that now matches.
|
||||
void set_property(const std::string& key, const std::string& value);
|
||||
// mark a service ready and respawn any wait-for dependents now unblocked.
|
||||
void mark_ready(const std::string& name);
|
||||
// spawn pending services whose dependencies are all ready; drop those
|
||||
// whose dependency died or vanished. Called after a ready or a reap.
|
||||
void reevaluate_pending();
|
||||
// leave the initramfs and hand control to the real init (2nd stage).
|
||||
// returns false on error; on success it does not return.
|
||||
bool run_switch_root(const Command& cmd);
|
||||
bool run_command(Command& cmd);
|
||||
// re-parse the .rc files and reconcile live services: removed services are
|
||||
// stopped, added ones registered, changed ones restarted. Called from
|
||||
|
|
|
|||
|
|
@ -4,6 +4,7 @@
|
|||
#include "../src/supervisor.cpp"
|
||||
|
||||
#include <cstdio>
|
||||
#include <chrono>
|
||||
#include <map>
|
||||
#include <memory>
|
||||
#include <sstream>
|
||||
|
|
@ -166,6 +167,58 @@ void test_parse_basic() {
|
|||
"unknown action command");
|
||||
// unknown directive at column 0
|
||||
CHECK_THROWS(parse_string("BROKEN = yes\n", fs), "unexpected directive");
|
||||
|
||||
// switch_root parses as a SwitchRoot command
|
||||
auto sr = parse_string("on boot\n switch_root /mnt/root /sbin/init --stage2\n",
|
||||
fs);
|
||||
CHECK_EQ(sr.actions.size(), 1u);
|
||||
CHECK_EQ(sr.actions[0].commands.size(), 1u);
|
||||
CHECK(sr.actions[0].commands[0].kind == Command::Kind::SwitchRoot);
|
||||
CHECK_EQ(sr.actions[0].commands[0].args.size(), 3u);
|
||||
CHECK_EQ(sr.actions[0].commands[0].args[0], std::string("/mnt/root"));
|
||||
CHECK_EQ(sr.actions[0].commands[0].args[1], std::string("/sbin/init"));
|
||||
CHECK_EQ(sr.actions[0].commands[0].args[2], std::string("--stage2"));
|
||||
|
||||
// `depends = A B ...` is a repeatable, multi-valued service option
|
||||
auto dp = parse_string(
|
||||
"service web /usr/sbin/httpd\n"
|
||||
" depends = net\n"
|
||||
" depends = db cache\n",
|
||||
fs);
|
||||
CHECK_EQ(dp.services.size(), 1u);
|
||||
CHECK_EQ(dp.services[0].depends.size(), 3u);
|
||||
CHECK_EQ(dp.services[0].depends[0], std::string("net"));
|
||||
CHECK_EQ(dp.services[0].depends[1], std::string("db"));
|
||||
CHECK_EQ(dp.services[0].depends[2], std::string("cache"));
|
||||
|
||||
// setprop/getprop action commands and `on property:` triggers
|
||||
auto pp = parse_string(
|
||||
"on boot\n"
|
||||
" setprop net.up 1\n"
|
||||
" getprop net.up\n"
|
||||
"on property:net.up=1\n"
|
||||
" log network is up\n",
|
||||
fs);
|
||||
CHECK_EQ(pp.actions.size(), 2u);
|
||||
// property trigger string is captured verbatim
|
||||
CHECK_EQ(pp.actions[1].trigger, std::string("property:net.up=1"));
|
||||
CHECK_EQ(pp.actions[0].commands[0].kind, Command::Kind::Setprop);
|
||||
CHECK_EQ(pp.actions[0].commands[0].args.size(), 2u);
|
||||
CHECK_EQ(pp.actions[0].commands[0].args[0], std::string("net.up"));
|
||||
CHECK_EQ(pp.actions[0].commands[0].args[1], std::string("1"));
|
||||
CHECK(pp.actions[0].commands[1].kind == Command::Kind::Getprop);
|
||||
|
||||
// `logfile = PATH` is an optional service option
|
||||
auto lf = parse_string(
|
||||
"service daemon /usr/sbin/daemon\n"
|
||||
" logfile = /var/log/daemon.log\n",
|
||||
fs);
|
||||
CHECK_EQ(lf.services.size(), 1u);
|
||||
CHECK_EQ(lf.services[0].logfile, std::string("/var/log/daemon.log"));
|
||||
auto nl = parse_string(
|
||||
"service plain /usr/sbin/plain\n",
|
||||
fs);
|
||||
CHECK_EQ(nl.services[0].logfile, std::string(""));
|
||||
}
|
||||
|
||||
void test_imports() {
|
||||
|
|
@ -262,6 +315,126 @@ void test_supervisor_helpers() {
|
|||
std::string("web running pid 42 oneshot class tools\n"));
|
||||
}
|
||||
|
||||
void test_crash_window() {
|
||||
// a service that crashes repeatedly inside the window gets throttled at
|
||||
// the threshold, then recovers once the window elapses.
|
||||
using Clock = std::chrono::steady_clock;
|
||||
using namespace std::chrono;
|
||||
|
||||
Service svc;
|
||||
svc.name = "boom";
|
||||
svc.crash_threshold = 2;
|
||||
svc.crash_window_secs = 10;
|
||||
|
||||
auto t = Clock::now();
|
||||
// first two crashes within the window are allowed
|
||||
CHECK(record_crash(svc, t) == CrashAction::Respawn);
|
||||
CHECK_EQ(svc.crash_count, 1);
|
||||
CHECK(!svc.throttled);
|
||||
CHECK(record_crash(svc, t + seconds(1)) == CrashAction::Respawn);
|
||||
CHECK_EQ(svc.crash_count, 2);
|
||||
CHECK(!svc.throttled);
|
||||
// third crash crosses the threshold -> throttle
|
||||
CHECK(record_crash(svc, t + seconds(2)) == CrashAction::Throttle);
|
||||
CHECK_EQ(svc.crash_count, 3);
|
||||
CHECK(svc.throttled);
|
||||
|
||||
// Once a fresh window opens (the crash was long after the last one), the
|
||||
// count resets but `throttled` stays set until an explicit start -- the
|
||||
// reaper stops respawning a throttled service, so the only recovery is a
|
||||
// manual `start`/`restart` (reset_crash_state below).
|
||||
CHECK(record_crash(svc, t + seconds(11)) == CrashAction::Respawn);
|
||||
CHECK_EQ(svc.crash_count, 1);
|
||||
CHECK(svc.throttled);
|
||||
|
||||
// a manual start resets the throttle, opening a fresh window.
|
||||
reset_crash_state(svc);
|
||||
CHECK_EQ(svc.crash_count, 0);
|
||||
CHECK(!svc.throttled);
|
||||
|
||||
// threshold=1 means a single crash then throttle.
|
||||
Service one;
|
||||
one.crash_threshold = 1;
|
||||
one.crash_window_secs = 10;
|
||||
CHECK(record_crash(one, t) == CrashAction::Respawn);
|
||||
CHECK(record_crash(one, t + seconds(1)) == CrashAction::Throttle);
|
||||
|
||||
// a degenerate (unset) window never throttles.
|
||||
Service none;
|
||||
none.crash_threshold = 0;
|
||||
none.crash_window_secs = 0;
|
||||
for (int i = 0; i < 100; ++i)
|
||||
CHECK(record_crash(none, t + seconds(i)) == CrashAction::Respawn);
|
||||
CHECK_EQ(none.crash_count, 0);
|
||||
CHECK(!none.throttled);
|
||||
}
|
||||
|
||||
void test_dependencies() {
|
||||
Config cfg;
|
||||
auto add = [&](const std::string& name, std::initializer_list<const char*> deps) {
|
||||
Service s;
|
||||
s.name = name;
|
||||
for (auto* d : deps)
|
||||
s.depends.emplace_back(d);
|
||||
cfg.services.push_back(std::move(s));
|
||||
};
|
||||
add("a", {});
|
||||
add("b", {"a"});
|
||||
add("c", {"b"});
|
||||
add("d", {"b", "c"}); // depends on a chain + a sibling
|
||||
add("e", {});
|
||||
|
||||
// linear chain: c -> b -> a
|
||||
std::vector<std::string> order;
|
||||
CHECK(resolve_dependencies(cfg, "c", order) == DepResolve::Ok);
|
||||
CHECK_EQ(order.size(), 3u);
|
||||
CHECK(order[0] == "a");
|
||||
CHECK(order[1] == "b");
|
||||
CHECK(order[2] == "c");
|
||||
|
||||
// declared order is preserved; shared deps appear once (diamond / DAG)
|
||||
order.clear();
|
||||
CHECK(resolve_dependencies(cfg, "d", order) == DepResolve::Ok);
|
||||
CHECK_EQ(order.size(), 4u);
|
||||
CHECK(order[0] == "a"); // b's dep
|
||||
CHECK(order[1] == "b"); // first declared dep of d
|
||||
CHECK(order[2] == "c"); // second declared dep of d (and c->b already done)
|
||||
CHECK(order[3] == "d");
|
||||
|
||||
// a service with no deps yields just itself
|
||||
order.clear();
|
||||
CHECK(resolve_dependencies(cfg, "e", order) == DepResolve::Ok);
|
||||
CHECK_EQ(order.size(), 1u);
|
||||
CHECK(order[0] == "e");
|
||||
|
||||
// unknown dependency
|
||||
Service ghost;
|
||||
ghost.name = "ghost";
|
||||
ghost.depends.push_back("nope");
|
||||
cfg.services.push_back(std::move(ghost));
|
||||
order.clear();
|
||||
CHECK(resolve_dependencies(cfg, "ghost", order) == DepResolve::Unknown);
|
||||
|
||||
// self-cycle
|
||||
Service self;
|
||||
self.name = "self";
|
||||
self.depends.push_back("self");
|
||||
cfg.services.push_back(std::move(self));
|
||||
order.clear();
|
||||
CHECK(resolve_dependencies(cfg, "self", order) == DepResolve::Cycle);
|
||||
|
||||
// mutual cycle a<->b
|
||||
Service m1, m2;
|
||||
m1.name = "m1";
|
||||
m1.depends.push_back("m2");
|
||||
m2.name = "m2";
|
||||
m2.depends.push_back("m1");
|
||||
cfg.services.push_back(std::move(m1));
|
||||
cfg.services.push_back(std::move(m2));
|
||||
order.clear();
|
||||
CHECK(resolve_dependencies(cfg, "m1", order) == DepResolve::Cycle);
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
int main() {
|
||||
|
|
@ -271,6 +444,8 @@ int main() {
|
|||
test_parse_basic();
|
||||
test_imports();
|
||||
test_supervisor_helpers();
|
||||
test_crash_window();
|
||||
test_dependencies();
|
||||
|
||||
std::printf("%d checks, %d failures\n", g_checks, g_failures);
|
||||
return g_failures == 0 ? 0 : 1;
|
||||
|
|
|
|||
166
tools/run_vm.py
166
tools/run_vm.py
|
|
@ -272,15 +272,19 @@ on boot
|
|||
start selinux-probe
|
||||
"""
|
||||
|
||||
def build_initramfs(init: Path, busybox: Path, rc_text: str, root_password: str | None,
|
||||
selinux: bool, keep: bool,
|
||||
bundles: list[tuple[str, Path]] | None = None) -> Path:
|
||||
if not shutil.which("cpio"):
|
||||
sys.exit("cpio not found (install cpio)")
|
||||
root = Path(tempfile.mkdtemp(prefix="bajia-root-"))
|
||||
try:
|
||||
for sub in ("etc/bajia", "bin", "sbin", "usr/sbin", "usr/bin",
|
||||
"dev", "proc", "sys", "run", "tmp"):
|
||||
# subdirectories laid out in every staging root tree (initramfs and disk root).
|
||||
COMMON_SUBDIRS = ("etc/bajia", "bin", "sbin", "usr/sbin", "usr/bin",
|
||||
"dev", "proc", "sys", "run", "tmp", "mnt")
|
||||
|
||||
def stage_root_tree(root: Path, init: Path, init_rel: str, busybox: Path,
|
||||
rc_text: str, rc_rel: str, root_password: str | None,
|
||||
selinux: bool, bundles: list[tuple[str, Path]]) -> None:
|
||||
"""Lay out the common bajia + busybox tree into `root`.
|
||||
|
||||
`init_rel` is where bajia lands ('init' for the stage-1 initramfs,
|
||||
'sbin/init' for the stage-2 root disk); `rc_rel` is where init.rc lands.
|
||||
The core payload is identical for both stages."""
|
||||
for sub in COMMON_SUBDIRS:
|
||||
(root / sub).mkdir(parents=True)
|
||||
|
||||
if selinux:
|
||||
|
|
@ -292,15 +296,17 @@ def build_initramfs(init: Path, busybox: Path, rc_text: str, root_password: str
|
|||
selcfg.write_text(f"SELINUX=permissive\nSELINUXTYPE={host_selinux_type()}\n")
|
||||
bundle_dynamic_libs(init, root)
|
||||
elif not is_static(init):
|
||||
print("warning: bajia is dynamically linked; /init will fail to exec "
|
||||
print("warning: bajia is dynamically linked; init will fail to exec "
|
||||
"inside the initramfs (error -2). Rebuild with --no-static "
|
||||
"unset (static is the default) or drop --no-build.")
|
||||
|
||||
if selinux:
|
||||
rc_text = rc_text + SELINUX_RC_PROBE
|
||||
|
||||
shutil.copy(init, root / "init")
|
||||
(root / "init").chmod(0o755)
|
||||
dst = root / init_rel
|
||||
dst.parent.mkdir(parents=True, exist_ok=True)
|
||||
shutil.copy(init, dst)
|
||||
dst.chmod(0o755)
|
||||
|
||||
ctl = BUILD / "bctl"
|
||||
if ctl.is_file():
|
||||
|
|
@ -312,7 +318,10 @@ def build_initramfs(init: Path, busybox: Path, rc_text: str, root_password: str
|
|||
for applet in BUSYBOX_APPLETS:
|
||||
(root / "bin" / applet).symlink_to("busybox")
|
||||
|
||||
(root / "etc" / "bajia" / "init.rc").write_text(rc_text)
|
||||
rcdst = root / rc_rel
|
||||
rcdst.parent.mkdir(parents=True, exist_ok=True)
|
||||
rcdst.write_text(rc_text)
|
||||
|
||||
(root / "etc" / "passwd").write_text(
|
||||
root_passwd_line(root_password) +
|
||||
"nobody:x:65534:65534:nobody:/:/bin/sh\n")
|
||||
|
|
@ -327,28 +336,104 @@ def build_initramfs(init: Path, busybox: Path, rc_text: str, root_password: str
|
|||
shutil.copy(src, dst)
|
||||
dst.chmod(0o644)
|
||||
|
||||
# The staging tree is owned by the host user and mkdtemp makes the
|
||||
# top dir 0700; GNU cpio preserves both, so without this the guest's
|
||||
# "/" would be mode 0700 owned by uid 1000 -- fine for root services,
|
||||
# but dropped-privilege services couldn't traverse it. Stamp owner
|
||||
# root:root (no host chown needed) and make the root traversable.
|
||||
# traversable by dropped-privilege services; the packers stamp ownership.
|
||||
root.chmod(0o755)
|
||||
|
||||
def pack_cpio(root: Path, keep: bool) -> Path:
|
||||
if not shutil.which("cpio"):
|
||||
sys.exit("cpio not found (install cpio)")
|
||||
# The staging tree is owned by the host user; --owner=0:0 and a traversable
|
||||
# root make every path root-owned inside the guest.
|
||||
p = run(["bash", "-c",
|
||||
"cd \"$1\" && chmod 0755 . && "
|
||||
"find . -print0 | cpio --null -o -H newc --owner=0:0",
|
||||
"cd \"$1\" && find . -print0 | cpio --null -o -H newc --owner=0:0",
|
||||
"bajia-initramfs", str(root)], stdout=subprocess.PIPE)
|
||||
if p.returncode != 0:
|
||||
sys.exit("cpio packing failed")
|
||||
initrd = Path(tempfile.gettempdir()) / "bajia-initrd.cpio.gz"
|
||||
initrd.write_bytes(gzip.compress(p.stdout))
|
||||
|
||||
if keep:
|
||||
print("initramfs root tree kept at:", root)
|
||||
return initrd
|
||||
|
||||
def pack_ext4(root: Path, keep: bool, size_mb: int = 128) -> Path:
|
||||
"""Pack `root` into a writable ext4 disk image via `mkfs.ext4 -d`.
|
||||
No loop mount needed, so it works unprivileged. The image is the second
|
||||
stage's real root filesystem, attached to the guest as a virtio-blk disk."""
|
||||
if not shutil.which("mkfs.ext4"):
|
||||
sys.exit("mkfs.ext4 not found (install e2fsprogs)")
|
||||
img = Path(tempfile.gettempdir()) / "bajia-stage2.ext4"
|
||||
if img.is_file():
|
||||
img.unlink()
|
||||
# mkfs.ext4 -d does not create the image file; pre-size a sparse file.
|
||||
r = run(["truncate", "-s", f"{size_mb}M", str(img)])
|
||||
if r.returncode != 0:
|
||||
sys.exit("truncate failed")
|
||||
r = run(["mkfs.ext4", "-q", "-d", str(root), "-F", str(img)])
|
||||
if r.returncode != 0:
|
||||
sys.exit("mkfs.ext4 failed")
|
||||
if keep:
|
||||
print("stage2 root tree kept at:", root)
|
||||
return img
|
||||
|
||||
def build_initramfs(init: Path, busybox: Path, rc_text: str,
|
||||
root_password: str | None, selinux: bool, keep: bool,
|
||||
bundles: list[tuple[str, Path]] | None = None) -> Path:
|
||||
root = Path(tempfile.mkdtemp(prefix="bajia-root-"))
|
||||
try:
|
||||
stage_root_tree(root, init, "init", busybox, rc_text,
|
||||
"etc/bajia/init.rc", root_password, selinux, bundles)
|
||||
return pack_cpio(root, keep)
|
||||
finally:
|
||||
if not keep:
|
||||
shutil.rmtree(root, ignore_errors=True)
|
||||
|
||||
def qemu_command(kernel: Path, initrd: Path, args: argparse.Namespace) -> list[str]:
|
||||
# first-stage config for --two-stage: bring up the basics, mount the real root
|
||||
# disk, and switch_root onto it. `root_dev` is the virtio-blk target (/dev/vda).
|
||||
TWO_STAGE_FIRST_RC = """\
|
||||
# generated by tools/run_vm.py --two-stage - first stage (initramfs).
|
||||
on early-init
|
||||
mount proc /proc proc
|
||||
mount sysfs /sys sysfs
|
||||
mount devtmpfs /dev devtmpfs
|
||||
|
||||
on init
|
||||
mkdir /mnt/root 0755
|
||||
mount {dev} /mnt/root {fstype}
|
||||
|
||||
on boot
|
||||
switch_root /mnt/root /sbin/init /etc/bajia/init.rc
|
||||
"""
|
||||
|
||||
def build_two_stage(init: Path, busybox: Path, second_rc: str,
|
||||
root_password: str | None, selinux: bool, keep: bool,
|
||||
bundles: list[tuple[str, Path]],
|
||||
root_dev: str = "/dev/vda",
|
||||
root_fstype: str = "ext4",
|
||||
root_size_mb: int = 128) -> tuple[Path, Path]:
|
||||
"""Return (stage1_initrd, stage2_root_img) for a classic two-stage boot.
|
||||
|
||||
stage1 is a minimal initramfs running bajia as /init with a generated
|
||||
first-stage config that mounts the real root and switch_roots onto it.
|
||||
stage2 is a writable ext4 disk image (the "real root"), bundled with its
|
||||
own copy of bajia at /sbin/init plus the full init.rc and services."""
|
||||
first_rc = TWO_STAGE_FIRST_RC.format(dev=root_dev, fstype=root_fstype)
|
||||
root1 = Path(tempfile.mkdtemp(prefix="bajia-stage1-"))
|
||||
root2 = Path(tempfile.mkdtemp(prefix="bajia-stage2-"))
|
||||
try:
|
||||
stage_root_tree(root1, init, "init", busybox, first_rc,
|
||||
"etc/bajia/init.rc", root_password, selinux, [])
|
||||
initrd = pack_cpio(root1, keep)
|
||||
stage_root_tree(root2, init, "sbin/init", busybox, second_rc,
|
||||
"etc/bajia/init.rc", root_password, selinux, bundles)
|
||||
img = pack_ext4(root2, keep, size_mb=root_size_mb)
|
||||
return initrd, img
|
||||
finally:
|
||||
if not keep:
|
||||
shutil.rmtree(root1, ignore_errors=True)
|
||||
shutil.rmtree(root2, ignore_errors=True)
|
||||
|
||||
def qemu_command(kernel: Path, initrd: Path, args: argparse.Namespace,
|
||||
root_img: Path | None = None) -> list[str]:
|
||||
qemu = args.qemu or shutil.which("qemu-system-x86_64") or "qemu-system-x86_64"
|
||||
display = args.display
|
||||
if display is None:
|
||||
|
|
@ -367,6 +452,9 @@ def qemu_command(kernel: Path, initrd: Path, args: argparse.Namespace) -> list[s
|
|||
"-display", display,
|
||||
"-serial", "stdio",
|
||||
]
|
||||
if root_img is not None:
|
||||
# second-stage root disk; virtio-blk (built into modern kernels) -> /dev/vda
|
||||
cmd += ["-drive", f"file={root_img},format=raw,if=virtio"]
|
||||
if args.nographic:
|
||||
cmd[cmd.index("-display") + 1] = "none"
|
||||
if args.serial_log:
|
||||
|
|
@ -386,7 +474,22 @@ def main() -> int:
|
|||
help="download busybox source (github.com/mirror/busybox), "
|
||||
"build it static, cache in ~/.cache/bajia")
|
||||
ap.add_argument("--config", type=Path,
|
||||
help="use this init.rc instead of the bundled test config")
|
||||
help="use this init.rc instead of the bundled test config "
|
||||
"(with --two-stage this is the *second stage* config)")
|
||||
ap.add_argument("--two-stage", action="store_true",
|
||||
help="boot a classic two-stage initramfs: a minimal stage-1 "
|
||||
"initramfs runs bajia as /init, mounts a real root disk "
|
||||
"and switch_roots onto it; stage-2 is a writable ext4 "
|
||||
"root running bajia from /sbin/init")
|
||||
ap.add_argument("--root-dev", default="/dev/vda",
|
||||
help="--two-stage: block device for the real root "
|
||||
"(default /dev/vda)")
|
||||
ap.add_argument("--root-fstype", default="ext4",
|
||||
help="--two-stage: filesystem type of the real root "
|
||||
"(default ext4)")
|
||||
ap.add_argument("--root-size", type=int, default=128,
|
||||
help="--two-stage: stage-2 root image size in MiB "
|
||||
"(default 128)")
|
||||
ap.add_argument("--bundle", action="append", default=[],
|
||||
metavar="REL=HOSTPATH",
|
||||
help="copy HOSTPATH into the initramfs at absolute REL "
|
||||
|
|
@ -458,12 +561,27 @@ def main() -> int:
|
|||
if ".." in [c for c in Path(rel).parts]:
|
||||
sys.exit(f"--bundle: REL must not contain '..': {rel}")
|
||||
bundles.append((rel, src))
|
||||
|
||||
if args.two_stage:
|
||||
if args.selinux:
|
||||
print("note: SELinux is bundled into both stages' roots")
|
||||
initrd, root_img = build_two_stage(
|
||||
init, busybox, rc_text, args.root_password,
|
||||
selinux=args.selinux, keep=args.keep_initramfs, bundles=bundles,
|
||||
root_dev=args.root_dev, root_fstype=args.root_fstype,
|
||||
root_size_mb=args.root_size)
|
||||
print("stage1 initramfs:", initrd,
|
||||
f"({initrd.stat().st_size / 1024:.0f} KB)")
|
||||
print("stage2 root disk:", root_img,
|
||||
f"({root_img.stat().st_size / 1024:.0f} KB)")
|
||||
else:
|
||||
initrd = build_initramfs(init, busybox, rc_text, args.root_password,
|
||||
selinux=args.selinux, keep=args.keep_initramfs,
|
||||
bundles=bundles)
|
||||
root_img = None
|
||||
print("initramfs:", initrd, f"({initrd.stat().st_size / 1024:.0f} KB)")
|
||||
|
||||
cmd = qemu_command(kernel, initrd, args)
|
||||
cmd = qemu_command(kernel, initrd, args, root_img)
|
||||
print("$", " ".join(cmd))
|
||||
return subprocess.run(cmd).returncode
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue