Upgrade Guide: v0.9.0 to v0.9.1¶
v0.9.1 is a fix release. There are no config changes and no wire changes, so mixed client and server versions keep working. What it fixes:
- #410 Root-owned
~/.localand~/.configin sheds with nested mounts. Existing sheds are repaired on their next start. - #284 Firecracker stop and delete latency with writable mounts.
- #409 VZ sheds no longer stop when the brew service restarts.
- #412 The first CLI command after a token-mode server restart no longer fails.
On macOS, the upgrade stops every running VZ shed once. Read Before you upgrade (macOS / VZ) first.
Before you upgrade (macOS / VZ)¶
This upgrade stops every running VZ shed, once
Upgrading a Mac from 0.9.0 to 0.9.1 stops all of its running VZ sheds. Nothing is lost on disk, but anything running in them ends.
Why: the 0.9.0 shed-server started vfkit in the launchd job's process group, and launchd
kills the whole group when the job exits. The brew services restart shed that applies 0.9.1
stops the old 0.9.0 server, so it still takes the VMs down. From 0.9.1 on, vfkit has its
own process group and restarts leave VMs running.
To control it, stop the sheds yourself first and start them afterwards:
shed stop <name> # for each running shed, before the upgrade
# ... upgrade the host (see below) ...
shed start <name> # afterwards
Two more things to know:
- The brew service needs Homebrew 6.0.13 or later for its new
stop_timeout. Runbrew updatefirst; if you have disabled auto-update, update explicitly. - A restart can now take up to about 30 s. The server shuts down gracefully, and connected desktops or SSH sessions hold the drain open. A faster drain is a follow-up.
Upgrade steps¶
macOS (Homebrew)¶
Linux server (apt)¶
Running Firecracker sheds keep running across this restart, as they have since 0.9.0.
Desktop: if you removed shed-host-agent¶
0.9.1 does not change the desktop app; this is a docs fix. In the default Automatic broker
mode the app picks between a running shed-host-agent daemon and its own in-process broker
once, at launch. If it chose the daemon and the daemon is later stopped or uninstalled, the
app does not fall back, and sheds lose ssh-agent and credential forwarding (ssh-add -l
reports "agent refused operation") until the app is relaunched. After removing the daemon,
quit and relaunch the app. A runtime fallback is tracked in
#411.
What changed¶
VZ VMs outlive the server (#409)¶
vfkit now runs in its own process group, so launchd no longer kills it when the
shed-server job exits. The brew service also sets stop_timeout 35, so launchd waits for the
server's graceful shutdown instead of sending SIGKILL partway through.
The trade-offs:
| Situation | Behavior from 0.9.1 |
|---|---|
brew services restart shed |
VMs keep running; the new server re-adopts them. |
brew services stop shed |
VMs keep running. Run shed stop first if you want them stopped. |
| Server cannot come back | VMs run unmanaged until it does. The next start re-adopts them (resumed message channel in the log). |
Ctrl-C on a foreground dev shed-server |
VMs are no longer stopped. |
Firecracker stop and delete (#284)¶
| Operation (running shed, writable mounts) | Before | 0.9.1 |
|---|---|---|
shed delete |
about 25 s | about 0.2 s |
shed stop |
25-30 s, then force-kill | about stop_timeout (10 s) |
- The shutdown hook and the pre-stop sync now run with mounts still live.
- On both backends, a client disconnect during
shed stopno longer turns the stop into a force-kill. shed deletenow refuses, and keeps the shed's resources, if its VM survives SIGKILL. Retry once the VM is gone.
Known limits:
- A graceful Firecracker stop still ends at
stop_timeout. The guest cannot receive Ctrl-Alt-Del; a fix is a follow-up. The hook and sync still run before the kill. - After a Firecracker server restart, running sheds' 9P mounts stay dead until
shed stop && shed start(#369).
Root-owned ~/.local and ~/.config (#410)¶
Directories a mount needs under the shed user's home are now created as shed. Existing
sheds are repaired on their next shed start.
A root-owned directory left by a mount you have since removed from config is not on any mount path, so it is not repaired. Fix it by hand inside the shed:
Add any other affected parents (for example ~/.local/share) to the command. Images built
from 0.9.1 also pre-create ~/.local, ~/.local/bin, ~/.local/share, ~/.local/state,
~/.config and ~/.cache owned by shed.
CLI re-mint (#412)¶
shed start, shed create, shed image pull, shed image push, shed snapshot create and
the prune commands now re-mint and retry once on an expired or invalid credential. The first
command after a token-mode server restart no longer fails with a 401.