Of sandboxes, filesystems, and failovers

· Originally published on X

Of sandboxes, filesystems, and failovers

In the last few months of building @mistledev, I’ve done a bunch of sandbox shopping and learnt a little bit more about what makes a good sandbox (for our use case).

Our requirements

Admittedly, these requirements have also been adjusted based on the limits of today’s providers. I think the future should be multi-provider, similar to multi-cloud.

My key takeaway as of now is that the filesystem is the lock-in. This is a problem when providers run out of capacity, have instabilities in different areas, or any other reason that leads to the inability to start or resume sandboxes. Imagine a scenario where your user does some work in a sandbox and revisits the same session a few days later only to find out that the sandbox can’t be resumed.

Every provider will promise a great uptime and I believe they will all be able to achieve it at some point in time. However, it’s also our responsibility to ensure we are not subject to a single point of failure.

The naive solution to this is to implement a file/dir-walker that tars everything up and exports to an object storage. You’d have to do it in a way that doesn’t affect the user’s workloads too much.

You can optimize this further by only storing the diffs over time so you can reconstruct the state when you need to.

When you take such an approach, you’re forced to pick specific directories that you care about because it’s just not practical to do this from root. Given this constraint, you’re basically forcing the user to operate in specific known directories. Not great.

If we take a step back, how are sandbox providers taking snapshots? First, when we talk about sandboxes / VMs, there’s the host-side and guest-side. Host-side is where the sandbox provider has its services on their worker nodes (the servers that host the sandboxes/VMs). Guest-side is where we, as sandbox users, operate (essentially within the sandbox). On the host-side, the sandbox provider has their own services that handle communication, control, and management of the underlying sandboxes. When a snapshot is taken, the provider packs up the sandbox’s rootfs and puts it into an object storage. When you start a new sandbox from the snapshot, it takes the snapshot out from the object storage and boots up a new sandbox from this rootfs snapshot. They’re able to do this efficiently because they own the underlying block device (btrfs, ZFS, etc.) while guest-side relies on tools that leverage filesystem APIs (tar, find, rsync, etc.). These filesystems are generally copy-on-write (CoW) and because they have access to underlying block storage, they can take snapshots by simply writing references to the actual physical blocks, as opposed to what’s possible in the guest-side.

Now, what if we could become the “host” in the sandbox and our users become the “guest”? What I’m driving at here is nested virtualization - a sandbox in a sandbox. This is where things start to get overengineered.

I experimented with nesting Icarus containers inside different sandbox providers with varying degrees of success - bearing in mind now it’s the Icarus containers that need to fulfil the requirements above. If successful, I’d basically get the benefits of host-side control without managing underlying infrastructure (scheduling, ensuring capacity, and many other things). I didn’t go through with this in the end because it’s extremely complex for what we wanted to achieve. FWIW, I did some benchmarking here and was surprised to find that there were minimal overheads from taking this approach.

I also experimented with @archildata - mounting an Archil disk into the sandbox root and let everything go through that but as of all network-based solutions, even with colocation, the latency adds up across many small file-writes. I want to caveat here that I admittedly did not spend enough time digging through this and am happy to stand corrected on the right approach to implementing this.

Note: I did not try classic FUSE mounts because we already know that it’s definitively worse for these workloads.

I think that the holy grail here is decoupled compute from disk, where both feel infinite and elastic. I can then only reach for more when I need it, and give back some when I no longer need it. I’m not a systems engineer so please take my words with a pinch of salt, given I would not have a full understanding of limitations today.

At the end of the day, all I really want is a reliable experience for agents + computers and it looks like decoupling compute / disk seems like the way to go here.