Here is the configuration it takes, on FreeBSD, to put a web server, an application, a database and a mail relay on one machine where none of them can see the others:
path = "/usr/local/jails/containers/$name";
host.hostname = "$name";
mount.devfs;
exec.start = "/bin/sh /etc/rc";
exec.stop = "/bin/sh /etc/rc.shutdown";
web { ip4.addr = 192.0.2.10; }
app { ip4.addr = 192.0.2.11; }
db { ip4.addr = 192.0.2.12; }
mail { ip4.addr = 192.0.2.13; }
Then sysrc jail_enable=YES and service jail start. Each of the four directories is a ZFS clone of one prepared template, which cost no disk space on the day it was made, and nothing had to be installed on the host: the parser ships with the base system and the confinement lives in the kernel.
On a Linux host the same four services are, more often than not, four containers, and before the first of them starts somebody installs an engine to run them.
The difference goes back twenty-six years, to the question of where the fence should live.
What Kamp decided in 1999
The problem arrived from a hosting provider. Customers wanted root on a shared machine, and the provider wanted that root to reach nothing beyond the customer's own corner of it. chroot had never been meant to settle that argument. It moves a process's idea of /, and a root user inside one can still load a kernel module or reconfigure the network; the ways out had been catalogued for years.
Poul-Henning Kamp's answer, written up with Robert Watson in 2000 under the title "Jails: Confining the omnipotent root", gave the kernel one new object, the prison, attached to the credential of every process inside it. Wherever the kernel asks whether something is permitted or visible, it now asks the prison as well. The paper puts the size of the change at roughly two hundred lines spread across some fifty files, plus about two hundred in two new kernel files. Four hundred lines, more or less, and FreeBSD 4.0 shipped them on 14 March 2000.
The paper also lists what a jailed root loses, and it reads like a security review somebody actually finished: a jailed root may not load kernel modules, touch any of the network configuration, interfaces, addresses or routing table, mount filesystems, create device nodes, open raw sockets, change most kernel parameters, clear the system's protective file flags or reach network resources that belong to someone else.
The list has grown with the system and kept its shape. In FreeBSD 15.1 the kernel defines 244 privileges in priv.h. A default jail receives 60 of them and can be granted 17 more by name in jail.conf. Refusal is the default, and the function that implements it can be read before the tea goes cold (it is long, but it is one switch): every request it does not recognise falls through to one line at the bottom of prison_priv_check(), which returns EPERM.
Ordinary work does not slow down for any of this. The prison is consulted at two kinds of moment, when a process asks for a privilege and when the kernel decides which processes and addresses it may touch, and reading a file or writing to a connected socket is neither, so a jailed service runs the same kernel code it would run on the bare host.
What Linux built instead, and why
The Linux kernel is one project, and the systems built around it are many. That shaped the answer from the start. Nobody in the kernel was in a position to ship a finished isolation facility with its own configuration file and management tool, because the kernel ships neither. So it shipped primitives, one at a time, and left the composition to whoever came next.
The mount namespace arrived in 2.4.19 in 2002. UTS and IPC followed in 2.6.19, PID and network in 2.6.24. The user namespace was merged in 2.6.23 and became fully usable only in 3.8, in 2013. Control groups, contributed by engineers at Google, landed in 2.6.24 in 2008; the cgroup namespace came in 4.6 and the time namespace in 5.6, in 2020. That makes eight namespaces according to namespaces(7), and around them cgroups for resource limits, seccomp for filtering system calls, capabilities for splitting up root, and a Linux security module on top.
The primitives are useful on their own, which is why they spread: systemd sandboxes services with them and browsers isolate their renderers with them, neither of which needed a container. What none of them gave an administrator was a jail. LXC assembled one out of them in 2008, and in 2013 Docker wrapped the result in a tool a developer could pick up on a Monday morning, a job no kernel mailing list was ever going to take on.
The assembly still has to happen somewhere, and today it happens in userland. runc, split out of Docker in 2015, takes a description of a container and turns it into the right sequence of clone, unshare, mount, cgroup and seccomp calls. Docker puts containerd and dockerd above it; Podman calls it, or its sibling crun, directly and without a daemon. Either way the fence is put together each time a container starts, by a program, from parts that belong to different subsystems of the kernel.
The matrix
Here are five forms of isolation side by side, with the figures that matter and where each one comes from.
chroot jail container VM microVM
(Docker, Podman) (KVM, bhyve) (Firecracker)
what confines path only one kernel 8 namespaces, own kernel on own kernel on
object cgroups, seccomp, virtual hardware minimal hardware
capabilities, LSM
kernel shared w. host yes yes yes no no
assembled by kernel kernel runc or crun, hypervisor VMM in userland
at every start
daemon no no Docker yes, no one VMM per guest
Podman no
root inside real root 60 of 244 uid 0 by default guest root guest root
privileges (Docker)
runtime cost vs host none none none without NAT 2 to 17 % compute < 5 MB memory
(KVM) per guest
what ships a directory ZFS snapshot, image, volumes disk image kernel and
data included excluded root filesystem
part of the system 1979 March 2000 separate install bhyve, 2014 separate install
Every row comes from a source, and here they are.
Runtime cost. IBM Research measured this properly in 2014, on a server with two Xeon E5-2665 processors running Ubuntu 13.10, with Docker 1.0 and KVM against native Linux. Linpack came to 290.9 GFLOPS in the container against 290.8 native. Memory bandwidth was identical and block I/O showed no overhead. Network throughput was identical as well, under one condition the paper spells out: "Containers that do not use NAT have identical performance to native Linux." Docker's default networking does use NAT, and in the paper's latency test it doubled the round trip. KVM, untuned, gave up 17 per cent on Linpack and half its random I/O operations.
A jail sits in the same class as the container on this row, for the same reason: no hardware is emulated and no second kernel runs, so ordinary work pays no tax. What separates the two is on the network side. A jail takes an address on an existing interface, or a complete network stack of its own through VNET, and its traffic reaches the wire without being translated on the way.
What assembles it. Counted from the public repositories on 5 October 2026, code lines only, without blank lines, comments, tests or vendored code, runc 1.5.2 comes to 16,465 lines. Above it, containerd 2.4.1 and moby, the engine behind dockerd, add 99,505 and 142,765. On FreeBSD 15.1, kern_jail.c and its header come to 4,548 lines, and the userland pieces, jail, jexec, jls and libjail, to 5,329.
So runc alone, the program whose task is to put the fence up, is larger than FreeBSD's jail implementation together with its tools. The stack a Docker installation brings for one isolated service comes to 258,735 lines, twenty-six times as much. The FreeBSD figure includes the kernel file in which the prison is implemented. The container figure includes nothing from the Linux kernel, where the namespaces and cgroups themselves live, so the comparison is a generous one to the container.
The VM column. Klara Systems put bhyve on FreeBSD 13.1 against KVM on Ubuntu 22.04 in November 2022, with the same guests on the same hardware. Storage under bhyve came out four to eight times quicker, whilst compute moved by about a quarter in either direction depending on the guest. The tester concluded that the storage wins outweigh the compute losses. A FreeBSD host that does need a virtual machine finds a competitive one in its base system.
The microVM. Firecracker, which AWS built for AWS Lambda and described at NSDI 2020, gives every guest a kernel of its own at under five megabytes of overhead. It exists because AWS was not prepared to run strangers' code side by side in containers, and the paper names the reason: with containers, AWS found itself "trading off between security and compatibility".
Build once, ship a snapshot
What made Docker famous was the shipping. An environment is built once and packed into an image, and the image that passed the tests is the image that runs in production, on whichever machine is asked to run it. On FreeBSD that job belongs to the thin jail, and it rests on two features ZFS has carried in the base system since 2008.
The first is the template. Following the FreeBSD Handbook, it is a dataset holding the base system at its current patch level, frozen by a snapshot:
# zfs create -p zroot/jails/templates/15.1-RELEASE
# tar -xf base.txz -C /usr/local/jails/templates/15.1-RELEASE --unlink
# freebsd-update -b /usr/local/jails/templates/15.1-RELEASE fetch install
# zfs snapshot zroot/jails/templates/15.1-RELEASE@base
The second is the clone, and the ZFS manual prices it with some precision: creating one "is nearly instantaneous, and initially consumes no additional space". Every thin jail begins life as a clone of the template. Software goes in from the host, and a snapshot of the result becomes the image, with a version in its name:
# zfs clone zroot/jails/templates/15.1-RELEASE@base zroot/jails/containers/web
# service jail start web
# pkg -j web install -y nginx
# zfs snapshot zroot/jails/containers/web@v1
Shipping goes through a pipe. The template crosses once. After that each jail travels as its difference from the template and lands on the other side as a clone of the same snapshot, which is the arrangement Docker builds out of layers and a registry:
# zfs send zroot/jails/templates/15.1-RELEASE@base \
| ssh prod zfs recv zroot/jails/templates/15.1-RELEASE
# zfs send -i zroot/jails/templates/15.1-RELEASE@base zroot/jails/containers/web@v1 \
| ssh prod zfs recv -o origin=zroot/jails/templates/15.1-RELEASE@base zroot/jails/containers/web
On the receiving host the jail needs its entry in jail.conf, and service jail start web starts the environment that passed the tests, block for block. There is no further step. The next release travels as zfs send -i @v1 zroot/jails/containers/web@v2, which carries the blocks that changed and nothing else, and @v1 stays on both machines for the afternoon somebody wants it back.
Two properties of the stream matter more than its size. A zfs receive aborts when the stream fails its checksum, and started with -s it keeps what arrived and resumes from there after a broken connection. And a snapshot freezes a dataset at one instant, so the data travels through the same pipe as the jail, in the state it had at that moment. PostgreSQL's documentation accepts exactly that: a consistent snapshot of the data directory starts up like a server that crashed, and replays its write-ahead log. Docker's reference for docker commit states its own position just as plainly: "Commits do not include any data contained in mounted volumes."
Patching follows the path the image took. The Handbook offers two ways for cloned jails: update each one in place with freebsd-update -j, or re-create them from a freshly patched snapshot of the template, with their data kept on datasets of their own so that the re-creation is, in the Handbook's word, cheap. The second is Docker's rebuild-and-redeploy, with the registry replaced by a pipe.
For the jail that has to go its own way, FreeBSD keeps the older form. A thick jail carries a complete base system of its own, and the Handbook calls its advantage independence: a thick jail can carry its own versions of everything, apart from the host and from every other jail. A customer system on its own release and its own patch schedule belongs there.
And a team that would rather keep its Containerfiles may keep them. FreeBSD has published official OCI images since 14.3, and Podman runs them. The Handbook adds the detail that settles the matter: "The underlying virtualization technology is still FreeBSD jails, with the same feature set."
What every image brings along
The shipping promise has a passenger nobody mentions on the box. Every image carries a userland of its own, a Debian here and an Alpine there, frozen at the moment somebody built it and maintained, if at all, by whoever that somebody was. Researchers at North Carolina State University pulled 356,218 images from Docker Hub in 2016 and counted the known vulnerabilities inside them: 199 per community image on average, 185 per official one, at least one rated high in more than 80 per cent of both, many images untouched for hundreds of days, and the flaws handed down from parent image to child. Seven years later Sysdig looked at the images its customers actually run in production and found high or critical vulnerabilities in 87 per cent of them.
The administrator of the host can patch none of it in place. A fix to an image is a rebuild, and the rebuild belongs to whoever wrote the Containerfile. A host with forty containers waits on up to forty such schedules before its userland is current, and the environment that was promised does arrive identical on every machine, with every unpatched library it was built with.
A jail's userland comes from one place. Every thin jail carries the base system of the template it was cloned from, which comes from one project with one security team and its advisories, and the administrator patches it from outside the jail with freebsd-update -j for the base and pkg -j for the packages. Thin jails built on NullFS go one step further, and the FreeBSD Handbook puts it plainly: "updating the template updates every jail that mounts it." Checking takes one line as well. FreeBSD's daily security run audits the installed packages against the project's vulnerability database by default, and security_status_pkgaudit_jails="*" in periodic.conf extends the audit to every jail on the host.
What each one is for
A matrix of mechanisms says what the parts are. The question a reader has on a Tuesday is what to use, so here are the jobs people actually do with isolation, and what covers each one on FreeBSD.
several services on one host, kept apart thin jail
customers or teams sharing a machine, each with root thin jail
a build that must not see the machine building it thin jail
a second version beside the first, for testing thin jail
shipping a service to another machine thin jail, zfs send
developing on a laptop against production thin jail inside a small VM
a system on its own release and patch schedule thick jail
a Linux binary nobody will port jail with the Linux layer
code from strangers you do not trust VM
a different kernel or operating system VM
Eight of the ten are jail jobs, six of them for a thin jail. The two that are not are exactly the two where Linux reaches for a virtual machine as well, and for the same reason.
The third line runs at the scale of the whole project: FreeBSD builds its entire package collection, tens of thousands of ports, with poudriere, and poudriere builds every package inside a jail. The first two lines are the hosting problem Kamp was asked to solve. The laptop line has a worked answer in my book, the Jail Ferry: a FreeBSD VM under QEMU, the production jail pulled into it with zfs send, the project directory shared over NFS, and the whole orchestration in 223 lines of POSIX shell. Docker Desktop, the commercial equivalent on the other side, hides a Linux VM of roughly four gigabytes behind a daemon and charges any company with 250 employees or ten million dollars in annual revenue for a subscription.
Where the escapes live
Kubernetes describes containers in its own documentation as offering "a weaker isolation boundary than virtual machines", and recommends sandboxing "when you are running untrusted code". Docker's documentation describes its default seccomp profile as one that "disables around 44 system calls out of 300+". Both statements are candid, and both concern a fence that a program assembles at runtime.
A whole class of escapes has lived in that program. runc published three container escapes on a single day, 5 November 2025, after CVE-2024-21626 the year before and CVE-2019-5736, which let a container overwrite the runc binary on the host. containerd has published ten advisories so far in 2026, four of them rated critical. Every one of those sits in the software that builds the isolation, which on the Docker side runs to a quarter of a million lines and on FreeBSD to 5,329.
The limit
Four deductions, and the first is the one a careful reader will be waiting for.
A jail shares the kernel, and the kernel has bugs. FreeBSD has published twelve security advisories against jails since 2004, four of them in 2026. The latest came on 29 September: three ways for a jailed process to step outside its filesystem root through descriptor handling, fixed in 15.1-RELEASE-p4 and the corresponding branches. At 39C3 in December 2025, ilja and Michael Smith presented roughly fifty distinct kernel issues reachable from inside a jail, many of which could crash the system or open a path beyond it. A host that runs jails needs its kernel patched like any other, and those advisories are the reason.
Strangers' code belongs in a VM, on both systems. Everything AWS and the Kubernetes documentation say about containers and untrusted tenants applies to jails as well, because both share the host's kernel.
Kubernetes does not run here. Its nodes run Linux or Windows, and the operators and managed clusters built around it assume as much. A team whose operations live there will not replace them with jail.conf, and a FreeBSD host makes no claim that it can.
bhyve has no live migration, and GPU passthrough is partial. Workloads that move running VMs between hosts, or need a graphics card in the guest, are better served by KVM today.
The bill
On most teams today, the default when two services need to be kept apart is a container engine, usually Docker, chosen because the workflow is familiar, and the price of that familiarity is rarely itemised.
Here is what that default costs on a single host. Two daemons running as root, with dockerd the process every container depends on; Podman removes the daemon and the need for root, and still assembles the namespaces anew at every start. A quarter of a million lines of userland to build a fence that the kernel on the other system builds in one call. Network latency doubled in IBM's test under Docker's default NAT, until somebody switches the bridge off. Images that leave the data behind by design, so the database travels by some other route. Images that each bring a distribution of their own, 87 per cent of them with a high or critical vulnerability in Sysdig's 2023 count, patched only when their authors get round to a rebuild. And an advisory list for the assembling software that runs, this year, to ten entries for containerd alone.
What the default buys over a jail is the ecosystem. A team that runs Kubernetes will keep it, and the limit above says so. A team that needs four services kept apart on one host is paying all of the above for the first line of that list of ten, and the answer ships with the base system, configured in nine lines.
Run ps on a FreeBSD host with a dozen jails and look for the process that holds them together: there is none, and there has not been one since March 2000.
A FreeBSD jail has been one kernel object since March 2000, with 60 of 244 privileges and nothing in userland holding the fence up. Docker builds the same fence from eight namespaces and four further mechanisms with a quarter of a million lines, and at runtime both cost nothing. A thin jail ships as a ZFS snapshot through a pipe, its data included, and eight of ten isolation jobs on FreeBSD are jail jobs.