Vivian Voss

What FreeBSD Charges at the Boundary

the bench marker freebsd benchmark kernel

The claim under test is older than most of the people repeating it. Separating the parts of a system into their own address spaces is said to cost so much at runtime that no serious system should do it. Tanenbaum and Torvalds had that argument in 1992, Torvalds restated it in 2006 when he called microkernels inherently slower, and it turns up today in every discussion where a monolith is defended by the cost of crossing a line.

It is worth testing now because the question has come back around from the other side. Rust went into the Linux kernel in 2022 and by this year it is shipping in production, including a Rust implementation of Android's Binder IPC driver and Apple's GPU driver in the Asahi work. That is the same problem with a different answer: keep the driver inside the kernel and let the type system exclude the class of fault, where the older answer moves the driver out and pays for every message.

One of those answers has a price you can measure with a stopwatch. So I measured it.

How this was measured

One C programme of about 120 lines, the same source on all three systems. It times four things: an empty function call through a genuine call path as the zero line, a system call that has no fast path to hide in, a round trip over a Unix domain socket between two processes, and a round trip over a pair of pipes between the same two processes. Warm-up runs are discarded, the round trips take 50,000 samples each, and the figures below are medians.

The system call is getppid, chosen because it does almost nothing: it reads one field and returns. Whatever it costs is the price of entering and leaving the kernel, with as little work as possible hiding inside. The call and syscall figures come from 201 blocks of 1,000 calls each, with the compiler prevented from optimising the empty call away.

Three machines. FreeBSD 15.1-RELEASE-p3 in a jail on bare metal, on a 13th Gen Intel Core i5-13500. macOS, Darwin 27.0, on an Apple M1 Max. And Debian 13 with Linux 7.1.8, running as a bhyve guest on that very same Intel machine.

Two honest notes before the numbers, because both of them changed what the numbers mean.

The first is a mistake I made. On Darwin the first run reported round trips as clean multiples of a thousand nanoseconds, which is the sound of a clock that cannot see any finer. CLOCK_MONOTONIC on macOS rounds to microseconds, so the timing had to move to mach_absolute_time. Anybody benchmarking across these systems should check that before trusting a single figure.

The second matters more, and it took a correction to get right. The Linux figures come from a guest, and my first instinct was to write them off as hypervisor overhead. That instinct was wrong: a plain system call raises no VM exit at all, so bhyve sits almost entirely outside this measurement. What the guest column actually shows is something else, and it took reading the processor feature bits on both sides to find it. Nothing was installed in that guest to get the numbers: the binary was cross-compiled on my desk with zig cc -target x86_64-linux-musl -static, copied in, run, and deleted.

The numbers

Median cost of one operation, in nanoseconds FreeBSD 15.1 i5-13500, bare metal macOS 27 M1 Max, bare metal Linux 7.1.8 bhyve guest, i5-13500 function call under 1 1 1 system call 52 149 816 unix socket round trip 3,193 3,500 15,635 pipe round trip 3,038 3,291 16,674 mach port round trip not available 2,750 not available 50,000 samples per round trip, 201 blocks of 1,000 calls for the two upper rows, medians reported

Read the first two rows together, because they carry the argument. On FreeBSD, the step from a function call to a system call is a factor of fifty, and the step from there to a process boundary is another factor of sixty. The multipliers differ on the other two machines, and the ordering does not. Crossing into the kernel is cheap. Crossing into another process is not, and the gap between those two is where every architecture argument about splitting a system actually lives.

Fifty-two nanoseconds is the fastest of the three, on the same generation of hardware as the Linux guest and on silicon two years older than the Mac.

The same three steps, three times over 1 ns 10 ns 100 ns 1 µs 10 µs function call system call process boundary logarithmic scale, so equal vertical distances mean equal multipliers FreeBSD macOS Linux guest 52 149 816

The one that surprised me

The last row of the matrix is the one I did not expect.

Mach ports are the message-passing mechanism of the Mach kernel, which sits underneath macOS to this day. The measurement that made microkernels a punchline was taken on Mach: a minimal eight-byte message cost roughly 115 microseconds round trip on the hardware of 1993, which is 115,000 nanoseconds. That number ended the conversation for a generation.

On an M1 Max, the same mechanism does a round trip in 2,750 nanoseconds. It is forty times quicker than the figure that buried it, and on that machine it beats both POSIX answers sitting next to it: the Unix socket at 3,500 and the pipe at 3,291.

Some of that is thirty-three years of silicon, and a good deal of it is not. The mechanism was reimplemented, the message path was shortened, and the thing that was held up as proof of slowness is now the quickest way across a process boundary on that machine. Folklore has a longer half-life than the implementation it came from.

Thirty-three years of the same argument 1968 Conway 1992 Tanenbaum vs Torvalds 1993 Mach 115 µs, L3 5.2 µs 1997 Dresden: five per cent 2022 Rust in the kernel 2026 2,750 ns The figure that settled the argument in 1993 has been wrong for at least a decade, and the argument carried on regardless

What the guest was never told

Eight hundred and sixteen nanoseconds for a system call looks damning, and the easy explanation is wrong.

A system call does not leave the guest. There is no VM exit, no trip through the hypervisor, nothing for bhyve to charge for. The number is what the Linux kernel in that guest genuinely spends on its own kernel entry path, and the reason is written in the processor's feature bits.

The host reads them and finds this:

IA32_ARCH_CAPS = 0x14c8fd6b <RDCL_NO, IBRS_ALL, SKIP_L1DFL_VME, MDS_NO, TAA_NO>

RDCL_NO is the processor stating that it is not affected by Meltdown. MDS_NO says the same about microarchitectural data sampling. FreeBSD reads that, believes it, and switches the countermeasures off: vm.pmap.pti sits at 0 and hw.ibrs_disable at 1 on this machine.

The guest cannot read any of it. Its /proc/cpuinfo carries no arch_capabilities flag at all, because the register is not passed through, so the kernel inside has no way to learn that this silicon is immune. Faced with an unknown processor, it does the correct and expensive thing:

meltdown                : Mitigation: PTI
mds                     : Mitigation: Clear CPU buffers
indirect_target_selection: Mitigation: Aligned branch/return thunks
l1tf                    : Mitigation: PTE Inversion

Page table isolation swaps the page tables on every entry and exit. Buffer clearing runs on every return to user space. The aligned thunks add instructions to every indirect branch on the way. Each of those lands on the same entry path that the benchmark walks two hundred thousand times, and together they are the difference between 52 and 816.

So the 816 nanoseconds are the price of a missing statement. The same kernel, on the same silicon, told what the silicon actually is, would be somewhere else entirely, and I would expect it in the same neighbourhood as the FreeBSD figure. Nobody here has made a mistake. The processor knows, the host knows, and the guest was never told.

Who decides what a boundary costs

That belief has an owner, and the owner is different in each case.

On FreeBSD the kernel, the userland, the boot loader and the documentation come out of one tree with one release. The decision about page table isolation was taken by people who could see every consumer of it in the same repository, and when it changes, everything that depends on it changes in the same commit. Conway had this in 1968: the shape of the organisation ends up as the shape of the product.

On the assembled side the same decision is distributed. The kernel is one project, the virtualisation layer is another, and the defaults that arrive on your machine are the sum of groups making locally correct choices without a shared view of the result. The 816 nanoseconds are what happens when a fact known at one layer never reaches the layer that would act on it, and filing that under Linux failings would put the blame a long way from where the gap sits.

That is the whole argument for an integrated system in one sentence, and it is why the fastest number in this table sits under the column where one team ships the lot.

The limit

Three deductions, and the first one takes the shine off my own headline.

The three columns are not a fair race. Two processors, two architectures, one of them virtualised. What survives comparison is the shape within each column: the distance from a function call to a system call, and from there to a process boundary. That ordering holds on all three, tens of nanoseconds into the kernel against thousands across a process, and it is what the argument rests on.

The Linux column needs a rerun on bare metal before anybody quotes it as a figure for Linux. It is an honest measurement of that guest and a poor one of the kernel, and I do not have a spare machine to settle it today.

And a microbenchmark is not a workload. getppid measures the boundary and nothing of the work behind it. On QNX, where the filesystem sits behind that boundary, the published whole-system figure is a penalty of five to ten per cent against a monolithic Linux, measured at Dresden in 1997. That is the number to argue with, and it is nowhere near the folklore.

The point

The old claim was that message passing is too slow for real systems. On this hardware a message between processes costs around three microseconds, a system call costs fifty-two nanoseconds on the system that decided it owed the processor nothing, and eight hundred and sixteen in a guest that decided otherwise.

The kernel type barely enters into it. What decides the bill is who is allowed to make that decision, and whether the person making it can see the whole machine while doing so.

On the machine that could, a call across the line costs fifty-two nanoseconds, and nobody had to convene three projects to agree on it.

Fifty-two nanoseconds on FreeBSD and 816 in a Linux guest on that same Intel chip, with macOS on Apple silicon at 149 in between. What separates the first two is one feature bit that the host could read and the guest was never given, and a system that ships as one tree had nowhere to lose it.