Vivian Voss

Every Read on QNX Is a Message

unix universe freebsd qnx microkernel

fusefs is in the base system on the machine I am writing this on. FreeBSD documents it in section four of the manual and ships it with the kernel, and I mount things through it without thinking twice: an archive, a remote directory, whatever needs to look like a filesystem for the afternoon.

Every read of a file inside such a mount leaves the kernel, travels to a user-space daemon, waits there, and comes back. I have chosen, on a monolithic system, to route a filesystem through message passing between processes, because writing a kernel module for the afternoon would be absurd.

That arrangement is how QNX does everything, to the byte.

Call read() on FreeBSD in the ordinary case and the processor changes privilege level, runs kernel code that already holds the filesystem and the driver, and comes back. One boundary, crossed twice. Call read() on QNX and the C library turns it into a message, the kernel hands that message to the filesystem manager, which is an ordinary process with no special rights, and the calling thread blocks until the reply arrives. The kernel itself knows nothing about files. It schedules threads, handles interrupts, manages some memory and passes messages, and that is close to the whole of it.

Two context switches where the other system has none. The price is paid on every call, and the argument about whether it is worth paying is older than most of the people having it.

What the kernel does not contain

The QNX Neutrino kernel in version 6.3.2 came to roughly 23,000 lines. Linux 2.6.0, released around the same period, was about 5.2 million. The ratio is not a rhetorical flourish; the two numbers count different things because the systems put different things in different places.

The same four jobs, on two sides of the line MONOLITHIC applications KERNEL scheduler, memory, interrupts filesystems disk drivers network stack, graphics one privileged address space a read crosses the line twice MICROKERNEL applications filesystem manager user space drivers, network graphics user space KERNEL scheduler, interrupts, memory, messages a read crosses it four times Neutrino 6.3.2: about 23,000 lines of kernel Linux 2.6.0 of the same period: about 5.2 million

The consequence people quote is the one that sounds like marketing until you have had the pager go off. A disk driver that corrupts memory it does not own takes a monolithic system with it, because it was running in the same address space as everything else. On QNX that driver is a process. You kill it and start it again, and the kernel never noticed. It can be attached to a debugger like any other program, because it is any other program.

The consequence people quote less often is the one the car industry actually buys: a kernel that small, doing that little, can be reasoned about well enough that somebody will certify the response time. Interrupts are disabled for hundreds of nanoseconds, where a large kernel is held to whatever its longest path happens to take. A bounded worst case comes out of a kernel small enough to have no long paths in it, and tuning a big one has never produced the same guarantee.

The bill that Mach sent

The history matters at this point, because for about a decade the microkernel argument was settled by one very bad implementation.

Mach came out of Carnegie Mellon in the eighties and became the microkernel everybody measured. On the hardware of the day, a minimal inter-process message of eight bytes cost about 115 microseconds round trip. That number ruined the idea for a generation. If the smallest unit of communication in your operating system costs that much, and your operating system is made of communication, the arithmetic ends the conversation.

Jochen Liedtke did not accept the arithmetic. His L3 kernel, measured against Mach on the same class of machine in 1993, did the same eight-byte message in 5.2 microseconds. His L4, three years later on a 486-DX50, did eight bytes in 5 and 512 bytes in 18, against Mach's 115 and 172. That is a factor of twenty-two, achieved by rewriting the path without redesigning the architecture, and it moved the question from whether microkernels can be fast to why one particular microkernel was not.

Four numbers the folklore never reconciled 115 µs Mach, eight-byte message round trip, 1993 5.2 µs L3, the same message same class of machine 5 % L4Linux against native whole-system, AIM, 1997 279 seL4 cycles per call 1 GHz Cortex-A9, today every figure from published work with public methodology; none of them measured here

What it costs now

The numbers have not stopped improving. seL4, the descendant of that line and the only kernel of its kind with a machine-checked proof of correctness, publishes its own measurements: a call() from one address space to another takes 279 cycles on a 1 GHz Cortex-A9, and 412 on a 3.1 GHz Haswell. Reply and receive cost about the same. On the Haswell that is roughly 130 nanoseconds for a full round trip between two processes.

For comparison, a plain system call on a modern monolithic kernel lands in the same neighbourhood once the side-channel mitigations are switched on. The gap that killed Mach has narrowed to something you have to look for.

Then there is the whole-system measurement, which is the one that ought to settle arguments and somehow never does. In 1997 Härtig and his group at Dresden put Linux on top of L4 and ran the AIM benchmarks against native Linux on the same hardware. Maximum throughput came out five per cent lower, with typical penalties between five and ten per cent. The same experiment on Mach-based MkLinux was far worse, which was rather the point they were making.

Five to ten per cent, for an operating system in which the kernel has been reduced to a message switch, is not the catastrophe the folklore remembers.

The part we already pay

Back to the mount I started with, because the uncomfortable part sits on this side of the fence.

A FUSE filesystem is a microkernel built by hand out of parts that were never designed for one. The kernel receives the read(), packages it, sends it to a user-space process, waits, and copies the answer back, on a kernel whose entire design rests on the opposite assumption.

One read, three routes FREEBSD, ORDINARY application kernel, fs and driver two crossings FREEBSD, THROUGH FUSE application kernel fuse daemon four crossings QNX application kernel routes fs manager four crossings, always

The measurements exist. Vangoor, Tarasov and Zadok took a pass-through FUSE filesystem to FAST in 2017 and reported that a straightforward implementation lost up to 60 per cent of throughput on a hard disk and up to 83 per cent on an SSD in the harshest case, small reads with many threads queueing behind a single-threaded daemon. A carefully optimised version still lost 27 per cent on 4 KB writes. Their summary is fair and worth repeating: the overhead is moderate for many workloads and near zero for some, and it eats CPU throughout.

So the question is not whether a general-purpose system will pay for user-space components. FreeBSD ships fusefs in the base system and I have used it happily for years. The question is what the payment buys. On QNX it buys a driver you can restart and a worst case somebody will certify. On my laptop it buys the convenience of mounting an archive without writing kernel code, and when the daemon dies the mount point goes stale and I get to explain to somebody why the directory has stopped existing.

Where the split actually sits

The honest framing has little to do with microkernel against monolith. It is a question of where a system puts its boundaries and what it charges at each crossing.

FreeBSD puts almost everything on the inside and charges nothing for the crossings that do not happen. The price is that the boundary between the network stack and the disk driver is a convention, held in place by review and by a single tree in which both were written by people who talk to each other. That is a real form of integrity, and it rests on a culture; the hardware has nothing to do with it.

QNX puts the boundaries in hardware and charges a message for each one. What it gets in return is a system in which the failure of a component stops at that component, which is why it runs in 275 million vehicles, in medical devices, in surgical robots and in industrial controllers, where a crash goes into an incident report.

Notice what the two answers have in common, because it is the part that decides whether either one works. Both systems are built by one group of people who ship kernel and userland together: QNX as a product out of one house in Ontario, FreeBSD as one tree with one release. A boundary is only worth what the people on both sides of it agree it is worth, and agreement is cheap when there is one team and expensive when there are forty upstreams. A system assembled from independent projects has the hardest time of the three, because it can enforce a boundary in hardware and still not know who is answerable for the thing on the other side of it.

Both answers are coherent, then, and they are coherent for the same underlying reason. They are priced differently, and the price is the thing the folklore keeps getting wrong.

The limit

Three deductions, and the first one is about my own numbers.

None of these figures were measured by me. They come from published work with public methodology: Liedtke's 1993 paper, the Dresden group's SOSP paper from 1997, the seL4 project's own performance page, the FAST 2017 FUSE study, and QNX's documentation. That is deliberate, and it also means every figure carries the hardware and the workload of its own experiment. An IPC benchmark measures IPC; it does not tell you what your application will do.

The comparison across decades is generous to nobody. A 486 from 1996 and a Haswell from 2013 have so little in common that the microsecond figures are useful only against their own contemporaries. The one comparison that survives is the one inside each experiment, which is why the Mach-against-L3 number matters and the absolute value does not.

And QNX pays in a currency my desktop does not accept. Message passing is synchronous by design: the sender blocks until the receiver replies. That is exactly what you want when you need to reason about timing, and it is a straitjacket when you want throughput. Nobody builds a database on it, and QNX has never suggested you should.

The bill

The microkernel bill has three lines, and two of them are smaller than advertised.

Per call, you pay two context switches, which cost a few hundred cycles on hardware anybody would recognise. Across a whole system, the best measured answer we have is five to ten per cent. And in engineering effort you pay a great deal, because the parts have to be written as processes that can survive each other's failure, which is harder than writing them as functions that cannot.

What you buy is a system whose failure modes are bounded and whose worst case can be put on a certificate. Whether that is worth five to ten per cent depends entirely on what happens when your filesystem driver takes a wrong turn: an annoyed sigh and a reboot, or an incident report with your name on it.

The joke, if there is one, is that the industry spent twenty years explaining that microkernels are too slow for real work, and then quietly put one in every car.

Two context switches on every read, paid by 275 million vehicles. Mach made that look absurd at 115 microseconds a message; Liedtke did the same message in 5.2, and Dresden measured a whole system at five per cent below native Linux. Every FUSE mount on my own machine pays the same price, and what it buys is convenience where QNX buys a certificate.