Vivian Voss

kTLS Sends the File Without Copying It

technical beauty scope freebsd networking ktls

I spent part of this morning watching nginx send the same two gigabytes twice, with DTrace counting the system calls as they went past. The first time the worker called write 131,080 times, once for every sixteen-kilobyte record, each one sealed inside the process and handed down as a finished envelope. The second time it called sendfile 2,746 times, and write four times for the handshake, and the file never entered the process at all. The kernel took it from the page cache, cut it into records, sealed them and put them on the wire.

The two configurations differ by one line.

That line exists on both systems. What makes it worth a Friday is that the identical switch, measured by people with no stake in my opinion, pays out at quite different rates depending on which kernel it is thrown in, and the reason has a date on it.

What the kernel already had

In 1998 David Greenman gave FreeBSD 3.0 a system call named sendfile. The thought behind it was unglamorous. A web server that reads a file into a buffer and then writes that buffer to a socket is copying data the kernel is already holding, twice, for no reason the kernel can see. Point the kernel at the file and at the socket, and it will lend the socket the pages it has in its cache: one copy on the way in from the disk, none on the way out.

Gleb Smirnoff rebuilt the call in 2016, with Netflix paying for the work, so that it no longer waited for the disk either. Pages that are not yet in memory get reserved in the socket buffer and marked with a flag that says not ready; the disk fills them in its own time; the flag clears when the data arrives. It is a small piece of machinery, it lives in the socket layer, and you will want to remember it, because it matters again in five years' time.

What the padlock took back

Then the web went encrypted, and the copies came back.

TLS turns a file into a sequence of records. Each record is a sealed body with a short header in front of it and an authentication tag behind, and the sealing has to happen somewhere. In every SSL library ever written it happens in the process, which means the server reads the file in, encrypts it, writes the result out, and makes one system call per record whilst it is at it. Three copies, and a careful piece of 1998 engineering sitting in the kernel with nothing to do, because the single thing the kernel could not do was seal an envelope.

Both kernels went and got it back, and both drew the line in the same place: the handshake stays up in the library with the private key, and the record layer moves down, so what went into ring zero was the envelope and no more than that. On FreeBSD the work came from John Baldwin and Drew Gallatin at Netflix, where it had been carrying video since 2016, and it reached the tree with 13.0 in April 2021.

The bill for leaving things as they were had been counted by then. Facebook had put a number on the cost by then, presented by Dave Watson at netdev in 2016: a tenth of its total processor time going on encryption and decryption, with another two per cent on the copying in and out of the kernel that made them possible. That is before anybody has served a frame of video. Which is roughly when somebody measures the result.

One socket option

The whole FreeBSD interface is a setsockopt. Once the handshake is done, the library hands the kernel a small structure holding the cipher suite, its keys, the protocol version and the record sequence number to start from, under the option TCP_TXTLS_ENABLE, and from that call onwards everything written to that socket is framed and encrypted on its way through. write still works. sendfile works again, which was the point of the exercise. A library that needs to send an alert afterwards uses sendmsg with a control message to set the record type, and that is the complete list of things an application has to learn.

sendfile on;
ssl_protocols TLSv1.3;
ssl_conf_command Options KTLS;

What the implementation did not do matters more than what it did. Nobody had to invent a mechanism for data that sits in a socket buffer whilst not yet being ready to send, because Smirnoff had built one in 2016 and it was still there, unchanged, doing its job for asynchronous sendfile. Kernel TLS takes an unencrypted record, marks it with that same flag, drops it on a queue, and a pool of threads, one per processor, seals records and clears flags as it works through them. The transport layer below never learns that TLS exists. A network driver that could already send pages it had not looked at could send TLS records without a line of change, and most of them could.

Two gigabytes over TLS 1.3: system calls made by the nginx worker without kernel TLS 131,080 write with kernel TLS 2,746 sendfile, plus four write calls for the handshake 2.1% of the original call count survives the one configuration line. Measured with DTrace on FreeBSD 15.0, AES-256-GCM, 16K records.

Three backends sit behind the one option. The software backend seals records in those kernel threads using the system's crypto framework. Two drivers, Chelsio's cxgbe and Mellanox's mlx5en, hand the unencrypted record to the card and let it seal the payload after the DMA, so the host never touches it. The kernel asks the card first and falls back to software if the card declines, and the application cannot tell which happened, which is the correct amount for an application to know.

All of it is 2,503 lines in sys/kern/uipc_ktls.c and another 935 in the crypto backend, counted without comments on the 15.0 tree. The line options KERN_TLS sits in GENERIC, which is to say in the kernel you installed, switched on by default since 15.0, and the structure the library hands over has not changed by a single byte since 13.0.

The same line, measured by other people

In November 2021 the nginx team published their own test of the feature. Same software, same cipher, same configuration line, run on two operating systems: Ubuntu 21.10 gained 13 per cent in median throughput, and FreeBSD 13.0 gained 29. Their interest was in nginx. They reported both numbers without comment, which is the most useful thing a vendor can do with an inconvenient one.

Intel measured the Linux side on its own at netdev the year before, with nginx on a 100-gigabit E810 and files of varying size, and found kernel TLS with sendfile ahead of userland by about 30 per cent once files reached ten megabytes, and behind it at one and four kilobytes, with the crossing point at sixty-four. Gallatin's figure for the same question on FreeBSD is blunter: switching the crypto from userland into the kernel, he told EuroBSDCon in 2022, “almost quadruples CPU efficiency”.

One switch, three laboratories, and a spread that nothing in the TLS specification accounts for. The difference is in what each switch found waiting for it underneath.

The chain underneath

Consider the path a film takes out of a Netflix server, and who built each piece of it.

One tree, seven pieces, twenty-three years 1998 sendfile Greenman pages lent, not copied 2016 not-ready flag Smirnoff async sendfile, Netflix 2019 unmapped mbufs, NUMA siloing, lagg Gallatin chain 6:1, Xeon 105 to 191 Gb/s 2021 kernel TLS, RACK and BBR Baldwin, Gallatin, Stewart reuses the 2016 flag Each piece could read the source of the one before it, because it sat in the same tree.

The pages come out of the cache without a copy, because of Greenman in 1998. They are allowed to sit in the socket buffer before they are ready, because of Smirnoff in 2016. They are described by mbufs that carry several unmapped pages at once, which Gallatin added for this traffic in 2019: the chain for a sixteen-kilobyte TLS record shrinks by roughly six to one, plain sendfile traffic by as much as nineteen to one, and Netflix measured five to twenty per cent less processor time on unencrypted workloads from that change by itself. They are sealed by kernel TLS, which Baldwin and Gallatin wrote on top of all three and reused the flag that was already there.

They are paced on the way out by the RACK and BBR stacks, which Randall Stewart wrote for FreeBSD with Netflix paying, shipping in 13.0 as loadable modules you select per connection. The algorithms themselves came out of Google's work on Linux, and the credit for them belongs there: Neal Cardwell's BBR landed in Linux 4.9 in December 2016, RACK was in 4.4 before that, and the FreeBSD manual page cites the Google papers where another project might have quietly claimed the idea. What FreeBSD contributed was the frame that lets a second and a third TCP stack live beside the first one in the same kernel and be chosen per socket, which arrived in 2017.

The connection's pages and the thread that seals them are kept on the same NUMA node as the port they will leave by, which is what Gallatin's siloing work arranged from 2019 onwards, once the ceiling on those boxes had stopped being the processor and become the path to memory. The numbers for that step alone are the ones I would put in front of anyone who thinks this is a matter of taste: a Xeon 4216 went from 105 to 191 gigabits a second, with traffic on the inter-socket fabric falling from 40 per cent of memory traffic to 13, and an EPYC went from 68 to 194. And when the egress port has to be chosen, lagg picks one on the right node, which Gallatin committed in May 2019 as use_numa.

That is seven pieces over twenty-three years, and whoever built each one could read the source of the piece beneath it, because it sat in the tree they were already working in. Each of those capabilities exists in the Linux world too, written by different groups shipping on different schedules into a kernel that none of them owned, so every one of them arrived as its own negotiation. Nobody there was ever in a position to decide the order.

The small change is a fair illustration. On FreeBSD the feature is in GENERIC, so it is in the kernel that came with the system. On Debian and Ubuntu, kernel TLS is CONFIG_TLS=m, a module that is not loaded until something asks for it, and the kernel only autoloads it for a caller holding CAP_NET_ADMIN, which a sensibly configured web server does not. That is what happens when the kernel, the library, the web server and the people who assemble them are four separate projects whose only meeting point is a distribution.

What it adds up to in a rack

Gallatin's EuroBSDCon figures are the arithmetic of that sequence. 200 gigabits a second from one server by 2020. In 2021, on a 32-core EPYC with four 100-gigabit ports, software kernel TLS reached 240 and stopped, limited by memory bandwidth; the cards then sealed records inline, memory traffic halved, and the box reached 380. By 2022 the prototype was at 731 gigabits on a dual-socket machine, limited by the network cards dropping output, with processor headroom to spare. The FreeBSD Foundation's case study puts kernel TLS on its own at a 50-gigabit improvement, with processor load falling as the throughput climbed. The slide that closed the 2021 talk said that all of this is upstream, and it is. A stock 15.0 kernel on a rented machine carries the same code path, which is the part that should interest anybody who does not work at Netflix.

There is no comparable published figure from the other side, and I looked. The highest I found for Linux is a laboratory result from Intel and Varnish in 2021: 366 gigabits a second from one processor and 509 from two, with Ubuntu 20.04, OpenSSL 1.1.1h, which has no kernel TLS at all, and 4,096 synthetic connections where Netflix carries roughly 400,000. In production, a video platform published 73 gigabits a second per server this August using Go and Linux kernel TLS. Netflix's own appliance page lists 200 gigabits for the storage box, running FreeBSD. One of those numbers came off a bench driven by wrk, the others off racks that carry paying traffic. The organisation with the most video on the public internet has published its working for six years running, and the operating system in those slides is not the popular one.

What arrives from the network

The objection is obvious and deserves answering properly: remote input now lands in the kernel.

It does, on the receive side. The kernel has been parsing bytes that arrived from strangers since 1983, and TCP is a parser of hostile remote input with forty years of practice at it. What kernel TLS adds on the way in is a five-byte header carrying a length, and an authentication tag, and no byte of payload is believed until the tag has been checked. The part of TLS that is genuinely nasty to parse, the handshake with its certificates and its key exchange, is precisely the part that stays out in the library, where a crash costs one process. It is the division IPsec has used since the nineties, keys from a daemon in userland and packets sealed in the kernel.

It is still surface, and the three advisories of 2026 are its honest accounting: every one of them on the receive side, in bookkeeping around records a peer had sealed with its own keys, which on a public server means anybody who cares to connect. They were found by outside researchers and by FreeBSD's own developers, and fixed within days. One thing I would change, and it is a small one: on my test box OpenSSL switched both directions on with the same line, so the kernel decrypted the client's request as well as sealing the reply. A server that serves files wants only the sending half and ought to be able to say so. The sysctl is a single switch for both.

On a smaller machine

My own figures come from a Ryzen 5 3600 with a one-gigabit Realtek port, nothing tuned, serving a gigabyte of random bytes to a client on the same box. They show the shape of the thing, and the numbers above give it its size.

With ssl_conf_command Options KTLS in nginx, the worker's share of ten gigabytes fell from 7.07 seconds of processor time to 0.26. The sealing moved to kernel threads on cores the worker was not using, and the whole machine, client included, spent eighteen per cent less time busy. A sysctl counted 329,429 records for five gigabytes, within half a per cent of five gigabytes divided by sixteen kilobytes, which is a pleasing way to confirm that the thing you measured is the thing you meant to measure.

The limit

Five things, and the first is the one a careful reader will have been waiting for.

The security record this year is not clean. The three advisories of 2026 were, in order, a local file overwrite in June where a decrypt in place reached pages belonging to a file on disk; a crash on a malformed CBC record at the end of that month; and this week a TLS 1.3 record that could panic the kernel with an off-by-one. All three sit on the receive side, the half of the feature a kernel should least want to own, and the transmit path this piece is about was not among them. That distinction is worth keeping, and it is also the reason to think twice before switching the receive side on where the sending half is all you need.

The comparison everybody reaches for is the Linux record, and it repays a close look. Three of its flaws arrived on a single day in February 2024, all of them in tls_sw.c, all scored 9.8 out of 10, and four more since have been specific to TLS 1.3 on the way in, which is the newest and least settled corner of the thing. The surfaces differ by construction as well: the record layer there accepts eight cipher suites, SM4 and ARIA among them, where this one accepts three and lets you switch the oldest of those off at runtime.

The same feature, two trees FreeBSD 15.0 Linux in the stock kernel options KERN_TLS in GENERIC CONFIG_TLS=m, CAP_NET_ADMIN cipher suites in the kernel 3, oldest switchable off 8, incl. SM4 and ARIA same nginx line, 2021 +29% median throughput +13% on Ubuntu 21.10 published single-server TLS video ~800 Gb/s at Netflix no comparable figure found mitigation without a reboot 2 runtime sysctls module will not unload Sources: nginx 2021, Gallatin at EuroBSDCon 2021 and 2022, Debian and Ubuntu kernel configs, ktls_create_session() in sys/kern/uipc_ktls.c, include/uapi/linux/tls.h.

The difference that matters at three in the morning is neither of those. Each of the three FreeBSD advisories can be defused with a sysctl before anybody schedules a reboot: kern.ipc.tls.rx_enable takes away the receive path on its own, kern.ipc.tls.enable takes away the lot, and both are writable at runtime. Linux offers no equivalent switch, the surface is the tls module itself, and a module with open sockets on it does not unload. The fix there is a new kernel and the reboot that comes with it, unless you pay for the livepatch service.

It buys processor time by spending memory bandwidth, and Intel's netdev paper is admirably direct about it: kernel TLS with sendfile used significantly more memory bandwidth than the userland path in their tests, which they called counter-intuitive and published anyway. On the Netflix boxes that trade is the whole point, since memory was the wall they hit at 240 gigabits. On a machine with one busy stick of RAM and a modest card, it may be no trade at all.

My own machine is a rented box with a one-gigabit port, and loopback is loopback. The client decrypted on the same processor, the card offloads nothing, and one worker served one stream. For what a 100-gigabit port does, quote the published figures above, which were produced by people who could measure it.

It helps workloads that look like a file, and Intel's numbers put a size on that: at one and four kilobytes the userland path was ahead, and the crossing point sat at sixty-four. A server that spends its day in handshakes gains nothing at all from a design that deliberately leaves handshakes where they were. Rekeying is out too, in both kernels, since the session keys can be set once per direction and that is that.

And the chain has a cost that the line count hides. Kernel TLS needs an SSL library built with the feature, a web server that knows the option, a driver that can send unmapped pages and a kernel compiled with it. Here they arrive together, in one upgrade, from one tree. Elsewhere they arrive one at a time, from four projects, and the version matrix is left as an exercise for the reader.

The sheet

Efficiency. 2,503 source lines in the core and 935 in the crypto backend, without comments, on 15.0; 1,521 and 576 at 13.0. Compiled into GENERIC, no module to load. A session structure of 256 bytes per direction and a 16-kilobyte buffer zone. On this machine, ten gigabytes over TLS 1.3 cost the nginx worker 0.26 seconds against 7.07 without it; 2,746 system calls for two gigabytes against 131,080. Independent figures: 29 per cent median throughput on FreeBSD 13.0 against 13 on Ubuntu 21.10 from the same nginx line, and “almost quadruples CPU efficiency” on the Netflix boxes. Runtime dependency: an SSL library for the handshake, and nothing else.

Security. Three advisories in 2026, all on the receive path, none on transmit; three Linux flaws at 9.8 in one day in tls_sw.c and four more since on TLS 1.3 inbound. Of the four questions: it runs in the kernel, so with every privilege there is; the receive side parses records from the network whilst the transmit side does not; the parser is the TLS record layer, accepting three cipher suites here against eight there; and the handshake, where the complexity of TLS actually lives, stays in userland with the private key and never enters. kern.ipc.tls.rx_enable takes away the receive half at runtime and kern.ipc.tls.enable the whole feature, which buys the time to schedule a reboot.

Longevity. FreeBSD 13.0, 13 April 2021; in production at Netflix since 2016; in every release since, and in GENERIC where other systems keep it in a module. OpenSSL 3.0 learnt to use it in September 2021, with Baldwin and Gallatin credited in the changelog beside Boris Pismenny for the Linux path, and nginx 1.21.4 followed in November. Linux has had its own since 4.13 in 2017. Nothing on macOS or OpenBSD.

Stability. The tls_enable structure the library hands to the kernel is byte for byte the same in 15.0 as in 13.0. The four socket options have kept their names and their numbers. The manual page runs to 259 lines today against 267 at 13.0, with 38 lines different between them, most of them a date and some new counters. Minimal configuration for nginx: one line.

Measured on 2 October 2026 on a spare AMD Ryzen 5 3600 running FreeBSD 15.0-RELEASE, nginx 1.30.4 and curl 8.22.0 from packages, OpenSSL 3.5.4 from the base system, one gigabyte of random data served from the page cache over 127.0.0.1 with TLS 1.3 and AES-256-GCM; worker on CPU 2 via cpuset, client on CPU 4; ten transfers per configuration, processor time from ps, system time from kern.cp_time, system calls counted with DTrace over two transfers, records from kern.ipc.tls.stats.ocf.tls13_gcm_encrypts.

Somewhere tonight a server is copying a film into a buffer so that it can copy it back out with a lock on it, and the kernel beneath it has had the pages all along.