Make servers
go brrr.
Slopio is a performance-oriented, thread-per-core runtime built on io_uring, with batteries included: kTLS, HTTP/2, gRPC, DNS, WebSocket and Postgres.
See how it worksBatteries included
- ๐kTLSTLS offloaded to the kernel
- ๐HTTP/2multiplexed streams
- ๐กgRPCprotobuf RPC over HTTP/2
- ๐งญDNSasync name resolution
- ๐WebSocketfull-duplex connections
- ๐Postgresdatabase driver
Features
- โ Full cancellationJust drop an operation: dropping it triggers its cancellation, with nothing to clean up by hand.
- โฑ๏ธ Really cheap timersTwo timer wheels keep timers cheap enough to put a timeout on every request without thinking twice.
- ๐ TelemetryOpenTelemetry support: export traces to the tools you already use.
- ๐ฒ Deterministic kernel simulationSwap the real kernel for a simulation with virtual time and fault injection. Same seed, same run, and libraries built on Slopio get it for free.
- ๐ฐ Backpressure via buffersThe buffer pool controls the inflow: no free buffers, no more reads, so data never piles up in memory.
- ๐งฉ NUMA-awareShape the runtime to your machine's NUMA topology, with very few dynamic allocations.
Why Slopio?
Tokio is the default async runtime in Rust for good reasons: it's mature, well documented, and almost every networking library is built on it. It's very practical. But it's designed for the general case, and on a server that has to squeeze every core, that generality has a cost.
Work-stealing leaves hardware on the table
Tokio balances load by letting idle threads steal tasks from busy ones. Since a task can resume on any thread, every spawned future must be Send + 'static, shared state ends up behind Arc and Mutex, and data keeps moving between cores, taking cache lines with it. Slopio pins one thread per core and keeps each task where it was spawned: no synchronization to pay for, and caches stay hot.
I/O traits shaped by epoll
Tokio's AsyncRead and AsyncWrite follow the readiness model of epoll and kqueue: wait until a socket is ready, then read into a buffer the caller lends for the duration of the call.
fn poll_read(
self: Pin<&mut Self>,
cx: &mut Context<'_>,
buf: &mut ReadBuf<'_>, // borrowed, only for this call
) -> Poll<io::Result<()>>
io_uring works the other way around: you submit an operation and the kernel completes it later, so the buffer has to stay alive until the kernel is done with it, even if the future was dropped in the meantime. A borrowed buffer can't guarantee that, and features like provided buffer rings and multishot reads, where the kernel picks the buffer, don't fit the trait at all. Using io_uring behind these traits means giving up most of what makes it fast.
Other runtimes, but no ecosystem
Rust already has thread-per-core and io_uring runtimes. The problem is everything around them: HTTP, gRPC, TLS, database drivers are written against Tokio's traits. On another runtime you end up with compatibility layers that bring back the costs you were trying to avoid, or you write the missing pieces yourself.
So I started Slopio
I've always been drawn to async and systems programming. With AI multiplying how much one developer can build, a runtime plus its whole ecosystem stopped being out of reach for a single person, so it was my time to recreate what I wanted.
Slopio is the runtime and its ecosystem, designed together around the completion model: kTLS, HTTP/2, gRPC, DNS, WebSocket and Postgres, all built for io_uring and thread-per-core from the start, so nothing has to be translated back to epoll-era abstractions.
One thread per core
Each core runs its own event loop on its own thread. Worker threads are all identical; special threads, also pinned to their own cores, take care of shared work. Threads never share tasks or state: they talk by sending messages through io_uring.
io_uring: batch the syscalls
With a readiness model like epoll, every read and write is its own syscall. With io_uring, operations are written into a submission ring shared with the kernel, submitted together, and their results come back in a completion ring: many operations, one syscall.
Provided buffer rings
A classic recv needs a buffer up front, so every pending read pins memory even if the socket stays silent. Slopio uses io_uring provided buffer rings instead: each worker registers a ring of buffers with the kernel, and each connection gets a single multishot recv, submitted once, with no buffer attached. The kernel picks a buffer only when data actually arrives, and the same request keeps producing completions as more data comes in. Once processed, the buffer goes back into the ring with no syscall and, since the ring belongs to a single worker, no lock.
Built for the machine
Slopio makes very few dynamic allocations: buffers are allocated up front and recycled through the rings instead of going back to the allocator. And because every thread is pinned, you can shape the runtime to your hardware's NUMA topology, keeping threads close to the memory they work on.