Workloads
Six workloads, each one part of an HTTP server’s job. Every server answers the same requests with the same bytes.
- Plaintext
GET /plaintext - Answers
Hello, World!(13 bytes) over kept-alive connections: the cost of HTTP itself.256 connections, kept alive. - Echo 4 KiB
POST /echo - A 4 KiB body, answered with the same bytes: reading bodies and writing larger responses.256 connections, kept alive, 4,096-byte body.
- Connection churn
GET /plaintext - A new connection for every request: accepting and closing, as fast as possible.64 connections, each request on a new one.
- Templates
GET /menu - A 12-row HTML table rendered from a template per request, every value escaped: Go’s
html/template, askama on axum, rocstache on roux, the platform’sHtmlon basic-webserver, a comptime template in Zig. Checked against the reference page.256 connections, kept alive. - SSE (Datastar)
GET /sse - A Datastar action: the page’s signals as JSON in the query, answered with a stream of ten Datastar events (the new count as a signal, the count’s element, eight log lines appended), each event a chunk of its own: Go flushing each, axum’s
Sse, basic-webserver’sSse.unfold!, roux’sSseeffects, fourneau’s streams. Checked against the reference stream, chunk by chunk.256 connections, kept alive. - RealWorld (Conduit)
/api/articles - A slice of RealWorld’s Conduit API on SQLite, the database in the loop: the article list (tags, favorites and authors joined, newest first), one article, a comment and a favorite, at once. Every server gets a fresh copy of the same seeded database (100 users, 1,000 articles) and opens it in WAL mode with
synchronous=NORMAL; its answers are checked against the seed, field by field. Go withdatabase/sqland modernc’s SQLite, axum with sqlx, roux with its typed statements, basic-webserver with itsSqlite, Zig with the SQLite library on fourneau.Open loop only: the list 50%, an article 30%, a comment 15%, a favorite 5%; 250 to 32,000 requests a second in all, the same rates for every server, each climbing until it answers under 90% of the rate. The chart is the mean over every request.
Each of the first five: 5 seconds of warmup (discarded), 20 seconds measured, three rounds. Definitions: race.json.
How the load is made: a closed loop
The load generator, oha, holds a fixed number of connections (256; 64 for churn), and each connection sends its next request only when the answer to the last one has arrived. That is a closed loop: the server’s own speed sets the pace. A faster server gets more requests, and the numbers are what each server does flat out.
In a closed loop, throughput and latency are two views of one thing. With N requests always in flight, Little’s law gives latency ≈ N ÷ throughput. At 256 connections and 50,000 requests a second, the typical request waits about 5 ms, mostly in queues, because the server is saturated by design. The latency figures are the latency of a saturated server, not of a lightly loaded one.
What a closed loop hides
- Coordinated omission. When a server stalls (a garbage-collection pause, a lock, a slow disk), the generator stalls with it: every connection is waiting, so no new requests are sent. The requests that real users would have sent during the stall are never made, so they are never counted as slow. p99 and p99.9 therefore understate what users arriving at a steady rate would see, and they understate it most for servers with pauses.
- Latency at a given load. Users do not slow down when a server does. The question “how slow is it at 20,000 requests a second?” is not asked here: every server runs at its own maximum, and the maximum differs.
- The knee. Real servers degrade sharply near capacity: queues grow and the tail explodes. A closed loop sits at capacity the whole time, so it never shows the climb.
Under load: an open loop
So after its rounds, each server climbs a ladder on one workload: requests offered at a fixed rate whatever happens, each measured from when it was meant to be sent, so a stall counts against every request it delays (wrk2’s method; oha’s -q with --latency-correction). The steps are 50%, 75%, 90%, 100% and 120% of that server’s own closed-loop median, each held long enough to settle. That is a ladder, not a ramp: a rate that keeps changing never settles long enough for its percentiles to mean that rate. The chart plots each step at its absolute rate, so a slow server’s comfortable half is never mistaken for speed. The 120% step is past saturation on purpose: there the queue grows for as long as the step lasts, and the latency says so.
A caveat measured, not assumed: oha pacing requests costs about half again the loader CPU of its closed loop. At the fastest servers’ rates the loader can become the limit. Each step records the loader’s CPU, and the chart draws a step where it was 85% busy or more as a hollow point: that point is about the loader, not the server.
Also worth knowing
- The loader runs on a bigger droplet than the server, and its CPU is recorded with every round, so a result where the loader was the limit can be told apart.
- Echo 4 KiB on the larger classes reaches the droplet’s network link before most servers’ limits: there, the fast servers tie. Each round records the network traffic, so this shows in the raw data.
- Throughput counts successful (2xx) answers only. Errors and other statuses are no work done.