LibreQoS Methodology

Test methodology

How the LibreQoS Internet Quality Test measures a connection, and how those measurements become the bufferbloat grade and per-application Quality of Outcome scores. Latency under load is the headline; raw speed is context.

What the test measures

In one paragraph

A speed test tells you how fast an idle link can move data. This test also asks the harder question: what happens to latency while the link is busy? The test deliberately loads your connection in each direction and measures responsiveness at the same time, then turns the result into an A–F bufferbloat grade and a set of application quality scores.

Throughput

The typical rate the link sustains in each direction while it is under load, measured as bytes transferred over the phase time.

Latency under load

How far round-trip time rises when the link is saturated. This is bufferbloat, and it is what makes a fast link feel slow.

Reliability

Packet loss and whether real-time traffic — calls, gaming, streaming — stays usable under that load.

The application scores follow the IETF Quality of Outcome (QoO) framework, which scores a connection by what it can actually do rather than a single number.

How a test runs

Phase sequence

A test runs in seven ordered phases in your browser. Two are calibration, two are scored (download and upload), and the rest are reference or diagnostic. A full run takes roughly 57–84 s, depending on how fast your link converges.

  1. Idle baseline 8–12 s

    Measures latency and loss with no test traffic. This is the reference point the load phases are compared against. It is not scored.

  2. Download warm-up 8–10 s

    Calibrates the download load for your link: it converges on block size, stream count, and a capacity tier. The first 4 s are ignored while it stabilizes. Not scored.

  3. Download 10–16 s

    Saturates the download direction while latency and loss are measured continuously. Throughput and loaded latency from this phase are scored.

  4. Upload warm-up 8–10 s

    Calibrates the upload load the same way. The first 4 s are ignored. Not scored.

  5. Upload 10–16 s

    Saturates the upload direction while latency and loss are measured. Throughput and loaded latency from this phase are scored.

  6. Bidirectional 10–12 s

    Downloads and uploads at the same time. This is a deliberately harsh stress test and is shown as a diagnostic. It never counts toward a grade or QoO score.

  7. Recovery 3–8 s

    Load stops and latency is watched as it settles. Diagnostic only.

Technical detail: warm-up is calibration, not a scored phase

The load has to be sized to the link. If the test used one fixed block size, it would either barely dent a fast connection or overwhelm a slow one. During warm-up an adaptive controller converges on the block size, stream count, and capacity profile that keeps the link busy for the whole phase.

Once warm-up finishes, the scored download and upload phases run non-adaptively on exactly the parameters it learned. They do not retune mid-phase and no static floor or bump is applied on top. That is the entire purpose of a warm-up: it makes the scored load stable and comparable between runs and between connections.

How throughput is measured

Bytes over time

During the scored download and upload phases the test opens several parallel connections to our servers and pushes as much traffic as the link will carry. The reported throughput is the total bytes transferred divided by the phase time, so it reflects what the link sustained rather than a momentary peak.

  • Download and upload are measured and reported separately.
  • Downloads and uploads use multiple parallel streams; uploads use larger blocks to keep request counts low.
  • Results also include percentile views, so an unusually steady or spiky phase is visible rather than hidden by one average.
Technical detail: block sizing and streams

The adaptive controller targets a short block duration (on the order of hundreds of milliseconds) and a bounded amount of data in flight. That is long enough to measure a real transfer and short enough to react to changing conditions without flooding the path. Parallel streams are spread across separate hostnames so they open independent congestion windows instead of multiplexing one.

On very fast links the test may add regional ISP or reference measurement endpoints as extra load once it is already using several streams (see “Where the traffic goes” below).

How latency under load is measured

Probes during load

Latency is probed continuously for the whole test, not once at the start. Each sample fires several small probes in parallel, discards failed or hung ones, removes statistical outliers, and reduces the rest to a single round-trip value. The phase is then summarized at the 95th percentile (p95).

  1. Probe. Up to four small round trips run in parallel, staggered by about 50 ms so a single unlucky path is not mistaken for the whole connection.
  2. Filter. Failed probes are dropped. Surviving samples get interquartile-range outlier removal.
  3. Reduce. The median of the filtered samples becomes one latency sample for that round.
  4. Summarize. Each phase reports p95 loaded latency, the tail that real-time apps actually feel.
Technical detail: why p95 and not an average

Averages hide the bad moments. A video call does not care that most packets arrived quickly if a steady minority were late enough to cause stutter, so the test reports the 95th percentile of loaded latency. Idle latency is kept as a reference; scoring uses loaded latency and the increase over idle.

Loss is derived from failed probe rounds, and the latency component is scored on a graded scale between each application’s “good enough” and “unusable” points (see the QoO section).

Bufferbloat and the A–F grade

Latency increase

Bufferbloat is the rise in latency caused by a busy link. We measure it as the loaded latency at the 90th percentile minus the idle latency at the 5th percentile, with differences under 2 ms treated as noise. Download and upload are graded separately, and the overall grade is the worse of the two — a good download cannot hide a bad upload. The bidirectional phase is not part of the grade.

A+under 5 ms
Aunder 30 ms
Bunder 60 ms
Cunder 200 ms
Dunder 400 ms
F400 ms and up

The grade describes responsiveness under load. It is separate from raw speed: a 1 Gbps link with an F grade will still stutter on a call while something else is transferring.

Technical detail: p90 loaded vs p5 idle

Comparing a loaded tail (p90) against a best-case idle reference (p5) makes the grade sensitive to the spikes users actually notice, instead of smoothing them away. The test also tracks mean-based increases for the detailed results table, but the letter grade uses the tail-based delta above.

Application quality: Quality of Outcome

Weakest link

Quality of Outcome (QoO) asks how a connection would feel for a specific application. Each app is scored on three dimensions: latency, loss, and throughput. Latency and loss are graded between a “good enough” point (100%) and an app-failure point (0%), while throughput is a pass/fail gate at the app’s minimum rate. The app score is the minimum of the three, so a fast connection with high latency under load does not score well for a video call.

QoO = min( latency, loss, throughput )

Latency

p95 loaded round-trip time, scored on a straight line between the app’s good and unacceptable thresholds. This is where bufferbloat shows up.

Loss

Mean packet loss during the loaded phase, scored the same way. Real-time apps feel loss long before throughput suffers.

Throughput

A pass/fail gate against the minimum rate the app needs in each direction. Meet it and throughput does not cap the score; miss it and the app scores zero.

Each app uses the worst result across its scored phases (for two-way apps that means the worse of download and upload), and the percentile view of throughput is compared against the app’s gate.

LibreQoS default application profiles
Application Scored phases Latency (good → bad) Loss (good → bad) Minimum throughput (down / up)
Web browsing Download + Upload 60 → 600 ms 0.5% → 5% 5 / 1 Mbps
Video streaming Download 90 → 1200 ms 1% → 4% 7 / 1 Mbps
Video conferencing Download + Upload 70 → 200 ms 0.8% → 3% 3.5 / 3.5 Mbps
Audio calls Download + Upload 90 → 250 ms 1% → 4% 0.3 / 0.3 Mbps
Online backup Upload 180 → 2000 ms 2% → 10% 1 / 20 Mbps
Real-time gaming Download + Upload 55 → 140 ms 0.5% → 2% 1 / 1 Mbps

The application classes and values above are LibreQoS implementation defaults, not thresholds prescribed by the IETF. They can be tuned per network or use case. The streaming minimum is the 1080p rate.

Excellent90–100no noticeable issues
Good80–89meets ISP target range
OK60–79noticeable under load
Poor0–59likely user impact

Where the traffic goes

Measurement hosts

Most of the load is served from our own Cloudflare-hosted download shards, with Cloudflare’s public speed-test host as a failover. Some tests can additionally use regional ISP or reference measurement endpoints, which puts load closer to the user and gives network operators their own vantage point. The mix of hosts that served a test is recorded so a fast result is never presented as if a single server produced it.

  • Warm-up always runs on our shards, so an endpoint’s capacity cannot distort the calibration that sets the scored load.
  • ISP/reference endpoints are only added once the test is already using at least five streams — they supplement the main load rather than replace it.
  • Endpoint candidates whose round-trip time exceeds 150 ms are dropped before the download phase begins.
  • Uploads only use own-network or same-country endpoints; a long-distance upload path would under-report the connection.
  • Optional peering probes run during the baseline phase as context only and never affect a score.

Limitations and what is not scored

Reading the result
  • Browser-based. Results reflect your device, browser, Wi-Fi, and the path to our servers. A weak device or browser can under-report very high speeds.
  • Keep the tab focused. Background tabs are throttled by the browser, which creates gaps; the test warns you before it starts and stops cleanly if the tab becomes inactive.
  • VPNs and proxies route traffic differently and can change both throughput and latency.
  • One destination. The test measures the path between you and our infrastructure, not to every service you use.
  • A snapshot. Conditions vary with time of day, congestion, and in-home activity. Treat a single run as a sample, not a verdict.
  • Diagnostics are not scored. Idle baseline, bidirectional, recovery, and peering probes inform context but never affect the grade or QoO scores.
  • Virtual Household is a separate test mode, reached from the Household tab, for household-level contention.

Standards and references

Where this comes from

Quality of Outcome is defined by the IETF IP Performance Measurement working group in draft-ietf-ippm-qoo. LibreQoS implements the framework’s weakest-link aggregation (the min() rule), its graded latency and loss scoring, and its pass/fail throughput gate. The per-application minimum rates and the bufferbloat A–F bands are LibreQoS choices, as the draft leaves application requirements to the implementer.

The numbers on this page — phase durations, grade bands, and application profiles — are rendered from the same behavior specification (v4) that the web test and the CLI both use, so this page cannot drift from the test.