Skip to content

multiload: fix hang when per-thread load bandwidth is below 1 MiB/s - #60

Merged
seranian merged 1 commit into
google:masterfrom
karanthv:fix-multiload-sub-mibps-hang
Oct 1, 2026
Merged

seranian merged 1 commit into
google:masterfrom
karanthv:fix-multiload-sub-mibps-hang

Conversation

@karanthv

@karanthv karanthv commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Problem

multiload hangs forever, with every worker thread spinning at 100% CPU, when any memory-load thread measures less than 1 MiB/s during a sample window.

In LOAD_MEMORY_SAMPLE_MIBPS, each load thread publishes its bandwidth as (uint64_t)mibps into t->x.count. That field is also the handshake: main polls until x.count != 0 to know the thread finished the sample. When mibps < 1.0, the cast truncates to 0, so:

the thread sets cur_sample = nxt_sample and never reports this sample again;
main spins forever in while (ready == 0) waiting for a non-zero count.

This can happen in practice when many threads share a memory node with very low bandwidth. For example, 80 threads on a node degraded to about 60 MiB/s works out to about 0.76 MiB/s per thread. It can also be triggered with a large injection delay:

multiload -H -t 1 -m 16m -l stream-triad-nontemporal-injection-delay -d 1500000 -n 2 -v
# before: hangs after "main: sample_no=0"

Fix

Load threads now publish bytes/s instead of MiB/s in x.count, and main converts the value back to MiB/s before aggregating.x.countis 64 bits wide, so there is no overflow risk at any real bandwidth. A value of 0 would require a single pass over the load buffer to take millions of seconds, so the handshake no longer depends on bandwidth being at least 1 MiB/s.

The fix also removes the per-thread truncation error (previously up to 1 MiB/s per thread, summed across all load threads in Total(MiB/s)).

The per-thread values (ML(n) at -v -v, and PerThread= at -v) are now printed with %.3f, so bandwidth below 1 MiB/s is visible instead of rounding to 0 or 1.

Chase threads are unchanged. Theirx.countremains an iteration counter.

Testing
Injection-delay repro above: before the fix it hangs; after the fix it exits 0 and reports Total(MiB/s) of about 0.58.

Load threads published (uint64_t)mibps in x.count, which is also the
sample-ready handshake polled by main. Below 1 MiB/s the cast truncated
to 0, so main waited forever. Publish bytes/s instead and convert back
to MiB/s in main. Also print per-thread values with %.3f.
@seranian
seranian merged commit 8b68e91 into google:master Oct 1, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants