Ray's Knowledge Base

A consumer that stops reading a bounded channel on error leaves its producers blocked forever

PitfallVerified 27 Sep 2026Holds anywhere
Pitfall. The symptom, what causes it, and the fix that was run and seen to work.

Symptom#

A pipeline with producers that send into a bounded channel and one consumer hangs forever at 0% CPU when the consumer fails, instead of reporting the error. In dftracer-utils, dftracer_view -o /dev/null on an indexed trace export hung forever when the output could not be opened, and the library could also abort.

Cause#

On the error, the consumer returned and stopped reading. The channel stayed full, so every producer waited in send for space that never came, and the task that waits for the producers never finished.

Fix#

On failure, the consumer must:

  1. Set a shared failed flag that the producers check, so they stop producing early.
  2. Keep reading (draining) the channel until all producers have closed it, and drop what it reads.
  3. Report the error after the drain.

Check every bounded-channel pipeline for this: find each place where a consumer can return early. Test it by forcing the consumer to fail (for example an output that cannot be opened) with a time limit on the test.

Evidence#

  • dftracer-utils commit 754d239f fix(views): fail an indexed export that cannot open its output added the drain and the flag in view_export_index.cpp, with a regression test.
  • The test timed out with the old consumer (no drain, no flag) and passed with the fix. The same command then failed in about 1 second with an IO error.
  • The same risk remained in dftracer_pgzip: it can block if every compressor worker fails. It was not fixed in that change.