A consumer that stops reading a bounded channel on error leaves its producers blocked forever
Symptom#
A pipeline with producers that send into a bounded channel and one consumer hangs forever at 0% CPU when the consumer fails, instead of reporting the error. In dftracer-utils, dftracer_view -o /dev/null on an indexed trace export hung forever when the output could not be opened, and the library could also abort.
Cause#
On the error, the consumer returned and stopped reading. The channel stayed full, so every producer waited in send for space that never came, and the task that waits for the producers never finished.
Fix#
On failure, the consumer must:
- Set a shared
failedflag that the producers check, so they stop producing early. - Keep reading (draining) the channel until all producers have closed it, and drop what it reads.
- Report the error after the drain.
Check every bounded-channel pipeline for this: find each place where a consumer can return early. Test it by forcing the consumer to fail (for example an output that cannot be opened) with a time limit on the test.
Evidence#
- dftracer-utils commit
754d239f fix(views): fail an indexed export that cannot open its outputadded the drain and the flag inview_export_index.cpp, with a regression test. - The test timed out with the old consumer (no drain, no flag) and passed with the fix. The same command then failed in about 1 second with an IO error.
- The same risk remained in
dftracer_pgzip: it can block if every compressor worker fails. It was not fixed in that change.