liorso opened a new issue, #51547:
URL: https://github.com/apache/arrow/issues/51547

   ### Describe the bug, including details regarding any error messages, 
version, and platform.
   
   An R process running open_dataset(<feather files>, format = "arrow") |> 
filter(...) |> collect() occasionally freezes forever: 0% CPU, every thread 
sleeping, no error, not interruptible. It happens on a tiny scan (6 static 
feather files, ~1.4 MB total, plus a schema file) after the identical call 
succeeded many times in the same process. We have hit it four times in 
production-like workloads over three months on three different datasets.
   
   Environment: arrow R 12.0.1.1 (libarrow 12.0.1), R 4.0.5, Ubuntu 24.04, 
x86_64, local ext4 filesystem, arrow::set_io_thread_count(2) and 
set_cpu_count(2) set at startup.
   
   We attached gdb to a frozen process and took two thread apply all bt dumps 
60 s apart (identical). Full stacks attached (arrow_deadlock_stacks.txt); the 
relevant threads:
   
   * R main thread: Table__from_ExecPlanReader -> RunWithCapturedR -> 
SerialExecutor::RunLoop, waiting for the collect.
   * IO-pool worker A (running the collect, since RunWithCapturedRIfPossible 
submits it to io_context.executor()): RecordBatchReader::ToTable -> 
ExecPlanImpl::StartProducing -> 
MergedGenerator<EnumeratedRecordBatch>::operator() -> FragmentToBatches -> 
IpcFileFormat::ScanBatchesAsync -> dataset::OpenReader -> 
ipc::RecordBatchFileReader::Open -> FutureImpl::Wait(). It holds the 
MergedGenerator state mutex (taken in State::PullSource()) while blocking on 
the footer read, which is an io::RandomAccessFile::ReadAsync submitted to the 
IO pool.
   * IO-pool workers B and C: both inside future-completion callbacks 
(MarkFinished -> MergedGenerator::InnerCallback::operator() -> 
arrow::util::Mutex::Lock()), blocked on that same mutex.
   * One idle CPU-pool worker; everything else (jemalloc, gomp, signal thread) 
idle.
   So: the thread holding the merged-generator lock needs an IO-pool task to 
complete; every IO-pool thread is blocked on that lock. Classic lock-ordering / 
blocking-wait deadlock, no progress possible.
   
   ## The three ingredients (all still present on main as of 2026-09-27)
   1. cpp/src/arrow/dataset/file_ipc.cc, IpcFileFormat::ScanBatchesAsync: the 
file is opened with OpenReaderAsync, and then re-opened synchronously in the 
continuation to apply the projected schema:
   ```
   auto reopen_reader = [self, options, 
source](std::shared_ptr<ipc::RecordBatchFileReader> reader)
       -> Future<std::shared_ptr<ipc::RecordBatchFileReader>> {
     ARROW_ASSIGN_OR_RAISE(auto options, GetReadOptions(*reader->schema(), 
*self, *options));
     return OpenReader(source, options);   // synchronous
   };
   ```
   This has been there since 5.0.0 (ARROW-11772); the file is byte-identical 
between 12.0.1 and main.
   2. cpp/src/arrow/ipc/reader.cc, RecordBatchFileReaderImpl::ReadFooter() is 
ReadFooterAsync(nullptr).status(): a blocking wait on an IO-pool read, so the 
synchronous Open above blocks a pool thread.
   3. cpp/src/arrow/util/async_generator.h, 
MergedGenerator::State::PullSource() holds the state mutex while calling the 
source generator ("so we don't pull sync-reentrantly"), so the blocking wait in 
(1) happens under the lock that every completion callback needs.
   The R package makes the pool one thread smaller than it looks: 
r/src/safe-call-into-r.h RunWithCapturedRIfPossible runs the top-level collect 
on io_context.executor(). With io_thread_count = 2 that leaves one IO thread 
for actual I/O. This is the situation acknowledged in #36121 ("we hijack the IO 
thread pool ... some Arrow code makes the usually safe assumption that there is 
at least one available IO thread"), where the resolution (#36304) was to warn 
for num_threads < 2. Our experience is that 2 deadlocks too, just rarely; and 
nothing prevents it at larger sizes if enough fragments open concurrently.
   
   ## Suggested fix
   Make the re-open asynchronous. OpenReaderAsync(source, options) already 
exists in the same file with the right signature, so reopen_reader can return 
OpenReaderAsync(source, options) instead of OpenReader(source, options). That 
removes the blocking Wait() from under the MergedGenerator lock and from the IO 
pool. Longer term, ReadFooter()'s blocking wait should not be reachable from 
pool threads at all.
   
   
   ### Component(s)
   
   C++, R


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to