@samitouri / QOSamiQemu / commits / a3fcbca0ef

fuse: Copy write buffer content before polling

aio_poll() in I/O functions can lead to nested read_from_fuse_export() calls, overwriting the request buffer's content. The only function affected by this is fuse_write(), which therefore must use a bounce buffer or corruption may occur. Note that in addition we do not know whether libfuse-internal structures can cope with this nesting, and even if we did, we probably cannot rely on it in the future. This is the main reason why we want to remove libfuse from the I/O path. I do not have a good reproducer for this other than: $ dd if=/dev/urandom of=image bs=1M count=4096 $ dd if=/dev/zero of=copy bs=1M count=4096 $ touch fuse-export $ qemu-storage-daemon \ --blockdev file,node-name=file,filename=copy \ --export \ fuse,id=exp,node-name=file,mountpoint=fuse-export,writable=true \ & Other shell: $ qemu-img convert -p -n -f raw -O raw -t none image fuse-export $ killall -SIGINT qemu-storage-daemon $ qemu-img compare image copy Content mismatch at offset 0! (The -t none in qemu-img convert is important.) I tried reproducing this with throttle and small aio_write requests from another qemu-io instance, but for some reason all requests are perfectly serialized then. I think in theory we should get parallel writes only if we set fi->parallel_direct_writes in fuse_open(). In fact, I can confirm that if we do that, that throttle-based reproducer works (i.e. does get parallel (nested) write requests). I have no idea why we still get parallel requests with qemu-img convert anyway. Also, a later patch in this series will set fi->parallel_direct_writes and note that it makes basically no difference when running fio on the current libfuse-based version of our code. It does make a difference without libfuse. So something quite fishy is going on. I will try to investigate further what the root cause is, but I think for now let's assume that calling blk_pwrite() can invalidate the buffer contents through nested polling. Cc: qemu-stable@nongnu.org Reviewed-by: Kevin Wolf <kwolf@redhat.com> Signed-off-by: Hanna Czenczek <hreitz@redhat.com> Message-ID: <20260309150856.26800-2-hreitz@redhat.com> Reviewed-by: Kevin Wolf <kwolf@redhat.com> Signed-off-by: Kevin Wolf <kwolf@redhat.com>

Hanna Czenczek committed Mar 9, 2026 at 16:08 UTC a3fcbca0ef643a8aecf354bdeb08b1d81e5b33e7
1 file changed +16 -1
block/export/fuse.c
+16 -1
@@ -301,6 +301,12 @@ static void read_from_fuse_export(void *opaque)
301 goto out;
302 }
303
304 + /*
305 + * Note that aio_poll() in any request-processing function can lead to a
306 + * nested read_from_fuse_export() call, which will overwrite the contents of
307 + * exp->fuse_buf. Anything that takes a buffer needs to take care that the
308 + * content is copied before potentially polling via aio_poll().
309 + */
310 fuse_session_process_buf(exp->fuse_session, &exp->fuse_buf);
311
312 out:
@@ -624,6 +630,7 @@ static void fuse_write(fuse_req_t req, fuse_ino_t inode, const char *buf,
630 size_t size, off_t offset, struct fuse_file_info *fi)
631 {
632 FuseExport *exp = fuse_req_userdata(req);
633 + QEMU_AUTO_VFREE void *copied = NULL;
634 int64_t length;
635 int ret;
636
@@ -638,6 +645,14 @@ static void fuse_write(fuse_req_t req, fuse_ino_t inode, const char *buf,
645 return;
646 }
647
648 + /*
649 + * Heed the note on read_from_fuse_export(): If we call aio_poll() (which
650 + * any blk_*() I/O function may do), read_from_fuse_export() may be nested,
651 + * overwriting the request buffer content. Therefore, we must copy it here.
652 + */
653 + copied = blk_blockalign(exp->common.blk, size);
654 + memcpy(copied, buf, size);
655 +
656 /**
657 * Clients will expect short writes at EOF, so we have to limit
658 * offset+size to the image length.
@@ -660,7 +675,7 @@ static void fuse_write(fuse_req_t req, fuse_ino_t inode, const char *buf,
675 }
676 }
677
663 - ret = blk_pwrite(exp->common.blk, offset, size, buf, 0);
678 + ret = blk_pwrite(exp->common.blk, offset, size, copied, 0);
679 if (ret >= 0) {
680 fuse_reply_write(req, size);
681 } else {