@samitouri / QOSamiQemu / commits / 544ddbb637

block: Never drop BLOCK_IO_ERROR with action=stop for rate limiting

Commit 2155d2dd introduced rate limiting for BLOCK_IO_ERROR to emit an event only once a second. This makes sense for cases in which the guest keeps running and can submit more requests that would possibly also fail because there is a problem with the backend. However, if the error policy is configured so that the VM is stopped on errors, this is both unnecessary because stopping the VM means that the guest can't issue more requests and in fact harmful because stopping the VM is an important state change that management tools need to keep track of even if it happens more than once in a given second. If an event is dropped, the management tool would see a VM randomly going to paused state without an associated error, so it has a hard time deciding how to handle the situation. This patch disables rate limiting for action=stop by not relying on the event type alone any more in monitor_qapi_event_queue_no_reenter(), but checking action for BLOCK_IO_ERROR, too. If the error is reported to the guest or ignored, the rate limiting stays in place. Fixes: 2155d2dd7f73 ('block-backend: per-device throttling of BLOCK_IO_ERROR reports') Signed-off-by: Kevin Wolf <kwolf@redhat.com> Message-ID: <20260304122800.51923-1-kwolf@redhat.com> Signed-off-by: Kevin Wolf <kwolf@redhat.com>

Kevin Wolf committed Mar 4, 2026 at 13:28 UTC 544ddbb6373d61292a0e2dc269809cd6bd5edec6
2 files changed +21 -2
monitor/monitor.c
+20 -1
@@ -367,14 +367,33 @@ monitor_qapi_event_queue_no_reenter(QAPIEvent event, QDict *qdict)
367 {
368 MonitorQAPIEventConf *evconf;
369 MonitorQAPIEventState *evstate;
370 + bool throttled;
371
372 assert(event < QAPI_EVENT__MAX);
373 evconf = &monitor_qapi_event_conf[event];
374 trace_monitor_protocol_event_queue(event, qdict, evconf->rate);
375 + throttled = evconf->rate;
376 +
377 + /*
378 + * Rate limit BLOCK_IO_ERROR only for action != "stop".
379 + *
380 + * If the VM is stopped after an I/O error, this is important information
381 + * for the management tool to keep track of the state of QEMU and we can't
382 + * merge any events. At the same time, stopping the VM means that the guest
383 + * can't send additional requests and the number of events is already
384 + * limited, so we can do without rate limiting.
385 + */
386 + if (event == QAPI_EVENT_BLOCK_IO_ERROR) {
387 + QDict *data = qobject_to(QDict, qdict_get(qdict, "data"));
388 + const char *action = qdict_get_str(data, "action");
389 + if (!strcmp(action, "stop")) {
390 + throttled = false;
391 + }
392 + }
393
394 QEMU_LOCK_GUARD(&monitor_lock);
395
377 - if (!evconf->rate) {
396 + if (!throttled) {
397 /* Unthrottled event */
398 monitor_qapi_event_emit(event, qdict);
399 } else {
qapi/block-core.json
+1 -1
@@ -5794,7 +5794,7 @@
5794 # .. note:: If action is "stop", a `STOP` event will eventually follow
5795 # the `BLOCK_IO_ERROR` event.
5796 #
5797 -# .. note:: This event is rate-limited.
5797 +# .. note:: This event is rate-limited, except if action is "stop".
5798 #
5799 # Since: 0.13
5800 #