Excessive IO caused by systemd-journald #40262
Description
opened on Jan 3on Jan 3, 2026
Last edited by XANi
systemd version the issue has been seen with
257.9
Used distribution
Debian 13
Linux kernel version used
6.12.57+deb13-amd64
Component
systemd-journald
Expected behaviour you didn't see
Log writes should be within order of magnitude of syslog
Unexpected behaviour you saw
VM doing ~50 IOPS when writing 2 lines of log per second
Steps to reproduce the problem
this is exactly same issue #15292 that was closed without good reason
Step 1. Use journald in mode where it writes to hard drive. The FS is XFS
Step 2. have constant stream of log entries going on a VM
Jan 03 13:37:01 cthylla haproxy[727]: 192.168.1.1:48550 [03/Jan/2026:13:37:01.392] f_www b_icinga/web 0/0/0/6/6 302 153 - - ---- 2/2/0/0/0 0/0 "GET / HTTP/1.0"
Jan 03 13:37:03 cthylla haproxy[727]: 192.168.1.1:36892 [03/Jan/2026:13:37:03.403] f_www b_icinga/web 0/0/0/7/7 302 153 - - ---- 2/2/0/0/0 0/0 "GET / HTTP/1.0"
Jan 03 13:37:05 cthylla haproxy[727]: 192.168.1.1:36904 [03/Jan/2026:13:37:05.416] f_www b_icinga/web 0/0/0/6/6 302 153 - - ---- 2/2/0/0/0 0/0 "GET / HTTP/1.0"
Jan 03 13:37:07 cthylla haproxy[727]: 192.168.1.1:36906 [03/Jan/2026:13:37:07.427] f_www b_icinga/web 0/0/0/6/6 302 153 - - ---- 2/2/0/0/0 0/0 "GET / HTTP/1.0"
Jan 03 13:37:09 cthylla haproxy[727]: 192.168.1.1:36912 [03/Jan/2026:13:37:09.439] f_www b_icinga/web 0/0/0/7/8 302 153 - - ---- 2/2/0/0/0 0/0 "GET / HTTP/1.0"
Jan 03 13:37:11 cthylla haproxy[727]: 192.168.1.1:36918 [03/Jan/2026:13:37:11.454] f_www b_icinga/web 0/0/0/6/6 302 153 - - ---- 2/2/0/0/0 0/0 "GET / HTTP/1.0"
Jan 03 13:37:13 cthylla haproxy[727]: 192.168.1.1:45832 [03/Jan/2026:13:37:13.465] f_www b_icinga/web 0/0/0/7/7 302 153 - - ---- 2/2/0/0/0 0/0 "GET / HTTP/1.0"
Jan 03 13:37:15 cthylla haproxy[727]: 192.168.1.1:45848 [03/Jan/2026:13:37:15.476] f_www b_icinga/web 0/0/0/6/6 302 153 - - ---- 2/2/0/0/0 0/0 "GET / HTTP/1.0"
Jan 03 13:37:17 cthylla haproxy[727]: 192.168.1.1:45856 [03/Jan/2026:13:37:17.488] f_www b_icinga/web 0/0/0/6/6 302 153 - - ---- 2/2/0/0/0 0/0 "GET / HTTP/1.0"
Jan 03 13:37:19 cthylla haproxy[727]: 192.168.1.1:45862 [03/Jan/2026:13:37:19.500] f_www b_icinga/web 0/0/0/6/6 302 153 - - ---- 2/2/0/0/0 0/0 "GET / HTTP/1.0"
Step 3. Observe the VM IO traffic.

I used VM as example because the complaint in #15292 was "iotop is not accurate" (I can believe that, it's before any OS write coaelscing) but this clearly shows traffic after every kernel mechanism was used. So no, it isn't "kernel making lotsa iops out of it", it's slow.
Journald just uses extremely inefficient format (also I've seen it corrupt on unclean reboot enough times to declare it's not even all that resilient) as files are also multiple times the size of what's actually written in them.
Activity
ValdikSS commented on Jan 5on Jan 5, 2026
See #15292 (comment)
Try to disable compression as a first measure.
XANi commented on Jan 5on Jan 5, 2026
Author
@ValdikSS no meaningful difference in my case. Log entry might be short enough to not trigger compression
mAd-DaWg commented on Jan 8on Jan 8, 2026
on Jan 8on Jan 8, 2026 · Hidden as off-topic
Last edited by mAd-DaWg
joaociocca commented on Apr 7on Apr 7, 2026
Last edited by joaociocca
not sure it this is in anyway helpful, sorry in advance as I'm not an expert on this... but I've been having some weird slowdowns, and they match times when btop shows IO topping 100%.
I found out iotop has an accumulation mode, and in just 15 minutes journald had almost 7GB of data written to disk - if there's anyway I can help diagnose this and get rid of this problem that has been happening for at least a week, I'd be really happy to.

As you can see in that Bluesky post linked, at first I thought it was something with my GPU, that was usually at 50% but when the slowdowns happened it dropped a lot... but even in that post it's possible to see btop showing the IO topped.
(as I end writing this, 22 minutes of iotop show almost 11GB)
ValdikSS commented on Jan 5on Jan 5, 2026
Last edited by ValdikSS
I've spent some time to debug this issue.
It's not trivial since journald works only on mmap'ed files which are hard to debug, and can't use regular I/O means.
I've tried several approaches, ended up with the following:
- Create ext4 loop device (in /tmp tmpfs), dedicated to be mounted to
/var/log/journal - Log
loop0I/O stat file along withsystemd-journaldcgroupio.statfor this device
My configuration
- journald with
Compress=no, SyncIntervalSec=10sinjournald.conf - systemd with
DefaultIOAccounting=yesinsystem.conf - sysctl:
vm.dirty_writeback_centisecs = 100, vm.dirty_expire_centisecs = 200 - ext4 journal file created with
-E 'lazy_journal_init=0' -E 'lazy_itable_init=0' - absolutely "silent" system, nothing is logged into journald
My test suite
I've made a small python script to represent 7th value of block stat file of loop0 device (write sectors, sector here is always 512 bytes), along with /sys/fs/cgroup/system.slice/systemd-journald.service/io.stat data for loop0 device.
⇒ blockwrites-journal-loop.py ⇐
I've wrote into syslog with logger -p info test.
The amount of journald data with all internal fields for this log message is about 752 bytes.
# journalctl -o json -a --since 19:16:12 --until 19:16:12.9 | wc -c
752
# journalctl -o json -a --since 19:16:12 --until 19:16:12.9 | cat
{"__CURSOR":"s=be2e6c538ca748578da943ea92371828;i=188;b=cc21d029f31c449795207b313de3452f;m=f7436510;t=65826db61ff7a;x=38e7ed08484d1d59","_MACHINE_ID":"8baf193516d544518355beda01863de4","__SEQNUM_ID":"be2e6c538ca748578da943ea92371828","_COMM":"logger","_HOSTNAME":"fedora","_GID":"1000","PRIORITY":"6","_SELINUX_CONTEXT":"unconfined_u:unconfined_r:unconfined_t:s0-s0:c0.c1023","_RUNTIME_SCOPE":"system","__REALTIME_TIMESTAMP":"1785773772898170","__MONOTONIC_TIMESTAMP":"4148389136","_TRANSPORT":"syslog","__SOURCE_REALTIME_TIMESTAMP":"1785773772898102","__SEQNUM":"392","SYSLOG_IDENTIFIER":"valdikss","_UID":"1000","MESSAGE":"test","SYSLOG_TIMESTAMP":"Aug 3 19:16:12 ","_BOOT_ID":"cc21d029f31c449795207b313de3452f","SYSLOG_FACILITY":"1","_PID":"14004"}
Here are the results:
$ ./blockwrites-journal-loop.py
19:36:21: Started with 15675392 bytes loop0, 3855360 bytes journald cgroup
⇒ TERM2: logger -p info test
⇒ JOURNALTERM: авг 03 19:36:39 fedora valdikss[16919]: test
19:36:39: loop0: 15682560 bytes (+7168 bytes, +7168 bytes since start) | journald cgroup: 3859456 bytes (+4096 bytes, +4096 bytes since start)
19:36:42: loop0: 15748096 bytes (+65536 bytes, +72704 bytes since start) | journald cgroup: 3924992 bytes (+65536 bytes, +69632 bytes since start)
19:36:45: loop0: 15752192 bytes (+4096 bytes, +76800 bytes since start) | journald cgroup: 3924992 bytes (+0 bytes, +69632 bytes since start)
19:36:47: loop0: 15754240 bytes (+2048 bytes, +78848 bytes since start) | journald cgroup: 3924992 bytes (+0 bytes, +69632 bytes since start)
19:36:50: loop0: 15761408 bytes (+7168 bytes, +86016 bytes since start) | journald cgroup: 3929088 bytes (+4096 bytes, +73728 bytes since start)
19:36:52: loop0: 15762432 bytes (+1024 bytes, +87040 bytes since start) | journald cgroup: 3929088 bytes (+0 bytes, +73728 bytes since start)
⇒ TERM2: logger -p info test
⇒ JOURNALTERM: авг 03 19:37:21 fedora valdikss[16934]: test
19:37:22: loop0: 15769600 bytes (+7168 bytes, +94208 bytes since start) | journald cgroup: 3933184 bytes (+4096 bytes, +77824 bytes since start)
19:37:24: loop0: 15814656 bytes (+45056 bytes, +139264 bytes since start) | journald cgroup: 3978240 bytes (+45056 bytes, +122880 bytes since start)
19:37:28: loop0: 15818752 bytes (+4096 bytes, +143360 bytes since start) | journald cgroup: 3978240 bytes (+0 bytes, +122880 bytes since start)
19:37:30: loop0: 15820800 bytes (+2048 bytes, +145408 bytes since start) | journald cgroup: 3978240 bytes (+0 bytes, +122880 bytes since start)
19:37:32: loop0: 15827968 bytes (+7168 bytes, +152576 bytes since start) | journald cgroup: 3982336 bytes (+4096 bytes, +126976 bytes since start)
19:37:34: loop0: 15828992 bytes (+1024 bytes, +153600 bytes since start) | journald cgroup: 3982336 bytes (+0 bytes, +126976 bytes since start)
⇒ TERM2: logger -p info test
⇒ JOURNALTERM: авг 03 19:37:58 fedora valdikss[16947]: test
19:37:59: loop0: 15836160 bytes (+7168 bytes, +160768 bytes since start) | journald cgroup: 3986432 bytes (+4096 bytes, +131072 bytes since start)
19:38:01: loop0: 15881216 bytes (+45056 bytes, +205824 bytes since start) | journald cgroup: 4031488 bytes (+45056 bytes, +176128 bytes since start)
19:38:05: loop0: 15885312 bytes (+4096 bytes, +209920 bytes since start) | journald cgroup: 4031488 bytes (+0 bytes, +176128 bytes since start)
19:38:07: loop0: 15887360 bytes (+2048 bytes, +211968 bytes since start) | journald cgroup: 4031488 bytes (+0 bytes, +176128 bytes since start)
19:38:09: loop0: 15894528 bytes (+7168 bytes, +219136 bytes since start) | journald cgroup: 4035584 bytes (+4096 bytes, +180224 bytes since start)
19:38:11: loop0: 15895552 bytes (+1024 bytes, +220160 bytes since start) | journald cgroup: 4035584 bytes (+0 bytes, +180224 bytes since start)
⇒ TERM2: for i in {1..10}; do logger -p info test; done
⇒ JOURNALTERM: авг 03 19:37:58 fedora valdikss[16947]: test
авг 03 19:38:32 fedora valdikss[16958]: test
авг 03 19:38:32 fedora valdikss[16959]: test
авг 03 19:38:32 fedora valdikss[16960]: test
авг 03 19:38:32 fedora valdikss[16961]: test
авг 03 19:38:32 fedora valdikss[16962]: test
авг 03 19:38:32 fedora valdikss[16963]: test
авг 03 19:38:32 fedora valdikss[16964]: test
авг 03 19:38:32 fedora valdikss[16965]: test
авг 03 19:38:32 fedora valdikss[16966]: test
авг 03 19:38:32 fedora valdikss[16967]: test
19:38:33: loop0: 15902720 bytes (+7168 bytes, +227328 bytes since start) | journald cgroup: 4039680 bytes (+4096 bytes, +184320 bytes since start)
19:38:35: loop0: 15980544 bytes (+77824 bytes, +305152 bytes since start) | journald cgroup: 4117504 bytes (+77824 bytes, +262144 bytes since start)
19:38:39: loop0: 15984640 bytes (+4096 bytes, +309248 bytes since start) | journald cgroup: 4117504 bytes (+0 bytes, +262144 bytes since start)
19:38:41: loop0: 15986688 bytes (+2048 bytes, +311296 bytes since start) | journald cgroup: 4117504 bytes (+0 bytes, +262144 bytes since start)
19:38:43: loop0: 15993856 bytes (+7168 bytes, +318464 bytes since start) | journald cgroup: 4121600 bytes (+4096 bytes, +266240 bytes since start)
19:38:45: loop0: 15994880 bytes (+1024 bytes, +319488 bytes since start) | journald cgroup: 4121600 bytes (+0 bytes, +266240 bytes since start)
⇒ TERM2: logger -p info test
⇒ JOURNALTERM: авг 03 19:39:02 fedora valdikss[16978]: test
19:39:03: loop0: 16002048 bytes (+7168 bytes, +326656 bytes since start) | journald cgroup: 4125696 bytes (+4096 bytes, +270336 bytes since start)
19:39:05: loop0: 16047104 bytes (+45056 bytes, +371712 bytes since start) | journald cgroup: 4170752 bytes (+45056 bytes, +315392 bytes since start)
19:39:09: loop0: 16051200 bytes (+4096 bytes, +375808 bytes since start) | journald cgroup: 4170752 bytes (+0 bytes, +315392 bytes since start)
19:39:11: loop0: 16053248 bytes (+2048 bytes, +377856 bytes since start) | journald cgroup: 4170752 bytes (+0 bytes, +315392 bytes since start)
19:39:13: loop0: 16060416 bytes (+7168 bytes, +385024 bytes since start) | journald cgroup: 4174848 bytes (+4096 bytes, +319488 bytes since start)
19:39:15: loop0: 16061440 bytes (+1024 bytes, +386048 bytes since start) | journald cgroup: 4174848 bytes (+0 bytes, +319488 bytes since start)
As you can see, either a single log message or 10 of similar log messages result in at least 55 KB of written data.
14 log messages resulted in 386 KB physical writes (according to block stat) and 319 KB writes from journald (according to cgroup block stat)
Comparison
Here I simulated what syslog goes: appending text to the text file, with fdatasync+fsync.
$ ./blockwrites-journal-loop.py
20:04:59: Started with 17602560 bytes loop0, 5547008 bytes journald cgroup
⇒ TERM2: echo 'авг 03 ZZ:XX:YY fedora valdikss[16978]: test' | dd conv=notrunc,fdatasync,fsync oflag=append of=testfile
20:05:34: loop0: 17609728 bytes (+7168 bytes, +7168 bytes since start) | journald cgroup: 5547008 bytes (+0 bytes, +0 bytes since start)
20:05:36: loop0: 17613824 bytes (+4096 bytes, +11264 bytes since start) | journald cgroup: 5547008 bytes (+0 bytes, +0 bytes since start)
⇒ TERM2: echo 'авг 03 ZZ:XX:YY fedora valdikss[16978]: test' | dd conv=notrunc,fdatasync,fsync oflag=append of=testfile
20:06:10: loop0: 17617920 bytes (+4096 bytes, +15360 bytes since start) | journald cgroup: 5547008 bytes (+0 bytes, +0 bytes since start)
20:06:13: loop0: 17618944 bytes (+1024 bytes, +16384 bytes since start) | journald cgroup: 5547008 bytes (+0 bytes, +0 bytes since start)
⇒ TERM2: echo 'авг 03 ZZ:XX:YY fedora valdikss[16978]: test' | dd conv=notrunc,fdatasync,fsync oflag=append of=testfile
20:06:33: loop0: 17623040 bytes (+4096 bytes, +20480 bytes since start) | journald cgroup: 5547008 bytes (+0 bytes, +0 bytes since start)
20:06:35: loop0: 17624064 bytes (+1024 bytes, +21504 bytes since start) | journald cgroup: 5547008 bytes (+0 bytes, +0 bytes since start)
As you can see, it's most of the time a single ext4 block for write + 1024 bytes of something, much smaller than journald.
21 KB written overall.
Conclusion
I'm not really sure what could be the reason, it requires more debugging. Journald writes much more data than expected, but I did not face any I/O hammering (many frequent writes), just a much larger written data than expected.
Further testing:
- If it's mmap issue, incorrectly marking the region as dirty somewhere
- If it's fallocated/hole'd files issue, I had issues with these on ext4 previously
systemd/src/journal/journald-manager.c
Lines 2094 to 2104 in 1652be7
| fd=open(fn, O_RDWR | |
| if (fd<0) | |
| return-errno; | |
| r=posix_fallocate_loop(fd, 0, size); | |
| if (r<0) | |
| returnr; | |
| p=mmap(NULL, size, PROT_READ | |
| if (p==MAP_FAILED) | |
| return-errno; |
ValdikSS commented last weekon Aug 3, 2026
birdie-github commented last weekon Aug 3, 2026
@ValdikSS mad kudos for the analysis, much appreciated.
On all my systems where I don't really care about logs persistence (after reboot) I've long switched to:
[Journal]
Storage=volatile
Sorry for the noise.
ValdikSS commented last weekon Aug 3, 2026
Last edited by ValdikSS
This is btrfs, chattr +C /var/log/journal:
# ./blockwrites-journal-loop.py
⇒ TERM2: logger -p info test
⇒ JOURNALTERM: авг 03 21:41:38 fedora root[13468]: test
21:40:42: Started with 8482816 bytes loop0, 442368 bytes journald cgroup
21:41:38: loop0: 8556544 bytes (+73728 bytes, +73728 bytes since start) | journald cgroup: 479232 bytes (+36864 bytes, +36864 bytes since start)
21:41:48: loop0: 8765440 bytes (+208896 bytes, +282624 bytes since start) | journald cgroup: 552960 bytes (+73728 bytes, +110592 bytes since start)
21:42:09: loop0: 8904704 bytes (+139264 bytes, +421888 bytes since start) | journald cgroup: 552960 bytes (+0 bytes, +110592 bytes since start)
⇒ TERM2: logger -p info test
⇒ JOURNALTERM: авг 03 21:42:55 fedora root[13486]: test
21:42:56: loop0: 8978432 bytes (+73728 bytes, +495616 bytes since start) | journald cgroup: 589824 bytes (+36864 bytes, +147456 bytes since start)
21:43:06: loop0: 9195520 bytes (+217088 bytes, +712704 bytes since start) | journald cgroup: 663552 bytes (+73728 bytes, +221184 bytes since start)
21:43:26: loop0: 9228288 bytes (+32768 bytes, +745472 bytes since start) | journald cgroup: 663552 bytes (+0 bytes, +221184 bytes since start)
21:43:27: loop0: 9400320 bytes (+172032 bytes, +917504 bytes since start) | journald cgroup: 663552 bytes (+0 bytes, +221184 bytes since start)
⇒ TERM2: for i in {1..10}; do logger -p info test; done
⇒ JOURNALTERM: авг 03 21:43:53 fedora root[13507]: test
авг 03 21:43:53 fedora root[13508]: test
авг 03 21:43:53 fedora root[13509]: test
авг 03 21:43:53 fedora root[13510]: test
авг 03 21:43:53 fedora root[13511]: test
авг 03 21:43:53 fedora root[13512]: test
авг 03 21:43:53 fedora root[13513]: test
авг 03 21:43:54 fedora root[13514]: test
авг 03 21:43:54 fedora root[13515]: test
авг 03 21:43:54 fedora root[13516]: test
21:43:54: loop0: 9474048 bytes (+73728 bytes, +991232 bytes since start) | journald cgroup: 700416 bytes (+36864 bytes, +258048 bytes since start)
21:43:58: loop0: 9506816 bytes (+32768 bytes, +1024000 bytes since start) | journald cgroup: 700416 bytes (+0 bytes, +258048 bytes since start)
21:44:04: loop0: 9785344 bytes (+278528 bytes, +1302528 bytes since start) | journald cgroup: 774144 bytes (+73728 bytes, +331776 bytes since start)
21:44:25: loop0: 9924608 bytes (+139264 bytes, +1441792 bytes since start) | journald cgroup: 774144 bytes (+0 bytes, +331776 bytes since start)
XANi commented last weekon Aug 3, 2026
Author
There is not much to debug here, the ondisk format is just extremely wasteful and verbose, like every single message repeats boot ID and vast amount of other duplicated data. It's not indexed in any sensible way either so it's expensive to write and expensive to read
ValdikSS commented last weekon Aug 3, 2026
Oh my, journald writes every line to disk, SyncIntervalSec is only a fdatasync+fsync delay, not a data write delay.
I assumed that journald writes data to volatile storage for SyncIntervalSec, and only then writes the data to persistent files. But no, it writes to persistent every time any data appears in log, and fsync's it only after SyncIntervalSec.
On ext4, you might want to enable lazytime mount option: https://lwn.net/Articles/621046/
birdie-github commented last weekon Aug 4, 2026
I've long switched to noatime, probably two decades ago when I found out that even running find results in a massive amount of writes. On the kernel side the issue has long been addresses with, correct, lazytime but I have zero applications that need precise access times, so why use/have/enable it in the first place?
XANi commented last weekon Aug 4, 2026
Author
@birdie-github ok but that fixes basically nothing in the bug, as most of the logs will be recent so relatime takes care of that




