A rsnapshot backup, that had been running for years, ran into an issue lately. The rsnapshot logs revealed there were problems during the rsync (using ssh) transfer:
[2026-09-30T01:31:00] ssh_dispatch_run_fatal: Connection to 192.168.7.11 port 22: message authentication code incorrect
[2026-09-30T01:31:00] rsync: connection unexpectedly closed (275529523 bytes received so far) [receiver]
[2026-09-30T01:31:00] rsync error: error in rsync protocol data stream (code 12) at io.c(232) [receiver=3.2.7]
[2026-09-30T01:31:00] rsync: [generator] write error: Broken pipe (32)
[2026-09-30T01:31:00] rsync error: unexplained error (code 255) at io.c(849) [generator=3.2.7]
[2026-09-30T01:31:00] rsync: connection unexpectedly closed (2425459 bytes received so far) [generator]
Let's have a closer look.
Rsnapshot uses rsync in the background - which uses SSH (by default) for data transfer.
Obviously the first check is to verify whether or not the backup server, where rsnapshot runs, can actually access the target server:
root@backup:~# ssh 192.168.7.11 'echo OK'
OK
Yes, the basics work. Now we can simulate a data transfer of a 4 GB file to a remote server:
root@backup:~# dd if=/dev/zero bs=1M count=4096 2>/dev/null | ssh 192.168.7.11 'cat >/dev/null'
client_loop: send disconnect: Broken pipe
This is basically confirming the error seen in the rsnapshot logs.
The tricky part is obviously finding the reason behind that error, that seems to have come all of a sudden. The OS of this backup server has been on Debian 12 for several months. However there has been a Kernel update the day before:
root@backup:~# cat /var/log/apt/history.log|grep linux-image -B 3
Start-Date: 2026-09-29 06:48:15
Commandline: /usr/bin/apt-get -y -o Dpkg::Options::=--force-confdef -o Dpkg::Options::=--force-confold -o DPkg::Lock::Timeout=60 dist-upgrade --auto-remove
Requested-By: ck (1000)
Install: linux-image-6.1.0-53-amd64:amd64 (6.1.187-1, automatic)
Upgrade: libssh2-1:amd64 (1.10.0-3+b1, 1.10.0-3+deb12u1), tzdata:amd64 (2026b-0+deb12u1, 2026c-0+deb12u1), libarchive13:amd64 (3.6.2-1+deb12u4, 3.6.2-1+deb12u5), liblzma5:amd64 (5.4.1-1+deb12u1, 5.4.1-1+deb12u2), libexpat1:amd64 (2.5.0-1+deb12u2, 2.5.0-1+deb12u3), zip:amd64 (3.0-13, 3.0-13+deb12u1), python3-httplib2:amd64 (0.20.4-3, 0.20.4-3+deb12u1), xz-utils:amd64 (5.4.1-1+deb12u1, 5.4.1-1+deb12u2), libssl3:amd64 (3.0.20-1~deb12u2, 3.0.22-1~deb12u1), linux-image-amd64:amd64 (6.1.180-1, 6.1.187-1), libevent-2.1-7:amd64 (2.1.12-stable-8, 2.1.12-stable-8+deb12u1), libpcre2-8-0:amd64 (10.42-1, 10.42-1+deb12u1), libevent-core-2.1-7:amd64 (2.1.12-stable-8, 2.1.12-stable-8+deb12u1), libpq5:amd64 (15.18-0+deb12u1, 15.19-0+deb12u1), openssl:amd64 (3.0.20-1~deb12u2, 3.0.22-1~deb12u1)
Could this have had an impact on the BNX network card?
root@backup:~# lspci -nn -s 04:00.0
04:00.0 Ethernet controller [0200]: Broadcom Inc. and subsidiaries NetXtreme II BCM5709 Gigabit Ethernet [14e4:1639] (rev 20)
A closer look into the Kernel ring buffer, specifically looking for bnx2, reveals something interesting:
root@backup:~# dmesg -T | grep -Ei \
'04:00.0|04:00.1|bnx2|AER|PCIe|Corrected|Non-Fatal|BadTLP|BadDLLP|Unsupported'
[Wed Sep 30 06:56:58 2026] sha1_ssse3 video wmi drm_display_helper cec aesni_intel rc_core drm_ttm_helper crypto_simd ttm cryptd intel_cstate drm_kms_helper acpi_ipmi intel_uncore i2c_algo_bit ipmi_si iTCO_wdt intel_pmc_bxt serio_raw ipmi_devintf pcspkr iTCO_vendor_support watchdog hpilo i7core_edac ipmi_msghandler evdev joydev acpi_power_meter pcc_cpufreq button sg drm fuse loop efi_pstore configfs ip_tables x_tables autofs4 ext4 crc16 mbcache jbd2 btrfs blake2b_generic zstd_compress raid10 raid456 async_raid6_recov async_memcpy async_pq async_xor async_tx xor raid6_pq libcrc32c crc32c_generic raid1 raid0 multipath linear md_mod hid_generic usbhid ata_generic dm_mod hid sd_mod t10_pi crc64_rocksoft crc64 crc_t10dif crct10dif_generic ata_piix libata uhci_hcd hpsa ehci_pci ehci_hcd crct10dif_pclmul scsi_transport_sas crct10dif_common crc32_pclmul psmouse usbcore crc32c_intel scsi_mod bnx2 lpc_ich usb_common scsi_common
[Wed Sep 30 06:56:58 2026] bnx2_start_xmit+0x31c/0x770 [bnx2]
Interestingly this is the exact same time that I tried the SSH transfer before:
root@backup:~# grep "ssh 192.168.7.11" /var/log/user.log
2026-09-30T06:56:43.464215+02:00 backup [audit ck/. as root/1239026 on pts/4/] /root: ssh 192.168.7.11 'echo OK'
2026-09-30T06:56:59.723615+02:00 backup [audit ck/. as root/1239026 on pts/4/] /root: dd if=/dev/zero bs=1M count=4096 2>/dev/null | ssh 192.168.7.11 'cat >/dev/null'
I blame the one second difference on the rsyslog handling and writing to log.
At this point we know: There seem to be issues with the bnx2 interface, or driver. But where to look deeper?
Let's look at the current interface options, that could interfere with data transfer and cause interface errors:
root@backup:~# ethtool -k bond1 | egrep 'tcp-segmentation|generic-segmentation|generic-receive'
tcp-segmentation-offload: on
tx-tcp-segmentation: on
generic-segmentation-offload: on
generic-receive-offload: on
Let's disable one by one and re-try the transfer over SSH.
First disable TSO (tcp-segmentation-offload):
root@backup:~# ethtool -K bond1 tso off
root@backup:~# dd if=/dev/zero bs=1M count=4096 2>/dev/null | ssh 192.168.7.12 'cat >/dev/null'
root@backup:~#
Hey, this worked without an error!
So it looks that the first option, TSO, was already the bad apple.
Now with data transfer working over SSH again, I manually launched the daily rsnapshot process. It ran through perfectly.
root@backup:~# tail -n 3 /var/log/rsnapshot.log
[2026-09-30T08:04:27] touch /backup/rsnapshot/daily.0/
[2026-09-30T08:04:27] rm -f /var/run/rsnapshot.pid
[2026-09-30T08:04:27] /usr/bin/rsnapshot -c /etc/rsnapshot.conf daily: completed successfully
As for the exact reason this is really tricky to say. The Kernel update the day prior of the backup certainly is a major hint. It seems that inside the Kernel some security fixes (between 6.1.180 and 6.1.187) touched the bnx2 driver, that could have caused this issue.
Disabling TSO on the affected interface helps as a workaround though.
To make this survive a reboot, I added this as post-up option in /etc/network/interfaces:
post-up /sbin/ethtool -K $IFACE tso off
No comments yet.
AI AWS Android Ansible Apache Apple Atlassian BSD Backup Bash Bluecoat CMS Chef Cloud Coding Consul Containers CouchDB DB DNS Databases Docker ELK Elasticsearch Filebeat FreeBSD Galera Git GlusterFS Grafana Graphics HAProxy HTML Hacks Hardware Icinga Influx Internet Java KVM Kibana Kodi Kubernetes LVM LXC Linux Logstash Mac Macintosh Mail MariaDB Minio MongoDB Monitoring Multimedia MySQL NFS Nagios Network Nginx OSSEC OTRS Observability Office OpenSearch PHP Perl Personal PostgreSQL PowerDNS Proxmox Proxy Python Rancher Rant Redis Roundcube SSL Samba Seafile Security Shell SmartOS Solaris Surveillance Systemd TLS Tomcat Ubuntu Unix VMware Varnish Virtualization Windows Wireless Wordpress Wyse ZFS Znuny Zoneminder