Skip to content

[Humble] Regression in tf2 0.25.21-0.25.23: process-wide stall after transient TF starvation (buffer never recovers) #995

Description

@ThyJoy

Bug report

Environment

  • ROS 2 Humble, Ubuntu 22.04 (jammy), amd64, Docker containers with network_mode: host
  • rmw_cyclonedds_cpp (reproduced with both 1.3.4 and 1.3.5; RMW ruled out, see bisection)
  • Nav2 1.1.20 from apt, running in component_container_isolated (amcl, costmaps, collision monitor, planner, controller, etc. in one process)
  • Affected: ros-humble-tf2 0.25.23-1jammy.20260907.213559 (+ family)
  • Healthy: ros-humble-tf2 0.25.20-2jammy.20260605.130708 (+ family)
  • Indoor delivery robots (differential drive), WiFi networking, TF tree at ~10–20 Hz (odom→base from EKF, map→odom from AMCL)

Symptom

With tf2 0.25.23, within 1–3 minutes of driving, the whole nav2 component container process enters a permanent stall of message delivery:

  • local costmap's TF listener stops ingesting /tf: getRobotPose keeps returning the pose from the stall instant (published footprint frozen at a stale pose while live TF moves on);
  • costmap filter logs Lookup would require extrapolation into the future ... the latest data is at time <T> where <T> is one frozen timestamp repeated hundreds of times per minute;
  • AMCL stops broadcasting map→odom (its last transform carries the same frozen stamp <T>);
  • notably, non-TF subscriptions of the same process stall too: the collision monitor stops reacting to its cmd_vel input topic while the publisher keeps publishing at 15–20 Hz.

A freshly started node in the same container sees live /tf and all topics normally. Only restarting the process recovers. Nothing is logged at the stall instant.

Trigger

The environment occasionally produces short TF starvation while driving (WiFi hiccups; map→odom lags ~5 s). With tf2 0.25.20 these episodes self-heal (we see two ~5 s bursts of the same extrapolation error, then normal operation). With 0.25.23 an episode appears to wedge the process permanently — the frozen stamps in the log match the moment of such an episode. This looks like a lock/wait introduced between 0.25.20 and 0.25.23 that can deadlock or permanently block an executor thread, starving the process's callbacks.

Bisection matrix (two identical robots, same configs; only image contents varied)

tf2 rmw_cyclonedds_cpp everything else result
0.25.20 1.3.4 June sync (reference robot) healthy, weeks of operation
0.25.23 1.3.5 Sept sync permanent stall in 1–3 min of driving, reproduced 4+ times
0.25.23 1.3.4 (pinned) Sept sync permanent stall (RMW ruled out)
0.25.20 (image from reference robot) 1.3.4 June sync binaries + current configs healthy, real delivery run
0.25.20 (pinned .debs only) 1.3.5 Sept sync healthy, real delivery run

Additional data point: the reference robot runs tf2_ros 0.25.22 with tf2 0.25.20 and is healthy — so the suspect range is the tf2 core changes in 0.25.21–0.25.23, or tf2_ros exactly 0.25.23.

Expected behavior

Transient TF starvation must not permanently stall the buffer or the hosting process; 0.25.20 recovers within seconds on identical hardware/network.

Workaround

Pinning the tf2 family to 0.25.20 (.debs from the 2026-07-02 snapshot) on top of an otherwise current Humble install.

Happy to test patches or provide detailed logs / reproduction help.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions