Skip to content

Deadlock between tf2_ros::Buffer::waitForTransform() and tf2::BufferCore::testTransformableRequests() when using MessageFilter #992

Description

@CHAIJZ888

Generated by Generative AI

Yes. This report was drafted with AI assistance, but all runtime logs, package versions, GDB backtraces, and workaround results come from the real affected system.

Operating System:

Ubuntu 22.04 (Jammy), x86_64

ROS version or commit hash:

ROS 2 Humble Installed binary package versions: ros-humble-geometry2: 0.25.23-1jammy.20260908.012033 ros-humble-tf2: 0.25.23-1jammy.20260907.213559 ros-humble-tf2-ros: 0.25.23-1jammy.20260907.224752

RMW implementation (if applicable):

rmw_fastrtps_cpp

RMW Configuration (if applicable):

Default Fast DDS configuration. No custom RMW XML configuration is used.

Client library (if applicable):

rclcpp (C++)

'ros2 doctor --report' output

ros2 doctor --report

Steps to reproduce issue

  1. Run ROS 2 Humble Nav2 with a local/global costmap containing
    nav2_costmap_2d::ObstacleLayer.

  2. Configure the obstacle layer with a sensor_msgs/msg/LaserScan
    observation source. The layer creates:

    • a tf2_ros::TransformListener using its dedicated executor thread;
    • a tf2_ros::MessageFilter<sensor_msgs::msg::LaserScan> running in
      the costmap executor thread.
  3. Continuously publish /tf and LaserScan messages. The issue appears
    when a scan requires MessageFilter to wait for a transform that is
    not immediately available.

  4. Stop Nav2 through its lifecycle shutdown path or send SIGINT.

  5. planner_server and controller_server sometimes fail to terminate.
    The launch system waits for 5 seconds and then escalates from SIGINT
    to SIGTERM.

The problem was repeatedly reproduced in both the local and global
costmaps.

A representative launch symptom is:

[ERROR] [planner_server]: process[planner_server] failed to terminate
'5' seconds after receiving 'SIGINT', escalating to 'SIGTERM'

[ERROR] [controller_server]: process[controller_server] failed to
terminate '5' seconds after receiving 'SIGINT', escalating to 'SIGTERM'

The lifecycle manager may also report:

Failed to change state for node: planner_server.
Exception: planner_server/change_state service client:
async_send_request failed.

The deadlock may be established before shutdown. Shutdown makes it
visible because Costmap2DROS::on_cleanup() waits for its executor
thread to finish.

Expected behavior

tf2_ros::MessageFilter and tf2_ros::TransformListener should be safe
to use from separate executor threads.

Receiving a transform while another thread registers a transformable
request must not cause a lock-order inversion.

During lifecycle cleanup, the costmap executor thread should terminate
and all Nav2 processes should exit cleanly after SIGINT without requiring
SIGTERM or SIGKILL.

Actual behavior

Two threads enter an apparent AB-BA lock inversion.

Thread A: costmap LaserScan / MessageFilter executor

The costmap executor is processing a LaserScan. It enters
tf2_ros::MessageFilter::add(), calls Buffer::waitForTransform(), and
blocks while trying to enter BufferCore::addTransformableRequest():

pthread_mutex_lock
tf2::BufferCore::addTransformableRequest(...)
tf2_ros::Buffer::waitForTransform(...)
tf2_ros::MessageFilter<sensor_msgs::msg::LaserScan>::add(...)
message_filters::Subscriber<sensor_msgs::msg::LaserScan>::...
rclcpp::Executor::execute_subscription(...)
rclcpp::executors::SingleThreadedExecutor::spin()

Thread B: TransformListener executor

At the same time, the dedicated TransformListener thread receives a TF
message. It enters BufferCore::testTransformableRequests() and blocks
inside the transform-ready callback:

pthread_mutex_lock
tf2_ros::Buffer transform-ready callback
tf2::BufferCore::testTransformableRequests()
tf2::BufferCore::setTransformImpl(...)
tf2::BufferCore::setTransform(...)
tf2_ros::TransformListener::subscription_callback(...)
rclcpp::Executor::execute_subscription(...)
rclcpp::executors::SingleThreadedExecutor::spin()

Main thread during cleanup

The main thread waits indefinitely for the costmap executor thread:

std::thread::join()
nav2_util::NodeThread::~NodeThread()
nav2_costmap_2d::Costmap2DROS::on_cleanup(...)
nav2_planner::PlannerServer::on_cleanup(...)

The same behavior was observed in controller_server.

Additional information

Suspected lock cycle

The installed geometry2 version is 0.25.23.

In tf2_ros::Buffer::waitForTransform(),
timer_to_request_map_mutex_ is held while calling
addTransformableRequest():

https://github.com/ros2/geometry2/blob/0.25.23/tf2_ros/src/buffer.cpp#L230-L258

addTransformableRequest() acquires
transformable_requests_mutex_:

https://github.com/ros2/geometry2/blob/0.25.23/tf2/src/buffer_core.cpp#L1182-L1252

In the opposite direction,
BufferCore::testTransformableRequests() holds
transformable_requests_mutex_ and
transformable_callbacks_mutex_ while directly invoking the registered
callback:

https://github.com/ros2/geometry2/blob/0.25.23/tf2/src/buffer_core.cpp#L1306-L1360

The callback created by Buffer::waitForTransform() attempts to acquire
timer_to_request_map_mutex_.

This appears to form the following cycle:

MessageFilter thread:
  timer_to_request_map_mutex_
    -> transformable_requests_mutex_

TransformListener thread:
  transformable_requests_mutex_
    -> transform callback
    -> timer_to_request_map_mutex_

Workaround and A/B result

As a diagnostic workaround, the LaserScan subscriber was connected
directly to the existing Nav2 obstacle-layer callback, bypassing only
tf2_ros::MessageFilter.

The scan was still projected to PointCloud2 and transformed by the
existing ObservationBuffer, so the rest of the costmap processing was
unchanged.

With the MessageFilter path bypassed:

  • planner_server terminated cleanly;
  • controller_server terminated cleanly;
  • the lifecycle manager terminated cleanly;
  • no SIGTERM escalation occurred.

This strongly isolates the failure to the asynchronous
MessageFilter -> waitForTransform -> addTransformableRequest path.

Disabling an unrelated custom voxel layer did not fix the problem while
the standard Nav2 ObstacleLayer LaserScan MessageFilter remained active.

Possible fix direction

Would it be acceptable for BufferCore::testTransformableRequests() to:

  1. determine completed requests while holding its internal mutexes;
  2. copy the callback and callback arguments;
  3. remove the completed request and callback from the internal
    containers;
  4. release all internal mutexes;
  5. invoke the callback only after the locks have been released?

Executing an external callback while holding
transformable_requests_mutex_ and
transformable_callbacks_mutex_ appears to be the part that completes
the lock cycle.

I can provide the complete GDB thread apply all bt output and the
corresponding Nav2 shutdown log. I can also help test a proposed patch on
the affected system.

planner_server.bt.txt

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions