HACETTEPE UNIVERSITY - DEPARTMENT OF COMPUTER ENGINEERING
A complete, full-stack autonomous navigation system for the resource-constrained Duckiebot platform utilizing a LAN-optimized ROS architecture.
Deploying multi-layered autonomy on resource-constrained embedded platforms like the Duckiebot presents a significant engineering challenge. These platforms rely on a basic sensor array—a monocular camera, wheel encoders, and an IMU—that lacks native depth perception and is highly susceptible to cumulative errors. Attempting to execute sophisticated perception models and high-level control loops locally saturates the onboard processing unit at approximately 98% utilization. This computational bottleneck induces severe system latency, preventing the robot from reacting to dynamic obstacles or maintaining real-time closed-loop stability. The core problem is orchestrating a complex navigation stack—encompassing localization, vision, path planning, and obstacle avoidance—on hardware that cannot natively support monolithic execution.
To deliver a complete, multi-layered autonomous driving system, we developed a dual-module navigation architecture centered on robust path planning and reactive control. The core of the system is divided into two specialized modules. The first module handles continuous reactive control, utilizing a PID-driven visual lane-following system operating concurrently with a YOLOv5 model for object detection and parameterized evasive maneuvers. The second module manages complex spatial routing, employing a hybrid planner where the A* algorithm generates optimal global paths across a static cost-map, while a custom Dynamic Window Approach (DWA) is developed for local path execution and dynamic obstacle clearance. To execute this sophisticated autonomy stack seamlessly on constrained hardware, we implemented a distributed Robot Operating System (ROS) network. By strategically offloading the computationally intensive perception and planning tasks to a remote workstation, we successfully reduced the onboard CPU load to ~35%, ensuring stable, real-time execution.
Module 1: Real-time, vision-based lane-following system driven by a PID controller and a fine-tuned YOLOv5 model, fully integrated into our custom development application with live map-based monitoring.
Module 2: Combines a vision-based PID lane follower with a remote fine-tuned YOLOv5 model for real-time object detection and dynamic obstacle avoidance.
Module 3: Extends the custom application's live monitoring to support multi-vehicle tracking, displaying multiple Duckiebots simultaneously on the real-time map interface.
Module 4: Implements Yolo-based obstacle detection to dynamically maintain a safe following distance with the Duckiebot in front.
Visual highlights from our development process, track testing, and architecture mapping.
The autonomous navigation stack is engineered around a bifurcated, multi-layered architecture. By decoupling high-level cognitive tasks from low-level hardware actuation via a distributed ROS network, the system achieves real-time closed-loop stability while evaluating complex perception and planning algorithms.
Initial high-frequency spatial tracking is achieved through encoder-based dead-reckoning derived from a differential drive kinematic model. The linear velocity (v) and angular velocity (ω) are computed to propagate the robot's global pose (x, y, θ) over discrete time steps (Δt):
Where R is the wheel radius, L is the wheelbase, and ωr, ωl represent the angular velocities of the right and left wheels, respectively.
To mitigate the inherent cumulative drift associated with wheel slippage, the localization pipeline integrates an absolute correction mechanism. Upon visual detection of map-aligned ARTag/AprilTag fiducial markers, the perception node calculates an SE(3) transformation matrix to overwrite the odometry frame, ensuring persistent global accuracy without the computational overhead of full Monocular SLAM.
The perception layer is divided into traditional computer vision for structural extraction and deep learning for semantic object detection. Lane boundaries are isolated within the HSV color space using adaptive thresholding, processed through a Canny edge detector and Probabilistic Hough Transform, and subsequently fed into a Histogram Filter to continuously estimate lateral offset (d) and heading error (θ).
To maintain continuous trajectory alignment within a lane, a Proportional-Integral-Derivative (PID) controller calculates a steering command (angular velocity, ω) based on the estimated errors:
Where e(t) is a weighted combination of the lateral and heading errors, and Kp, Ki, Kd are the empirically tuned controller gains.
Concurrently, a fine-tuned YOLOv5s model handles dynamic object detection. To overcome the camera's limited field of view, we engineered a persistent obstacle memory mechanism. Detections are projected via ROS TF2 into global map coordinates and maintained via a moving average temporal filter. This custom node assigns semantic priors (safety margins) to classes and retains spatial memory of obstacles even after they leave the camera frame.
Goal-oriented navigation relies on a dual-layered planning architecture. Global routing is resolved using the A* algorithm over a discretized directed graph generated from the topological road map, utilizing Bezier curves for smooth intersection transitions and Manhattan distance heuristics.
For local execution, a custom, lightweight Dynamic Window Approach (DWA) operates at 10 Hz, forward-simulating kinematically feasible trajectories. Trajectories are evaluated against a multi-objective cost function to select optimal motor commands:
Where Cpath ensures lane adherence, Cobs applies a repulsive potential field for obstacle clearance, Rvel incentivizes forward progression, and Chead aligns the final angle with the global path.
This formulation enforces strict obstacle repulsion—applying infinite cost penalties for safety margin violations—while optimizing for path adherence and forward progression.
To monitor and control the distributed architecture, a custom cross-platform interface was developed for Web and Mobile environments. Utilizing mDNS for local network discovery, the frontend leverages Three.js to render a hardware-accelerated 3D environment.
Through a WebSocket bridge, the application receives high-frequency telemetry, including pose data, velocity, YOLOv5 bounding boxes, and the live camera feed. This enables operators to view live spatial tracking, input target coordinates for the A* planner, or seamlessly seize manual control via virtual joysticks without interrupting the underlying control loops.
The entire autonomy stack operates over a distributed Robot Operating System (ROS) network to prevent embedded hardware saturation. By strategically isolating low-level hardware actuation on the edge device and offloading computationally intensive perception (YOLOv5s) and trajectory planning nodes to a remote workstation, the system successfully reduces the onboard CPU load from ~98% to ~35%. This execution split ensures that the high-frequency PID steering controllers and safety deceleration modules maintain strict real-time responsiveness.
The implementation of our distributed navigation stack yielded highly stable, real-time autonomous performance, successfully bridging the gap between theoretical algorithms and resource-constrained edge hardware.
Initial monolithic testing revealed that executing perception and planning algorithms locally saturated the Duckiebot's CPU at approximately 98%, inducing severe thermal throttling and control instability. By strategically offloading these cognitive nodes via our ROS architecture, the onboard CPU load was reduced to a stable ~35%. A minimal communication latency was maintained, proving entirely sufficient to guarantee real-time closed-loop stability.
The PID-driven visual lane-following system successfully minimized lateral and heading errors to maintain smooth trajectory alignment. Crucially, the custom obstacle memory node decoupled spatial awareness from the camera's instantaneous field of view. This allowed the safety module to execute a smooth linear deceleration profile and halt the robot safely, even when dynamic obstacles temporarily exited the frame during tight maneuvers.
Operating on the discretized topological cost-map, the A* algorithm provided highly stable and computationally efficient global reference paths. The custom DWA local planner, operating at a high frequency of 10 Hz, successfully sampled kinematically feasible velocities. However, testing revealed a look-ahead tracking bug that currently limits dynamic path execution, highlighting a clear focal point for future development.
As anticipated, relying exclusively on wheel encoders resulted in significant cumulative spatial drift. By integrating periodic SE(3) pose corrections via ARTag fiducial markers, the system achieved precise global accuracy. This hybrid approach provided the spatial reliability necessary for complex routing while avoiding the massive computational overhead and latency associated with full Monocular SLAM.
The cross-platform Web/Mobile application successfully validated the distributed system's integration. The Three.js-rendered 3D environment accurately visualized live telemetry, YOLOv5 bounding boxes, and projected DWA trajectories. The system demonstrated high fault tolerance, allowing operators to seamlessly toggle between autonomous coordinate dispatching and manual overrides without disrupting the underlying control loops.
For future continuations of this project, the immediate priority should be resolving the DWA look-ahead tracking bug to stabilize dynamic path execution. Once resolved, the architecture could be expanded into a comprehensive "Live City" paradigm. Potential upgrades include translating visual markers for strict traffic rule compliance, leveraging DWA simulations for autonomous parking, and utilizing the existing telemetry stack to explore multi-robot swarm logic.
Supervisor: Özgür Erkent
2210356060
2210356102
2210765036
2210765027
2210765018