Document Type : Technical Paper

Author

Chabahar Maritime University

10.22044/jadm.2026.18083.2984

Abstract

Reliable reference tracking in an inverted-pendulum system requires
accurate motion control while keeping the pendulum upright and limiting
control effort. This study compares three control architectures for a non
linear spring-coupled two-cart inverted pendulum: a model-based feedfor
ward linear-quadratic regulator (FF-LQR), Proximal Policy Optimization
that directly commands the applied force (PPO-balanced), and Twin De
layed Deep Deterministic Policy Gradient that learns a bounded correction
to FF-LQR (Residual TD3). The learned components use the same ob
servation space, reward function, nominal training target, and evaluation
protocol, while all three controllers share the same total-force limit. Each
learned method is trained with three independent seeds. Under nominal
analytical conditions, FF-LQR provides the most accurate tracking, PPO
balanced produces the smallest pendulum excursions with the lowest con
trol effort, and Residual TD3 balances tracking accuracy, upright regula
tion, and actuation. The frozen controllers are then evaluated without re
training in a dynamics-matched MuJoCo model under actuator lag, friction,
measurement noise, observation delay, parameter variations, and their com
bined effect. Under the combined condition, FF-LQR and Residual TD3
complete all trials, whereas the PPO-balanced policies terminate early. Rel
ative to FF-LQR, Residual TD3 reduces the tracking RMSE from 0.101
m to 0.072 m while also reducing pendulum excursion and control effort.
The results suggest that bounded residual learning can improve an existing
model-based controller while keeping the learned correction within a clearly
defined limit.

Keywords

Main Subjects