Classical control design usually starts from a mathematical model: we write equations describing how a system behaves, and then use those equations to design a controller.
But today we often have something else in abundance: data.
A robot can repeatedly interact with its environment. A building can record years of operational data. An industrial process can continuously measure its inputs and outputs. This raises a natural question: How can we learn a good control strategy directly from experience, without first identifying a complete model of the system?
The central idea behind reinforcement learning/ data-driven optimal control is to find an optimal policy
π* = arg minᵤ E[ Σ γᵏ ℓ(xₖ,uₖ) ],
which minimizes the cumulated cost incurred by the controlled system. The goal is therefore not simply to track a reference objective, but to learn how to behave optimally over time.
At the heart of many reinforcement-learning algorithms lies the Bellman principle: the best decision now is the one that combines its immediate cost with the best possible future policy,
V*(x) = minᵤ { ℓ(x,u) + γ E[V*(x⁺)] }.
Much of my research in this area has focused on solving this equation by using data on how the system interacts with the environment, despite the curse of dimensionality and uncertainties.
Linear programming formulations offer a conceptually elegant approach to solve the infinite-horizon, model-free, nonlinear, continuous space version of such equation. One appealing feature of this approach is that the Bellman equation can be replaced by inequalities. Once those inequalities are generated from experimental data, learning an optimal controller becomes an optimization problem. We can therefore move directly from observed transitions of the system to a controller.
However, in addition to the curse of dimensionality, their practical use is limited by the often poor scalability of the exact formulation, and the difficulty to obtain bounded solutions for a reasonable amount of data.
First, collecting enough data can be expensive. We show that, for linear systems, a relatively small but sufficiently rich experiment [1] can be used to generate many additional Bellman inequalities offline. In other words, one carefully designed experiment can contain much more information than the individual measured transitions may initially suggest. We then expand this idea to affine systems [2], extending the famous Willems' Lemma to this class of dynamical systems, and design estimators for the Bellman inequalities when the dynamics is stochastic.
Second, the resulting optimization problems can become very large or even unbounded. We develop relaxed formulations [3] that substantially reduce their size and require less data collection, especially for stochastic systems. We then focus on the geometry [4] behind these optimization problems to understand when a finite solution is guaranteed. More recently, we use moment-matching techniques [5] to extend boundedness guarantees to nonlinear problems and polynomial features.
We also work on the algorithms underlying dynamic programming itself. The mini-batch Bellman operator [6] provides a tunable compromise between sequential algorithms, which can converge quickly, and fully parallel algorithms, which exploit modern computing hardware. In a complementary direction, our PAGE-PG algorithm [7] reduces the variance of policy-gradient estimates, one of the main reasons why reinforcement learning can require so many interactions with the environment.
[1] On the synthesis of Bellman inequalities for data-driven optimal control
A. Martinelli, M. Gargiani,J. Lygeros
IEEE Conference on Decision and Control (CDC), 2021
[2] Data-driven optimal control of affine systems: A linear programming perspective
A. Martinelli, M. Gargiani, M. Draskovic, J. Lygeros
IEEE Control Systems Letters, vol. 6, pages 3092-3097, 2022
[3] Data-driven optimal control with a relaxed linear program
A. Martinelli, M. Gargiani and J. Lygeros
Automatica, vol. 136, art. 110052, 2022
[4] Data-driven optimal control via linear programming: Boundedness guarantees
L. Falconi*, A. Martinelli*, J. Lygeros
IEEE Transactions on Automatic Control, 70(3):1683-1697, 2025
[5] Bounded Linear Programs for Data-Driven Optimal Control via Moment-Matching
A. Martinelli, L. Pezzetti, N. Schmid, F. Dörfler, J. Lygeros
IEEE Control Systems Letters, vol. 10, pp. 2065-2070, 2026
[6] Parallel and flexible dynamic programming via the mini-batch operator
M. Gargiani, A. Martinelli, M. Ruts Martinez, J. Lygeros
IEEE Transactions on Automatic Control, 69(1):455-462, 2024
[7] PAGE-PG: A simple and loopless variance-reduced policy gradient method with probabilistic gradient estimation
M. Gargiani, A. Zanelli, A. Martinelli, T. Summers, J. Lygeros
International Conference on Machine Learning (ICML), PMLR 162:7223-7240, 2022