Dozhdikov A.V. and Sitkovskiy A.M. (2026) Evolutionary multi-agent reinforcement learning for crisis-aware demographic policy optimization. Front. Big Data 9:1842233. doi: 10.3389 ... Dozhdikov A.V. and Sitkovskiy A.M. (2026) Evolutionary multi-agent reinforcement learning for crisis-aware demographic policy optimization. Front. Big Data 9:1842233. doi: 10.3389/fdata.2026.1842233ISSN 2624-909XDOI https://doi.org/10.3389/fdata.2026.1842233Posted on site: 12.08.26Òåêñò ñòàòüè íà ñàéòå æóðíàëà URL: https://www.frontiersin.org/journals/big-data/articles/10.3389/fdata.2026.1842233/full (äàòà îáðàùåíèÿ 12.08.2026)AbstractDemographic systems face unprecedented challenges from simultaneous crises. Conventional statistical demography techniques and agent–based models often struggle to capture nonlinear inter–regional interactions during periods of severe socio–economic disruption. To address this, we propose MADDPG–EVO–DGM, a hybrid algorithm that integrates multi–agent deep reinforcement learning with evolutionary optimisation and meta–learning principles to model regional demographic processes under multiple crisis scenarios. Each region is treated as an autonomous agent learning to steer demographic policy levers, while periodic evolutionary “boosters” overcome local optima via population–based perturbations of actor network parameters. Additionally, a Darwin–Gödel Machine–inspired meta–learning mechanism adapts the booster triggers, enabling self–improvement in the learning process. We evaluate MADDPG–EVO–DGM on a simulation environment calibrated with real demographic data for eight federal regions of the Russian Federation over the period 2000–2024 and subject to ten concurrent crisis scenarios (e.g., pandemic, geopolitical conflict, economic collapse). Experiments demonstrate significantly faster convergence and improved performance over a baseline MADDPG: the hybrid approach achieves a higher final average reward (252.57 vs. 243.07) and 3.4 × lower convergence variance (σ = 0.24 vs. 0.80), indicating more reliable training. It also exhibits qualitative performance jumps of +68% during evolutionary phases and maintains 35%–45% greater resilience under crisis shocks compared to the baseline. To our knowledge, this is the first application of multi–agent reinforcement learning to large–scale demographic modeling under crises, opening new possibilities for evidence–based, crisis–resilient population policy design. Code, data, and logs are provided to ensure reproducibility.