Skip to content

Research / Machine generated

Conservative Offline RL with Uncertainty-Aware Policy Improvement

Published with an anonymised author line — the document prints Anonymous authors / Paper under review. It is reproduced here exactly as generated.

Year
2026
Length
8 pages

How to read this

Written end to end by an agent. Published unedited, as evidence of what the system produces. It has not been reviewed, and no claim in it has been checked by a person. It is here because the interesting artefact is the process, not the result: this is what the system produces when it is pointed at a research question and left to run.

Abstract

We study conservative offline reinforcement learning with uncertainty-aware policy improvement under a tight compute budget. The goal is to combine conservative value regularization with ensemble-based uncertainty penalties and evaluate when such coupling improves mean performance, stability, and calibration. We design four hypothesis-driven experiments, including conservatism– uncertainty sweeps, checkpoint stability analysis, dataset-quality regime comparisons for implicit Q-learning, and correlation-based calibration of uncertainty penalties. Because full simulator access is unavailable, we report a transparent simulation-based validation using logged classic-control proxies that preserve the benchmark structure and metrics. The results show consistent gains in mean normalized score for uncertainty-augmented conservative learning, improved checkpoint stability, and positive uncertainty–density correlations, while variance reductions and dataset-quality effects are mixed. These findings motivate follow-on experiments on full D4RL benchmarks and provide a reproducible evaluation scaffold for conservative offline RL under strict resource constraints.