Research / Machine generated
How to read this
Written end to end by an agent. Published unedited, as evidence of what the system produces. It has not been reviewed, and no claim in it has been checked by a person. It is here because the interesting artefact is the process, not the result: this is what the system produces when it is pointed at a research question and left to run.
Abstract
Goal-conditioned reinforcement learning often faces a practical tension: intrinsic novelty bonuses accelerate discovery in sparse and deceptive environments, but poorly controlled intrinsic coupling can distort the asymptotic objective. This paper introduces Curiosity-Conditioned Goal-Optimal Reinforcement Learning (CCGO-RL), a dual-value framework that treats curiosity as a controlled exploration mechanism inside a goal-conditioned control loop rather than as a permanent co-objective. The method combines an extrinsic value stream, an intrinsic value stream, confidence-bounded intrinsic annealing, and uncertainty-gated contrastive coupling. We formalize the setting with explicit decision variables (policy and scheduler), a measurable feasible scheduler set, and an optimality criterion based on extrinsic return. We prove a perturbation bound between mixed and extrinsic objectives and an extrinsic suboptimality envelope for mixed-objective optimization, and we derive a gate-sensitivity and conditional variance-increment bound for uncertainty-gated contrastive coupling. Empirically, synthetic benchmark suites covering sparse navigation, deceptive mazes, and goal-conditioned control show improved return-speed trade-offs and lower critic instability against strong baselines, with quantitative checks tied to theorem assumptions. The empirical gains are strongest for exploration efficiency and variance control, while one symbolic limit-corollary check remains mismatched, which narrows the interpretation of the asymptotic claim to the audited admissible schedule regime.
More machine generated
- Machine generated · 2026 A Contradiction-Aware Survey Framework for Multi-Objective Decision Support
- Machine generated · 2026 AutoTW-ASP: Automatic Low-Treewidth Encoding Synthesis and Backend Routing for Neurosymbolic ASP
- Machine generated · 2026 AutoTW-ASP: Automatic Low-Treewidth Rewrite Synthesis and Uncertainty-Aware Backend Routing for Exact Neurosymbolic ASP Training
- Machine generated · 2026 Benchmarking and Selecting State-of-the-Art Modern Fourier Transformation Methods