Research / Preprint
Abstract
Training models to act as agents that can effectively navigate and perform actions in a complex environment, such as a web browser, has typically been challenging due to lack of training data. Large language models (LLMs) have recently demonstrated some capability to navigate novel environments as agents in a zero-shot or few-shot fashion, purely guided by natural language instructions as prompts. Recent research has also demonstrated LLMs have the capability to exceed their base performance through self-improvement, i.e. fine-tuning on data generated by the model itself. In this work, we explore the extent to which LLMs can self-improve their performance as agents in long-horizon tasks in a complex environment using the WebArena benchmark. In WebArena, an agent must autonomously navigate and perform actions on web pages to achieve a specified objective. We explore fine-tuning on three distinct synthetic training data mixtures and achieve a 31% improvement in task completion rate over the base model on the WebArena benchmark through a self-improvement procedure. We additionally contribute novel evaluation metrics for assessing the performance, robustness, capabilities, and quality of trajectories of our fine-tuned agent models to a greater degree than simple, aggregate-level benchmark scores currently used to measure self-improvement.
Cite this paper
@misc{patel2024large,
author = {Patel, Ajay and Hofmarcher, Markus and Leoveanu-Condrei, Claudiu and Dinu, Marius-Constantin and Callison-Burch, Chris and Hochreiter, Sepp},
title = {Large Language Models Can {Self-Improve} At Web Agent Tasks},
howpublished = {arXiv:2405.20309},
eprint = {2405.20309},
archivePrefix = {arXiv},
year = {2024},
date = {2024-10-01},
pagetotal = {27},
url = {https://www.dinu.at/research/large-language-models-can-self-improve-at-web-agent-tasks},
}Ajay Patel, Markus Hofmarcher, Claudiu Leoveanu-Condrei, Marius-Constantin Dinu, Chris Callison-Burch, Sepp Hochreiter. (2024, October 1). Large Language Models Can Self-Improve At Web Agent Tasks. arXiv:2405.20309. 27 pages. https://www.dinu.at/research/large-language-models-can-self-improve-at-web-agent-tasks
TY - RPRT AU - Patel, Ajay AU - Hofmarcher, Markus AU - Leoveanu-Condrei, Claudiu AU - Dinu, Marius-Constantin AU - Callison-Burch, Chris AU - Hochreiter, Sepp TI - Large Language Models Can Self-Improve At Web Agent Tasks PB - arXiv PY - 2024 DA - 2024/10/01/ SP - 27 UR - https://www.dinu.at/research/large-language-models-can-self-improve-at-web-agent-tasks N1 - arXiv:2405.20309 ER -
More preprint
- Preprint · 2026 Software Commoditization in the AI-First Economy: From Implementation Moats to Contribution Markets
- Preprint · 2025 HyDRA: A Hybrid-Driven Reasoning Architecture for Verifiable Knowledge Graphs
- Preprint · 2025 Primality Testing via Circulant Matrix Eigenvalue Structure: A Novel Approach Using Cyclotomic Field Theory
- Preprint · 2025 Ringdown Bounds on UV-Regularized Black-Hole Cores