Navigating WebAI cover frame
Paper · 2023 · Master Thesis & ACM SAC 2024

Navigating
WebAI

Training agents to complete web tasks with large language models and reinforcement learning.

June 2022·Master Thesis·Maastricht University

A hybrid of supervised and reinforcement learning over the MiniWoB benchmark, with a hard look at whether prior models actually understand HTML or merely memorize where targets tend to live.

Navigating WebAI: Training Agents to Complete Web Tasks with Large Language Models and Reinforcement Learning

Abstract

Recent advancements in language models have demonstrated remarkable improvements in various natural language processing (NLP) tasks. Supervised learning (SL) approaches have achieved impressive performance while utilizing significantly less training data compared to previous methods. However, these SL-based models fall short when compared to reinforcement learning (RL) approaches, which have shown superior results. In this paper, we propose a novel approach that combines SL and RL techniques over the MiniWoB benchmark.

Furthermore, we address important limitations of previous works that claimed their models could understand HTML content. Our analysis reveals that these models tend to memorize the distributions of target elements rather than truly understanding the underlying HTML structure. To address this issue, we propose several methods aimed at rectifying this shortcoming and present a new baseline of results.

By integrating SL and RL techniques, we aim at leveraging the strengths of both approaches and overcome their individual limitations. Our experiments demonstrate that our proposed method outperforms previous SL approaches on some tasks using a fraction of the data, while also bridging the performance gap with RL models. Our best baseline models in SL achieve 43.58% average accuracy and combined with a multimodal RL approach scores 36.69% accuracy. The combination of these approaches presents a promising direction for future web navigation, and the study of individual tasks puts light on their limitations.

In conclusion, this paper introduces a hybrid approach that combines SL and RL techniques for language modeling over the MiniWoB benchmark. We highlight the limitations of previous works regarding the understanding of HTML content and propose effective methods to address this issue. The presented results showcase the effectiveness and limitations of our approach and set a new baseline for further exploration in the field of pre-trained language modeling for computer tasks.


Let's connect on LinkedIn.