Uploaded July 2025 | Updated September 2026, 2 weeks ago
Recent work have found success by separating high-level planning from low-level execution, which leads to better alignment between high-level goals and environment-specific actions. However, planning remains inherently difficult for tasks like web navigation, often due to complex environmental dynamics and insufficient training data. To address these issues, Amir Gholami -- Research Scientist, UC Berkeley -- proposes Plan-and-Act, a novel framework that incorporates explicit planning into LLM-based agents. Plan-and-Act consists of a Planner model that generates high-level plans to achieve user goals and an Executor model that translates these plans into precise, environment-specific actions. By explicitly separating planning from execution, Plan-and- Act improves alignment between the high-level reasoning and low-level actions, which enables consistent decision-making and adaptability to changes in the environment. To train the Planner effectively, the team behind the framework introduce a synthetic data generation method that annotates ground-truth trajectories with feasible plans, augmented by diverse and extensive examples to enhance generalization. They then evaluate PLAN-AND-ACT using web navigation as a representative long-horizon planning environment, demonstrating a state-of-the-art 57.58% success rate on the WebArena-Lite benchmark as well as a text-only state-of-the-art 81.36% success rate on WebVoyager.
Chapters
00:00 - Why agents fail at long-horizon tasks
01:15 - Introducing Plan-and-Act
03:40 - LPUs: Treating LLMs like CPUs
06:00 - The LLM Compiler (ICML 2024)
08:10 - Planning remains hard for LLMs
09:50 - Using synthetic data to train better planners
11:25 - Step-by-step: how Plan-and-Act is trained
13:00 - Benchmark results: WebArena, WebVoyager
14:30 - Real-world demo: Narada AI for enterprise workflows
17:00 - Automating tools like Concur, Salesforce, Hubspot
19:20 - Vision for agent abstraction beyond SaaS
21:00 - Summary and enterprise invite
đź§ Related Papers:
Plan-and-Act: arxiv.org/abs/2503.09572
LLMCompiler: arxiv.org/abs/2312.04511
LLM2LLM: arxiv.org/abs/2403.15042
đź’ˇ Want to join the conversation? Attend upcoming events at the Arize AI community:
👉 arize.com/community
Recent work have found success by separating high-level planning from low-level execution, which leads to better alignment between high-level goals and environment-specific actions. However, planning remains inherently difficult for tasks like web navigation, often due to complex environmental dynamics and insufficient training data. To address these issues, Amir Gholami -- Research Scientist, UC Berkeley -- proposes Plan-and-Act, a novel framework that incorporates explicit planning into LLM-based agents. Plan-and-Act consists of a Planner model that generates high-level plans to achieve user goals and an Executor model that translates these plans into precise, environment-specific actions. By explicitly separating planning from execution, Plan-and- Act improves alignment between the high-level reasoning and low-level actions, which enables consistent decision-making and adaptability to changes in the environment. To train the Planner effectively, the team behind the framework introduce a synthetic data generation method that annotates ground-truth trajectories with feasible plans, augmented by diverse and extensive examples to enhance generalization. They then evaluate PLAN-AND-ACT using web navigation as a representative long-horizon planning environment, demonstrating a state-of-the-art 57.58% success rate on the WebArena-Lite benchmark as well as a text-only state-of-the-art 81.36% success rate on WebVoyager.
Chapters
00:00 - Why agents fail at long-horizon tasks
01:15 - Introducing Plan-and-Act
03:40 - LPUs: Treating LLMs like CPUs
06:00 - The LLM Compiler (ICML 2024)
08:10 - Planning remains hard for LLMs
09:50 - Using synthetic data to train better planners
11:25 - Step-by-step: how Plan-and-Act is trained
13:00 - Benchmark results: WebArena, WebVoyager
14:30 - Real-world demo: Narada AI for enterprise workflows
17:00 - Automating tools like Concur, Salesforce, Hubspot
19:20 - Vision for agent abstraction beyond SaaS
21:00 - Summary and enterprise invite
đź§ Related Papers:
Plan-and-Act: arxiv.org/abs/2503.09572
LLMCompiler: arxiv.org/abs/2312.04511
LLM2LLM: arxiv.org/abs/2403.15042
đź’ˇ Want to join the conversation? Attend upcoming events at the Arize AI community:
👉 arize.com/community










