Why it matters
This research suggests that coding agents can automate significant engineering effort typically required for TAMP. Builders may find these agents capable of developing reusable planning solutions, potentially accelerating development in robotics and AI systems that require complex physical reasoning and task execution.

What changed

This research investigates the capability of coding agents to automate generalized Task and Motion Planning (TAMP). TAMP problems are inherently complex due to the tight coupling of discrete decisions with continuous geometric, kinematic, and dynamic constraints. Existing generalized TAMP methods often demand substantial TAMP-specific engineering. The study explores whether coding agents can synthesize programs that generalize across different problem instances, thereby reducing the need for manual engineering. The agents were tasked with interacting with a simulator to develop a program within a fixed synthesis budget. This program was then evaluated on unseen instances without further LLM involvement during the testing phase.

Agents and Environments

Two primary coding agent backends were evaluated: Claude Code (Opus 5) and Codex (utilizing GPT-5.6 Sol and GPT-6 Astra). These agents were tested on 28 simulated environments drawn from the KinDER benchmark and PDDLStream domains. The environments included kinematic and dynamic tasks in both 2D and 3D, with object counts exceeding those in the original benchmarks. The agents were provided with task descriptions and simulator access, and in some settings, also with environment source code.

Performance Comparison

The synthesized programs were compared against hand-engineered planners, a one-shot generation approach, and an LLM-based generalized planning baseline. Across the 16 environments where a planner was available, all three agent configurations outperformed the planners. Claude Code achieved an average success rate of 82%, while Codex with GPT-6 Astra reached 95% and with GPT-5.6 Sol achieved 56%. These figures are notably higher than the planners' average of 47%. The agents' programs maintained higher success rates and used an order of magnitude less computation per instance as object counts increased, whereas the planner's performance degraded.

Agentic Behavior and Novel Strategies

Analysis of the agents' interaction logs and synthesized programs revealed that they used targeted probing to infer environment dynamics and developed strategies that exploited regularities across instances. Examples include identifying blocking obstacles before planning a route and adapting fixed manipulation sequences. The agents also discovered novel strategies, such as non-prehensile maneuvers and unique uses of the environment layout, which were not present in prior published work on these benchmarks. This suggests a capacity for genuine physical reasoning beyond memorization.

Why it matters for builders

This work presents a significant potential shift in how generalized TAMP problems are approached. Instead of requiring deep domain expertise and extensive manual engineering for each new problem instance or environment, builders may be able to leverage coding agents to automate the generation of generalized planning solutions. This could drastically reduce development time and complexity, making advanced robotic and AI planning more accessible.

Practical impact

Builders working on robotics, automated systems, or any domain requiring complex task and motion planning should consider evaluating coding agents for their TAMP challenges. The research provides a strong baseline for using agents like Claude Code and Codex in this domain. Developers can explore using these agents to synthesize planning programs, especially for problems with varying object counts or configurations, and compare their performance against existing TAMP solvers. The release of the agents' code and prompts by the researchers offers a starting point for experimentation.

Caveats and source limits

The study was conducted on simulated environments, and performance in real-world physical systems may differ. While the agents outperformed hand-engineered planners and LLM-based baselines, some complex dynamic 3D environments, particularly those involving pouring or sweeping many small objects, remained largely unsolved in the main experimental setting. The research paper does not provide specific pricing for the evaluated models or details on their availability for direct integration into custom TAMP systems beyond the research context. Further investigation is needed to assess the robustness and scalability of these agentic approaches in diverse real-world scenarios.

Sources

Written with AI assistance from the linked sources; every claim below was checked against them automatically. How we produce articles.

Claim check: 5/5 supported claims - 5 evidence links - 90% avg confidence
  • Coding agents, including Claude Code (Opus 5) and Codex (GPT-5.6 Sol, GPT-6 Astra), can automate generalized Task and Motion Planning (TAMP) by synthesizing programs that generalize across problem instances.supported - arxiv.org
  • All three evaluated coding agent configurations outperformed hand-engineered planners, one-shot generation, and an LLM-based generalized planning baseline in mean success rates on generalized TAMP problems.supported - arxiv.org
  • Claude Code averaged 82% success, and Codex with GPT-6 Astra averaged 95% success on generalized TAMP tasks where planners were available, compared to the planners' 47% average success.supported - arxiv.org
  • As object counts increase, the agents' synthesized programs maintain higher success rates and use an order of magnitude less computation per instance compared to traditional planners.supported - arxiv.org
  • Coding agents discovered novel strategies for TAMP, including non-prehensile maneuvers and unique uses of the environment layout, not previously documented in related benchmarks.supported - arxiv.org

Caveats

  • The study was conducted in simulation.
  • Performance metrics are based on simulated environments.
  • These results are from simulated environments and specific benchmark configurations.
  • The novelty is based on the researchers' assessment of prior published work.
  • Single-source caution: verify critical details at the linked source.
Radar score 80/100 - how it was calculated
Reliability80
Freshness50
Novelty77
Technical84
Developer87
Ecosystem68
Confidence96
  • Reliability 80: Research metadata source
  • Freshness 50: Fresh research date
  • Novelty 77: Research implementation signal
  • Technical 84: Research technical evidence
  • Developer 87: Research developer relevance
  • Ecosystem 68: Research implementation signal
  • Confidence 96: Claims have reliable evidence
Share
XLinkedInHacker News

Related articles

Agents - Sep 29, 2026AgentWorks v1.25.133: AI Agents for Work AutomationAgentWorks is an open-source framework for creating AI agents that automate tasks. The latest release, v1.25.133, supports agents for Claude, ChatGPT, and Cursor, designed to pursue defined goals.AI Coding - Sep 29, 2026Pi Herdsman: Orchestrates Parallel Coding Agents with Nested DelegationPi Herdsman is a new extension for the Pi and Herdr AI development environments, enabling asynchronous subagents and fleet orchestration for parallel coding tasks. It allows for nested delegation, background work, and supervision of multiple agents within a coordinated hierarchy.Agents - Sep 29, 2026Nous AI Agent Framework Introduces Minsky's Society of MindThe Nous AI agent framework has been released, implementing Marvin Minsky's Society of Mind principles and decision intelligence for structured memory and learning. It aims to create agents that can think, learn, and grow beyond simple stateless responses.Agents - Sep 29, 2026Vulkgryph/Forge v0.5.2 ReleaseVulkgryph/Forge has released version v0.5.2, an autonomous AI coding agent designed to run locally. This agent supports various LLM backends, including Claude, ChatGPT, and any OpenAI-compatible endpoint.Agents - Sep 29, 2026Tektonix Autonomous Coding Agent v0.9.0-rc6Tektonix, an autonomous coding agent built on LangGraph and deepagents, has released version v0.9.0-rc6. This agent is designed to plan, build, verify, and ship production changes with an integrated human-in-the-loop review.AI Tools - Sep 29, 2026AgenticOS: Open-Source Platform for Building and Managing AI AgentsAgenticOS is a new open-source, self-hosted platform designed for building, running, and governing AI agents within an organization. It provides a unified environment for managing agent skills, context files, automations, and budgets, with a focus on auditability and control.