What changed
This research investigates the capability of coding agents to automate generalized Task and Motion Planning (TAMP). TAMP problems are inherently complex due to the tight coupling of discrete decisions with continuous geometric, kinematic, and dynamic constraints. Existing generalized TAMP methods often demand substantial TAMP-specific engineering. The study explores whether coding agents can synthesize programs that generalize across different problem instances, thereby reducing the need for manual engineering. The agents were tasked with interacting with a simulator to develop a program within a fixed synthesis budget. This program was then evaluated on unseen instances without further LLM involvement during the testing phase.
Agents and Environments
Two primary coding agent backends were evaluated: Claude Code (Opus 5) and Codex (utilizing GPT-5.6 Sol and GPT-6 Astra). These agents were tested on 28 simulated environments drawn from the KinDER benchmark and PDDLStream domains. The environments included kinematic and dynamic tasks in both 2D and 3D, with object counts exceeding those in the original benchmarks. The agents were provided with task descriptions and simulator access, and in some settings, also with environment source code.
Performance Comparison
The synthesized programs were compared against hand-engineered planners, a one-shot generation approach, and an LLM-based generalized planning baseline. Across the 16 environments where a planner was available, all three agent configurations outperformed the planners. Claude Code achieved an average success rate of 82%, while Codex with GPT-6 Astra reached 95% and with GPT-5.6 Sol achieved 56%. These figures are notably higher than the planners' average of 47%. The agents' programs maintained higher success rates and used an order of magnitude less computation per instance as object counts increased, whereas the planner's performance degraded.
Agentic Behavior and Novel Strategies
Analysis of the agents' interaction logs and synthesized programs revealed that they used targeted probing to infer environment dynamics and developed strategies that exploited regularities across instances. Examples include identifying blocking obstacles before planning a route and adapting fixed manipulation sequences. The agents also discovered novel strategies, such as non-prehensile maneuvers and unique uses of the environment layout, which were not present in prior published work on these benchmarks. This suggests a capacity for genuine physical reasoning beyond memorization.
Why it matters for builders
This work presents a significant potential shift in how generalized TAMP problems are approached. Instead of requiring deep domain expertise and extensive manual engineering for each new problem instance or environment, builders may be able to leverage coding agents to automate the generation of generalized planning solutions. This could drastically reduce development time and complexity, making advanced robotic and AI planning more accessible.
Practical impact
Builders working on robotics, automated systems, or any domain requiring complex task and motion planning should consider evaluating coding agents for their TAMP challenges. The research provides a strong baseline for using agents like Claude Code and Codex in this domain. Developers can explore using these agents to synthesize planning programs, especially for problems with varying object counts or configurations, and compare their performance against existing TAMP solvers. The release of the agents' code and prompts by the researchers offers a starting point for experimentation.
Caveats and source limits
The study was conducted on simulated environments, and performance in real-world physical systems may differ. While the agents outperformed hand-engineered planners and LLM-based baselines, some complex dynamic 3D environments, particularly those involving pouring or sweeping many small objects, remained largely unsolved in the main experimental setting. The research paper does not provide specific pricing for the evaluated models or details on their availability for direct integration into custom TAMP systems beyond the research context. Further investigation is needed to assess the robustness and scalability of these agentic approaches in diverse real-world scenarios.
Sources
Claim check: 5/5 supported claims - 5 evidence links - 90% avg confidence
- Coding agents, including Claude Code (Opus 5) and Codex (GPT-5.6 Sol, GPT-6 Astra), can automate generalized Task and Motion Planning (TAMP) by synthesizing programs that generalize across problem instances.supported - arxiv.org
- All three evaluated coding agent configurations outperformed hand-engineered planners, one-shot generation, and an LLM-based generalized planning baseline in mean success rates on generalized TAMP problems.supported - arxiv.org
- Claude Code averaged 82% success, and Codex with GPT-6 Astra averaged 95% success on generalized TAMP tasks where planners were available, compared to the planners' 47% average success.supported - arxiv.org
- As object counts increase, the agents' synthesized programs maintain higher success rates and use an order of magnitude less computation per instance compared to traditional planners.supported - arxiv.org
- Coding agents discovered novel strategies for TAMP, including non-prehensile maneuvers and unique uses of the environment layout, not previously documented in related benchmarks.supported - arxiv.org
Caveats
- The study was conducted in simulation.
- Performance metrics are based on simulated environments.
- These results are from simulated environments and specific benchmark configurations.
- The novelty is based on the researchers' assessment of prior published work.
- Single-source caution: verify critical details at the linked source.
Radar score 80/100 - how it was calculated
- Reliability 80: Research metadata source
- Freshness 50: Fresh research date
- Novelty 77: Research implementation signal
- Technical 84: Research technical evidence
- Developer 87: Research developer relevance
- Ecosystem 68: Research implementation signal
- Confidence 96: Claims have reliable evidence