CoMAP: Co-Evolving World Models and Agent Policies for LLM Agents
Language agents often lack predictive capabilities in interactive environments. Existing textual world models are fixed after training, failing to adapt to distribution shifts from evolving policies. CoMAP co-evolves world models and agent policies via closed-loop interaction, dynamically updating models with self-distillation, improving long-horizon tasks, offering new optimization without external rewards.
01 ABSTRACT
The paper proposes CoMAP, where world models predict future state feedback, agents reflect and refine actions, and resulting on-policy trajectories update models via self-distillation. Experiments show superiority over baselines in embodied, web, and tool benchmarks, e.g., +16.75% relative with Qwen3-4B. Authors claim co-evolution improves prediction accuracy and long-horizon decision-making.
02 KEY FINDINGS
- CoMAP co-evolves world models and agent policies via closed-loop without external rewards
- World model predicts future feedback; agent does future-aware reflection; self-distillation updates model
- Outperforms baselines on embodied, web, and tool benchmarks, e.g., +16.75% relative with Qwen3-4B
- Co-evolution improves world model prediction accuracy and long-horizon decision-making over time
AI GENERATED SUMMARY / DISCOVERED BY ARXIV CS.CL