ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context
Abstract
Embodied agents now take on ever longer tasks. For long tasks, knowing only whether a task finally succeeds or fails says little; the steps along the way matter. Progress Reward Models (PRMs) score how far a task has come at every step, and serve as dense rewards, verifiers and monitors. Yet in long tasks the current frame alone often cannot tell how far the task has come, because progress depends on what happened before. We call this problem context-dependent progress estimation. Existing benchmarks on progress estimation mostly focus on short tasks whose progress can be read from the current observation, and whether PRMs can estimate progress when context is needed remains underexplored. We therefore build ContextProgress-Bench, with 24 manipulation tasks for 120 episodes. The benchmark covers three settings: (i) State Recall, where information needed for progress appeared earlier but is not in the current frame; (ii) Sequence Tracking, where steps follow a fixed order, so progress requires knowing which steps are done and which comes next; and (iii) Recurrence Disambiguation, where look-alike frames sit at very different progress. We then run a paired diagnosis: each PRM keeps the same input format in both runs, and in one run its instruction integrates the right context. Even PRMs that read the entire history get lost in estimating progress, yet with the right context the same five models cut their progress error by 77-82%. Embodied PRMs are thus not incapable of progress estimation, but lost without the right context. We therefore propose ProgressCompass, an autonomous agentic loop that reorients an existing PRM and uses current general-purpose VLMs to supply the context the PRM needs. Wrapped in the loop, the same frozen PRM cuts its progress error by 63% and raises its rank agreement by 76%. With such a compass, PRMs estimate progress far better on longer, more complex tasks.
Community
๐จ VLAs struggle on long, context-dependent tasks, so dense progress, a score at every step, matters for them.
๐ค Can progress reward models label dense progress for these context-dependent tasks?
๐งญ Our work shows why they fail: not because they are blind, but because they get lost in the context. Given the right context, the same progress reward models are reliable again.
๐งต So we propose ProgressCompass, a training-free agentic system: off-the-shelf progress models + VLMs, guiding progress estimation in long tasks.
๐ ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context
๐ https://andyzworks.github.io/progresscompass/
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- ARS: Agentic Reward System for Robot Learning (2026)
- Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents (2026)
- StateTape: Action-Conditioned Evidence Lifecycle Modeling for Long-Horizon Coding Agents (2026)
- One Rollout Is All You Get: Fully Test-Time Adaptation for GUI Agents (2026)
- How Strongly Should Task State Influence an LLM Agent? (2026)
- The Unreliable Progress Bar: Can LLM Agents Reliably Report Task Progress Throughout Execution? (2026)
- Aim Short to Reach Far: Your Frozen World Model Can Plan Better Than You Think (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.36684 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper