Running on Zero MCP 5 Valen Visual Decisions 👁 5 Visual question answering with candidate probabilities
Paint-Anything: Unified Any-Color Control for Image Generation and Editing Paper • 2609.20816 • Published 12 days ago • 57
OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation Paper • 2609.22069 • Published 11 days ago • 37
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 11 days ago • 136
Grounded Skill Synthesis from Code at Scale for Agentic Intelligence Paper • 2609.05571 • Published 25 days ago • 119
RNGBench Collection [EMNLP2026] Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games • 2 items • Updated 14 days ago • 2
WorldReward: Reward Modeling for Camera-Conditioned World Models Paper • 2609.03952 • Published 26 days ago • 27
Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models Paper • 2608.25518 • Published Aug 26 • 60
JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution Paper • 2608.25593 • Published Aug 26 • 69
One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows Paper • 2608.19741 • Published Aug 20 • 12