computer-using-agents

Tag

Cards List
#computer-using-agents

@omarsar0: // Automating SKILL.md Generation // Increasingly, mining sessions is one of the best ways to improve your agents. Open…

X AI KOLs Following · 2026-06-19 Cached

This paper from MIT and Harvard explores automating SKILL.md generation by mining GUI interaction trajectories, finding that clusters are readable but do not improve policy performance across domains.

0 favorites 0 likes
#computer-using-agents

@dair_ai: Outstanding paper on computer-using agents. (bookmark it) Computer-using agents drive real software through the screen,…

X AI KOLs Following · 2026-06-17 Cached

PreAct compiles successful agent runs into small state-machine programs, enabling 8.5-13x faster replay on repeated tasks without per-step language model calls, with runtime screen checks to ensure correctness.

0 favorites 0 likes
#computer-using-agents

PreAct: Computer-Using Agents that Get Faster on Repeated Tasks

arXiv cs.AI · 2026-06-17 Cached

PreAct compiles successful task runs of computer-using agents into small state-machine programs, allowing fast replay (8.5–13× faster) on repeated tasks by skipping per-step language model calls, while verifying screen states at each step and falling back to the agent when mismatches occur.

0 favorites 0 likes
#computer-using-agents

SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows?

arXiv cs.AI · 2026-05-18 Cached

SaaS-Bench is a new benchmark built on 23 deployable SaaS systems across six professional domains, containing 106 long-horizon tasks for evaluating computer-using agents. Experiments show that even the strongest models complete fewer than 4% of tasks end-to-end, highlighting significant limitations in current agent capabilities.

0 favorites 0 likes
← Back to home

Submit Feedback