Tag
Omi Desktop is an AI companion for Mac that summarizes meetings, analyzes screen content, generates tasks, and provides proactive nudges to enhance productivity, positioning itself as a more natural alternative to AI assistants like ChatGPT or Claude.
The article discusses the importance of integrating QA into RL task generation through an iterative process to improve AI pipelines and model evaluation, emphasizing the value of intuition in eval design.
Proposes a labeled-data-free meta-learning method that generates tasks by assigning soft labels from pre-trained models to unlabeled data, avoiding computationally expensive model inversion. Achieves up to 104x speedup and 8.4-36.4% accuracy improvements over state-of-the-art DFML methods.
Introduces PROPEL, a solver-amortized framework that trains a lightweight activation probe to predict solver pass rates, enabling efficient training of task generators for RL without costly solver rollouts. The method improves generation at the learnable frontier across math, code, and software-engineering tasks.
OpenComputer presents a framework for creating verifiable software environments for computer-use agents, integrating state verifiers, self-improving verification layers, task synthesis, and evaluation systems across 33 desktop applications. Experiments show its verifiers align better with human judgment than LLM-as-judge, and frontier agents struggle with end-to-end completion.