CogEvol: Towards Efficient and Reliable Learning Environment Generation
Summary
CogEvol is a family of models that efficiently generate structured learning artifacts like slides and interactive HTML pages in a single pass using supervised fine-tuning and reinforcement learning with vision-language rewards, reducing cost and improving reliability.
View Cached Full Text
Cached at: 09/01/26, 11:52 AM
Paper page - CogEvol: Towards Efficient and Reliable Learning Environment Generation
Source: https://huggingface.co/papers/2608.30968 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
CogEvol is a family of models that generate structured learning artifacts in a single pass using supervised fine-tuning and reinforcement learning with vision-language rewards, achieving high quality with far fewer parameters and lower cost.
We present CogEvol, a family of models trained specifically for Learning Environment Generation: turning a course brief into a finished learning artifact (structured-JSON slides or self-contained interactive HTML pages) in a single pass. Across 220k production requests, CogEvol completes a slide in a median of 17 seconds and an interactive page in 59, replacing minutes-long multi-turn agent scaffolding. Reliability is enforced rather than hoped for: a production-grounded data pipeline turns real failures into 53,687 verifiedSFTsamples, and a hybrid rule-plus-VLM rewarddrivesGRPO-based RL, hardened after we caught and fixed a reward-hacking episode that produced visually convincing but unplayable games. CogEvol-27B scores 83.7 on slide quality and 63.7 on a 500-case interactive-HTML benchmark with 26.9x fewer parameters than flagship coding models, and, in collaboration with the OpenMAIC team, serves their live production traffic. CogEvol-4B is released openly under the Apache 2.0 license at https://github.com/CogEvol/CogEvol-4B; external flagships are measured on the same suites under the identical harness. Scaffold editing cuts interactive-page generation cost by a further ~76%, and the full stack runs on domestic Ascend accelerators at application-level parity with A800 GPUs, lowering the unit cost of AI-native education at scale.
View arXiv pageView PDFGitHub1Add to collection
Get this paper in your agent:
hf papers read 2608\.30968
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper3
#### CogEvol/CogEvol-4B Text Generation• 5B• Updatedabout 9 hours ago • 1 • 14
#### CogEvol/CogEvol-4B-Q4_K_M-GGUF Text Generation• 4B• Updatedabout 9 hours ago • 2
#### prithivMLmods/CogEvol-4B-GGUF Text Generation• 4B• Updated42 minutes ago • 1
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.30968 in a dataset README.md to link it from this page.
Spaces citing this paper1
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation
GenEvolve is a self-evolving image generation framework that uses tool-orchestrated trajectories and visual experience distillation to iteratively improve generative capabilities, achieving state-of-the-art performance.
EvoOptiGraph: Weakness-Driven Coevolution via Graph-Based Structural Generation for Optimization Modeling
EvoOptiGraph is a framework for automating optimization modeling from natural language using graph-based evolutionary generation to create diverse training data and co-evolve the model with weakness-driven reinforcement learning, achieving state-of-the-art results on multiple benchmarks.
EAGER: Enhancing Generative Event Extraction via Reinforcement Learning with Verifiable Rewards
EAGER is a reinforcement learning framework that improves generative event extraction through fine-grained verifiable rewards and schema-contrastive advantage estimation, outperforming prompting, fine-tuning, and prior RL methods on seven benchmark datasets.
Self-Evolving Deep Research via Joint Generation and Evaluation
Researchers from HKUST, ByteDance, and UCL propose SCORE, a co-evolutionary training framework that jointly trains an LLM as both a deep research report generator and an evaluator, using a meta-harness to dynamically adjust evaluation difficulty and prevent reward saturation. Experiments show consistent improvement in open-ended research report quality.
Edu-Theater: A Data-Efficient Agent Framework for Scalable Learner Behavior Simulation through Staging Roll-Call
Edu-Theater is a data-efficient agent framework that uses LLM-powered generative agents to simulate learner behavior in educational settings. It employs a cohort-aware roll-call paradigm to infer learner states with fewer data and computational resources, achieving higher simulation accuracy.