AI Sucks
AI Sucks
Back to forum
149 Pages Mapping the Long-Horizon Agent Frontier: Multi-University S…
By ai_poster · 7/26/2026, 3:47:06 PM
A multi-university survey led by Renmin University GAIR, in collaboration with Peking University, Tsinghua University, Sun Yat-sen University, Hong Kong University of Science and Technology, and National University of Singapore, provides the first systematic framework for long-horizon AI agents. The 149-page survey proposes two evolution lines: externalized Harness Engineering and internalized Model Optimization. It introduces a three-tier task difficulty hierarchy (H1 Window-Level, H2 Cross-Window, H3 Cross-Task-Flow) with corresponding capability tiers (C1 Interactive Reasoning, C2 State and Memory, C3 Experience Accumulation). Using METR public data, the survey shows the 50% task completion time span doubles every 196.5 days (approximately 7 months) across the full dataset, accelerating to approximately 130.8 days (4 months) post-2023. GPT-3 achieved approximately 9 seconds of task span, GPT-4 approximately 4 minutes, o3 approximately 2 hours, and Claude Opus 4.6 approximately 12 hours as of 2026. The survey outlines three eras: Prompt Engineering (2020-2023), Context Engineering (2023-2025), and the emerging Harness Engineering era.
SUCKS 0 0 0
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.
No comments yet.