Prime Agent: A Self-Improving RLM Harness

2026-08-24Artificial Intelligence

Artificial IntelligenceComputation and LanguageSoftware Engineering
AI summary

The authors created Prime Agent, a tool that helps language models handle very long tasks by letting them use external memory and agents working together. It keeps track of everything the model does and lets smaller subagents communicate and work in parallel. This setup helps the model perform much better on difficult tasks, like coding and game simulations, by managing processes without failing easily. The authors show that Prime Agent improves performance significantly and works well across various challenging scenarios.

language modellong-horizon agencyRecursive Language ModelIPython REPLsubagentsagent communicationcontext processingresource managementautonomous agentscoding workflows
Authors
Seth Karten, Alex L. Zhang, Kevin Thomas, Sebastian Müller, Elie Bakouch, Daniel Auras, Mika Senghaas, Fares Obeid, Konstantin Dunas, Johannes Hagemann, Sami Jaghouar
Abstract
Language models are sequential processors, but long-horizon agency requires external information and computation beyond model weights and active context. Prime Agent is an open-source harness for long-horizon evaluation and coding-agent workflows. A persistent IPython REPL follows the Recursive Language Model abstraction for programmatic context processing and test-time compute, while Continual Harness preserves histories, memories, skills, prompts, and subagent specifications across trajectories. Recursive subagents coordinate through direct agent-to-agent communication, and the Agents View lets humans inspect and manage daemon-backed sessions. Prime Agent standardizes execution, recovery, verification, and resource accounting while leaving strategy construction to the model. This low-friction, expressive membrane prevents harness failures from becoming model failures and pushes measurement toward the model's true maximal underlying capability. Prime Agent raises ARC-AGI-3 RHAE Best@1 from 30% to 95.5% and matches or exceeds native and popular harnesses across long-context coding, GPU-kernel generation, emulator construction, and autonomous nanoGPT speedruns. On Factorio, we find refinement allows for continuous technology progression and dedicated subagents enable parallelized work. Code is available at https://github.com/PrimeIntellect-ai/prime-agent.