EverMind has released Raven, an Apache 2.0-licensed system positioned as a 'harness of harnesses': it autonomously constructs and evolves agent harnesses for specific models and domains, while a Host Agent decomposes tasks into a graph and coordinates specialized agents to produce a single final result. The GitHub repository has already gathered around 5,000 stars and 1,700 commits, release 0.2.3 is dated September 27, 2026, and the project was published alongside a technical report on arXiv and the THRESHOLD case, where Raven ran 42 rounds of an autonomous cycle in approximately 4 days. Key benchmark figures currently belong to EverMind's own runs, so the system still requires independent verification.

image
image

What happened

EverMind published the Raven project in the EverMind-AI/Raven repository: an open Apache 2.0 license, around 5,000 stars and 1,700 commits as of the first public trace, and current release 0.2.3 dated September 27, 2026. Inside are EverOS memory, SkillForge with a corpus of 114,190 skills in SkillCorpus, four built-in agents (Raven-Research, Raven-Code, Raven-Design, Raven-Oncall), and presets for connecting 13 external CLI agents, including Claude Code, Codex, OpenCode, GitHub Copilot, Qwen Code, Hermes Agent, and OpenClaw. The central demonstration artifact is the THRESHOLD case: from a human brief, Raven ran 42 rounds of the 'planning — development — verification' cycle in approximately 4 days and independently built a website, poster, PPTX presentation, and a playable Godot 4 FPS with an arena and boss fight; release artifacts from September 25, 2026 are in the repository and are manually verified. The repository tables and technical report claim 76.5% for Raven-Research on the DeepResearch Mixed aggregate versus 68.9% for DeepSeek-Harness and 67.2% for MiroFlow, Pass@1 of 0.8762 for Raven-Code with Opus-5 on DataAgentBench, Raven-Design leading on PresentBench, and results on SWE-bench Verified and SWE-bench Pro.

Context

A harness is the layer around a language model: planners, tools, skills, and memory that turn a model into a working agent. Traditionally, this layer is built manually, and it is often this layer, not the model itself, that limits the final quality of an agent. The technical report arXiv:2609.33439 dated September 27, 2026, establishes a different approach: harness design shifts from manual engineering to autonomous construction and evolution. In Raven, this is implemented through four decoupled agent modules (Memory, Planning, Capability, Action): the Evolver and Curator components diagnose failures, test changes against benchmarks, and accept only verified edits. As a result, each 'model + harness' pair becomes a composable unit that the Host Agent assembles into an All-Domain Collaboration Network: the goal is decomposed into a task graph (DAG), specialized agents are selected for specific subtasks, dependencies are coordinated, and results are integrated into a single final artifact. Architecturally, such a network of specialized agents opposes monolithic systems like DeepSeek-Harness and MiroFlow.

Why this matters for the industry

For the industry, this is an attempt to shift value from the layer of manual framework development to the layer of orchestrating ready-made agents: instead of building their own planner and memory, teams connect CLI agents they already use through ready-made presets and get orchestration as a configuration task. For startups and product teams, this drastically reduces the barrier to entry into multi-agent development — they need to verify configuration, artifact acceptance, and experiment boundaries, not a framework from scratch. Open benchmark tables together with release artifacts from a 42-round autonomous run give the industry a measurable reference point for where the boundary of long-horizon autonomy without a human cycle currently lies. If harness evolution through Evolver and Curator shows a reproducible delta on external benchmarks and independent replications appear, Raven could become an open base for meta-orchestration and harness comparison over the same model, and product value in this market would shift to brief design, eval sets, and result acceptance tools.

Why this matters for users

The cost of verification is zero: Raven is free and installs with a single command curl -fsSL https://raven.evermind.ai/install.sh | bash on macOS, Linux, or WSL2; an alternative path is git clone and docker compose -f docker/docker-compose.yml up, after which the WebUI opens on localhost:18793. A realistic scenario today is a closed pilot on non-sensitive tasks: research automation, internal tooling, auto-prototyping from a brief, or a package of presentation materials, always with manual result verification and readiness to change configurations between releases. For those already using CLI agents, it is easiest to evaluate Raven as a host agent over familiar tools connected through presets, and to test the 'harness + model' combination on their own tasks. The level of autonomy is visually assessed: the attached release artifacts in the repository are checked by eye, without retelling.

What is still unknown / limitations

All key benchmark figures (including 76.5% for Raven-Research on DeepResearch Mixed and 0.8762 Pass@1 for Raven-Code on DataAgentBench) are EverMind's own runs on benchmarks it selected; there are no independent replications yet, and the run methodology is not fully disclosed. The self-development module Curator is marked as experimental in the repository, the project is in pre-alpha stage, and interfaces and configurations change rapidly between releases. The THRESHOLD case is a single illustrative example, not a systematic proof of stable autonomy on long horizons; how much depth of harness evolution is reproducible on external tasks and benchmarks not belonging to EverMind is currently unknown.

Sources

Author

Look at AI, editorial team