
Raven: The Harness of Harnesses for Composable Agentic Intelligence
Get every episode summarized
Each time Daily Paper Cast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
About this episode
“Today's paper is from the Hugging Face Daily Paper List of September 30th, 2026, with 254 upvotes. The title is Raven, the harness of harnesses for composable, agentic intelligence.”From the transcript
🤗 Upvotes: 254 | cs.AI, cs.CL, cs.GT, cs.MA, cs.NE
Authors:
EverMind AI
Title:
Raven: The Harness of Harnesses for Composable Agentic Intelligence
Arxiv:
http://arxiv.org/abs/2609.33439v1
Abstract:
As large language models advance, AI agents are moving beyond isolated, domain-specific tasks toward long-horizon, cross-domain workflows. This transition exposes two challenges: increasing harness complexity makes manual design difficult to scale, while tighter coupling to specific domains limits the generality of a single harness. The central question thus shifts from how to engineer a stronger harness for one domain to how to autonomously construct specialized harnesses, improve them through experience, and orchestrate them across domains. We introduce Raven, \emph{The Harness of Harnesses}, an open-source multi-agent ecosystem that automatically constructs and evolves modular harnesses for specific models and domains, treating each executable model--harness pair as a composable unit of intelligence. To support an \emph{All-Domain Collaboration Network}, its Host Agent decomposes goals, matches subtasks to specialized agents, coordinates execution dependencies, and integrates results, while a host archive and EverOS preserve experience across tasks and Skill Forge makes that experience available as reusable procedures. Our theory establishes sufficient conditions for such composition to expand reliable task coverage beyond that of the available individual agents under a shared resource budget. On complex and long-horizon tasks, Raven significantly outperforms the state-of-the-art agent systems, pushing the frontier of composable agentic intelligence.
Get every episode summarized
Each time Daily Paper Cast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
344 searchable segments. Every word is indexed and playable.
Full transcript
Daily Paper Cast — Raven: The Harness of Harnesses for Composable Agentic Intelligence. Machine-transcribed; use the interactive transcript above to jump the player to any line.
Welcome to Daily Papercast. Today's paper is from the Hugging Face Daily Paper List of September 30th, 2026, with 254 upvotes. The title is Raven, the harness of harnesses for composable, agentic intelligence. It's authored by Chuanru Hu and Dijon Shui, with Chuanru Hu as the corresponding author, representing Ever Mind AI. Let's dive into the introduction. Advances in large language models, including instruction following and code generation, have provided a foundation for agents that interpret user goals and act through software. Indeed, language models can even learn to invoke external tools. Agent systems organize these capabilities into multi-step interactions with external environments. This includes reasoning combined with actions and observations, which allows agents to gather information and adapt their decisions based on feedback.
Specialized execution interfaces shape what agents can accomplish. For example, in automated software engineering, these developments have shown how agents with different expertise and execution mechanisms can work toward a shared goal. Complex goals often require several specializations within one workflow. For instance, developing and operating a first-person shooter game involves requirements research, gameplay programming, visual design, integration testing, and sustained operation. Each of these stages involves different tools and criteria, and the outputs of one stage support the next. So an agent's capability depends not only on its model, but also on its harness, the tools, context management, skills, execution policies, and recovery mechanisms around the model. These mechanisms dictate how an agent applies its models' capabilities within a domain. Exactly. And this leads to two main challenges. First, as we support more tools and execution conditions, the number of interacting design
choices increases, making manual harness development challenging to scale across various models and domains. And second, specialized harnesses often encode assumptions about their tools, working context, and expected outputs. Mectoring capabilities requires careful matching of sub-tasks to suitable executors and ensuring compatible handoffs. Multi-agent failures often stem from interagent misalignment and inadequate task verification. Mechanisms like multi-agent conversation and role-based workflows help to organize collaboration. However, the benefit of such compositions depends on task structure, local capabilities, and the cost of coordination. The central problem then becomes constructing and improving specialized harnesses while ensuring their composability across different domains. To address these challenges, the paper introduces Raven, the harness of harnesses, and open-source multi-agent ecosystem. This system treats each executable model, harness pair, as a unit of composition.
Raven includes the native specialists, Raven research, Raven code, Raven design, and Raven encall, as well as independently developed agents like Claude code, codex, and Hermes agent. Execution adapters connect these agents to a shared orchestration interface while preserving their tools and internal execution policies. Within this system, the host agent organizes collaboration by matching sub-tasks to registered capabilities and mapping these dependencies as a directed, acyclic graph or DAG. Each node represents an agent, and the edges denote dependencies between their outputs and subsequent tasks. Moreover, Raven combines harness adaptation with the reuse of experience across tasks, whereas prior methods adapt agents by retaining insights from past executions or distilling interactions into reusable skills, Raven modular harnesses expose execution policies that external revolvers can modify within defined boundaries. Building on previous work, harness bank, the evolution process diagnoses failures, proposes
candidate harnesses, and evaluates their behavior with a fixed task model. Persistent memory complements these policy changes with a host archive retaining user context and an optional ever OS back end providing semantic access to past experiences. Furthermore, Skill4 allows for procedural reuse by retrieving task relevant procedures from local skills, memory derived skills, and skill hub, a corpus and retrieval design based on prior work called SkillCorpus. In addition to the system design, the theoretical analysis in the paper characterizes when composition can expand the capabilities of the available agents. For a specified agent pool and task family, it formalizes capability as reliable task coverage under a common resource budget. To dive into specifics, the theory establishes sufficient conditions under which complementary local capabilities, compatible handoffs, and bounded planning and execution errors allow the composed system to solve tasks that the individual agents cannot reliably solve
alone under the same budget. Essentially, Raven aims to improve task performance on complex and long horizon tasks, significantly outperforming state-of-the-art agent systems. This pushes the frontier of composable, agentic intelligence. Indeed, and this concludes the introduction section of the paper. Alright, let's dive into the methods detailed in the paper. The method section covers several key components of the Raven ecosystem, including its architecture for multi-agent collaboration, harness self-evolution, memory systems, and skill retrieval processes. Starting with multi-agent collaboration, Raven's design includes a host agent, and an agent that decomposes tasks, matches sub-tasks to specialized agents, coordinates execution dependencies, and integrates results. Exactly. The host agent uses an orchestration guidance tool which helps to compose a graph for complex tasks. It starts by loading the orchestration guidance on demand when a request needs multiple
agents. This is done to avoid incurring an input cost for each turn that doesn't require orchestration. Once the host determines that a request requires orchestration, it loads the full set of rules necessary to construct a graph. These rules specify node wiring, placeholder syntax, concurrency limits, and whether single delegation is preferable. If the guide is missing, the system provides feedback to direct the host to relevant rules. After constructing the graph by working backward from the requested deliverable, each node is assigned a unique identifier and an executor able to produce the output. The host agent then adds dependencies indicating which nodes output is required as input for another node. Admission validation follows. The runtime admits the graph after five groups of checks. Format, graph structure, agent capability, agent status, and environment. These checks ensure the nodes are well formed, dependencies are resolvable, paths are confined
within permitted routes, agents are registered and enabled, and that the process does not violate dispatch budget or recursion rules. Validation stops at the first error and returns a message naming the offending node or field, often suggesting a repair. Once the graph is admitted, each node proceeds through the runtime, moving through states like pending, running, exception, completed, failed, skipped, or canceled. As run wants, all dependencies are settled, and output or raised errors return to an independent model called judge. The judge must validate the completion of nodes and either accept them as accomplished or mark them as exceptions for host intervention. Clarification requests are sometimes made by workers seeking information omitted in the prompt. The host can attempt to answer these from its current conversation and memory or forward them to the user if necessary. Under execution, a run report is generated summarizing intermediate outcomes, including completed nodes, failures, cancellations, skips, node output files, and terminal outputs.
Reports enable the host to synthesize the final deliverable from the nodes outputs. Nodes are key here. Artifacts and memory flow separately. Artifact completion releases dependencies while memory recording proceeds asynchronously. Local-based node records allow successors to read up stream outputs in line or by reference paths without duplicating artifacts and prompts. Cross-graph reuse is supported with node identifiers and shared context blocks, allowing later runs to reference earlier outputs without executing the node again. This reduces redundant subtasks and limits host context references. Group memory is also notable. Shared group memory converts observations into assessments that inform future orchestration and worker invocations. Round summaries, agent assessments, and user feedback provide a collective record that informs planning and retrieval during subsequent tasks. This memory management strategy ensures persistence, context transfer, and execution synchronization,
supporting effective agent collaboration over long periods and multiple tasks. Now, let's move on to the second component, Harness Self-Evolution. Raven's method for harness adaptation builds on Harness Bank. It exposes the model, Harness pairs through modular strategy interfaces, evaluates their performance, diagnoses failures, and evolves Harness's by proposing candidates and screening them under fixed protocols. Failures are monitored, and new candidate Harness's are proposed to address diagnosed issues. These candidates undergo a strict screening process, evaluating validity, activation, and paired gain, retaining only the strongest performers for further evaluation. These candidates are categorized by the types of modifications they include prompt, knowledge, runtime, and configuration edits. Selected candidates are then stored in a Harness Gene Bank for reuse and further improvement. A task agent executes tasks, and a Volvo agent proposes modifications, and a fixed evaluator
measures outcomes. Modular Harness design allows for changes without affecting the underlying model, focusing on improving how the agent applies its model capabilities within domains. Evaluation uses a fixed protocol specifying model versions, task sets, and environment conditions. Successful interventions are validated for their causal contribution to performance improvements. Complex tasks guide the selection process, but do not provide family-wise guarantees. Memory systems help agents preserve interaction histories and procedural knowledge. Critic trace formation groups interactions into segments, which undergo semantic consolidation into MemScenes, identifying trends and persistent information. Reconstructive retrieval, recalling episodes and facts, uses BM25 and Dense Search combined through reciprocal rank fusion. SkillForge retrieves procedures, updates skills from execution cases, sorts reusable insights, and appends validated skills to the memory store.
The retrieval ensures agents have access to prior experiences to inform subsequent tasks. Each agent session records useful methods and fixes to maintain continuity and adaptability across tasks. That concludes the method section. Let's move on to the Experiment and Results section to understand how Raven performs. Yes, the paper evaluates Raven's effectiveness in organizing specialist work, improving execution with a fixed model and reusing procedural knowledge. Let's start with the Multi-Agent Orchestration. The evaluation uses the Multi-Agent Orchestration Benchmark, or MAOB, which comprises 140 requests modeled on occupational tasks. Each task is paired with a reference directed A Cyclic Graph, or DAG, a four specialist subagents. The specialists cover research, coding, content creation, and on-call tasks. The Benchmark measures agreement between the hosts' proposed graph and a reviewed reference
graph. This allows us to evaluate specialist selection and dependency planning independently of downstream execution quality. Exactly. The metrics used include NodeF1, EdgeF1, Partial Order Accuracy, and Exact Match Rate. Raven's orchestration performance is compared to two baseline harnesses, Claude Code and Hermes Agent, under two backbones, QN38, two seven billion, and deep seek V4 Flash 731. Raven consistently outperforms both baselines across all metrics and backbones. For instance, with QN38, two seven B, Raven achieves a NodeF1 of 0.923 compared to 0.776 for the strongest baseline. The EdgeF1 also shows improvements, with Raven reaching 0.950 under the same backbone. Raven's Exact Match Rate is 0.711, with QN38, two seven B, offering a significant 10.4%
point improvement over the strongest baseline. With deep seek V4 Flash 731, Raven achieves an exact match rate of 0.867, a 10.5% point improvement over the strongest alternative. These metrics show Raven's strong capability in specialist selection and accurate dependency planning, significantly outperforming alternatives. The results highlight the robust orchestration capabilities of Raven. Next is Harness Self-Evolution. Raven builds on Harness Bank and evaluates the diagnosis, search, and screening procedures on tasks withheld from evolution. Harness and benchmarks are used, including Terminal Bench 2, EVO Agent Bench covering 5 domains and AppWorld. The evolved Harness consistently improves held out pass at 1 across all benchmarks. For example, Harness Evolution raises the score from 36.1% to 45.4% on Terminal Bench 2, and from 41.3% to 56.7% on AppWorld.
The improvements demonstrate the effectiveness of Harness Adaptation and the transfer ability of gains across different tasks. It's impressive how the evolved Harness achieves such notable gains, confirming the utility of Raven's approach to modular Harness Evolution. Moving on to Raven Research, which answers deep research questions using evidence from the live web. Evaluated on deep research mixed, it combines questions from browse comp, frames, humanities last exam, and expench deep search. On deep seek V4 Flash, Raven Research achieves a pooled accuracy of 76.5%, outperforming all compared systems. With Q336, 35B, A3B, and Q335, 397B, A17B, it also leads with 56.3%, and 59.3% accuracy respectively. In addition to accuracy, Raven Research balances inference costs well.
For example, on deep seek V4 Flash, its meme cost per question is $0.024, achieving higher accuracy at competitive costs. Then we have Raven Code, which focuses on repository-aware software engineering. It resolves tasks from SWE Bench Pro and Verified, Repository Development from WorkBuddy Bench, and whole repository migration from SWE Refactor. Raven Code achieves high scores across benchmarks. On SWE Bench Pro, it resolves 15 more tasks than Claude Code with Q32827B. On SWE Refactor, Raven Code outperforms leaderboard entries by 9.5 points with deep seek V4 Flash. Its combined performance shows strong results on complex repository tasks and targeted bug fixing. Data Agent Bench further confirms Raven Code's strengths, achieving the highest pass at one of 0.8762 among leaderboard entries for analytical database queries.
Finally, Raven Design handles visual design and inspection. Present Bench overall scores show Raven Design outperforming baselines in both tested backbums. Its scores 80.2 with Claude Opus 5 and 72.9 with GPT5.6 Luna. Raven Design achieves consistent gains across visual design benchmarks, dashboard, SVG, and GDPVAL tasks. It led in all settings, demonstrating its effectiveness in visual artifact generation and professional deliverables. And we see similar effectiveness in managing long running jobs with Raven On Call. For AI4AI, Raven On Call achieves better pre-training quality and lower costs under Claude Opus 5 and deep seek V4.1 Flash. It also solves more tasks with lower resource use in the AI4S scientific benchmark. Across all experiments, Raven consistently outperforms comparable systems, demonstrating the utility of its orchestration, adaptation, research, coding, design, and sustained task
management capabilities. This concludes the experiment section of the paper review. Now, let's delve into the related work section to see how Raven builds on and compares with existing methods in the field. The paper highlights four main areas of related work that Raven integrates in advances, multi-agent collaboration and orchestration, agent harness self-evolution, long-term memory for agents, and skill acquisition, retrieval and evolution. Starting with multi-agent collaboration, different frameworks have been developed to organize agents within a shared framework. Autogen and MetagPT are notable examples that specialize through configurable message exchanges and role prompts. Yes, Autogen allows developers to compose agents through messages and conversation patterns, while MetagPT assigns roles for collaborative software development. These frameworks support specialization through communication configurations, but often
lack an explicit typed graph that can be validated before execution. Other systems like MacNet represent interaction with directed acyclic graphs, and magenta one maintains a task ledger, assigning agents dynamically based on changes during execution. Mass router, meanwhile, learns to route requests among large language models in multi-agent setups. This demonstrates that coordination gains depend heavily on the interaction between task structure and system architecture. Kim at all found that increasing agent count alone doesn't predict collaboration performance. Efficient coordination is essential, especially in complex task environments. Ravens' approach to multi-agent orchestration includes a host agent that delegates tasks through explicit executable graphs. Ensuring dependencies are well-formed and validated before execution. This separates planning from execution, which allows for better coordination across diverse agents. Moving on to agent harness self-evolution.
Previous methods like GEPA and ADAS focus on evolving prompts and agent designs. GEPA evolves prompts through reflective feedback, while ADAS automates the design of agentic systems. These approaches optimize specific parts of agent execution without providing a comprehensive mechanism for evolving the harness itself. Open-ended evolution systems like the Darwin-Gotal Machine and Growing Harness address broader aspects by evolving the agent's own code or its harness around a frozen model. They work within the constraints of an offline optimizer or reusable scaffolding respectively. Harness Bank introduces a semantic gene bank approach, allowing modular harness evolution with gated verification. This supports extensive evaluation and comparison across various harness modifications categorized by prompt, knowledge, runtime, and configuration edits. Raven builds on these methods by integrating harness evolution directly into its multi-agent system. The modular design allows changes to the harness without altering the underlying model.
The fixed evaluator and evolve agent ensure that only the most effective harnesses are retained. Next, we consider long-term memory for agents. Memory systems range from context management and reflection models like generative agents and MemGPT to scalable systems like Memory OS and Mirix, which organize memory across different dimensions and substrates. These systems illustrate varied approaches to managing memory. Generative agents simulate human behavior and reflect on episodes, whereas Mirix includes procedural memories alongside other types in a cohesive system. However, the difference between task artifacts and memory records often remains blurred in practice. Raven leverages Evaros to structure memory as discrete episodes, facts and reusable skills, enabling efficient recall and update processes. By separating artifact completion from memory recording, Raven ensures clear boundaries between what an agent needs immediately and what it retains for future use.
Lastly, skill acquisition, retrieval, and evolution are critical for effective agent performance. Memory systems like Voyager and XPL learn executable code and natural language insights, respectively, while others like skill corpus curate extensive catalogs of public skills. Skill corpus, for example, aggregates, deduplicates, and evaluates public skills, providing a curated database for agent retrieval. Skill router enhances this retrieval with specialized models ensuring appropriate skill routing at skill. Rave and skill weaver represent the latest steps in procedural skill evolution, updating knowledge dynamically based on execution experiences. They, however, often focus on within task gains rather than a broader multi-agent system perspective. Raven skill forage extends this by integrating local and ever-OS-derived skills with skill hub, enabling the agent to retrieve and utilize reusable procedures. Modular updates from execution cases ensure skills remain relevant and improve over time.
In summary, Raven stands on the shoulders of previous work while forging new paths in multi-agent collaboration, harness evolution, agent memory, and skill utilization, making it a comprehensive solution for complex, long horizon tasks. This concludes the related work section of the paper. As we near the end of today's episode, let's summarize the key contributions and takeaways from the paper. The paper presents Raven, the harness of harnesses, an innovative multi-agent ecosystem designed to tackle the complexity of harness development and the composition of specialized capabilities across domains. A standout feature of Raven is its advanced multi-agent orchestration. The system's host agent efficiently decomposes tasks, matches sub-tasks to specialized agents, and ensures execution dependencies are well coordinated. Raven's orchestration capabilities significantly outperform existing systems, as demonstrated
by its superior exact match rate on the MAOB benchmark. Another core contribution is harness self-evolution. Building on the foundation of harness bank, Raven's modular approach allows for continuous improvement and effective adaptation of agent harnesses. The rigorous evaluation and screening processes ensure only the most effective harness modifications are implemented. Raven also excels in task performance across various domains. Its specialized agents, Raven research, Raven code, Raven design, and Raven on call demonstrate remarkable efficiency and accuracy in their respective tasks. From deep research to repository aware software engineering, visual design, and sustained operation tasks. Moreover, Raven's memory systems and skill forage facilitate procedural reuse and skill improvement. By integrating local skills, EVROS derived skills, and skill hub, Raven ensures that agents leverage past experiences to inform and enhance future tasks.
In essence, Raven stands as a comprehensive solution, pushing the frontiers of composable agentech intelligence. It combines effective task orchestration, modular harness evolution, and advanced memory and skill management to tackle complex long-term tasks successfully. Thank you for joining us on this episode of Daily Papercast. We hope you found our discussion on Raven insightful and thought-provoking. Be sure to tune in again for more deep dives into the latest advancements in AI research and technology. Until next time, stay curious and keep exploring. Goodbye!
More episodes
More from Daily Paper Cast

MaLiang-Harness: A Programmable Path to Image and Video Generation
Daily Paper Cast

PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation
Daily Paper Cast

Omni-IO Skills: Harnessing Your Agent Omni-Native
Daily Paper Cast

What Makes World Action Models Generalize? An Empirical Study of Test-Time Futur...
Daily Paper Cast