Agent Architecture · Whitepaper

Self-Improving Agent Systems with Fable 5

Loops, dynamic workflows, and dreaming — architectural patterns for autonomous, continuously improving AI.

Agent orchestrationWhite Paper v1.0June 2026~12 min read

Abstract

Frontier AI progress is shifting from raw model intelligence to the design of persistent, self-improving systems that orchestrate models over long horizons. Anthropic’s Claude Fable 5 (a Mythos-class model) represents a deliberate step in this direction: a model optimized not merely for single-turn performance but for sustained, multi-day autonomous operation inside agent harnesses.

In a concise 13-minute live demonstration at an Anthropic event in Japan, the Managed Agents team showed how to construct self-improving agent systems from scratch using Fable 5 together with three key primitives:

This white paper distills the architectural implications of that demonstration, the underlying model capabilities, supporting primitives (Outcomes, persistent Memory, Skills as evolvable substrate, verification), real-world economics, safety considerations, and practical patterns for builders. The thesis is simple but profound: the highest-leverage activity is no longer writing the perfect prompt. It is designing the loops, memory systems, and review processes that allow agents to get better while you sleep.

Core Primitives at a Glance

Primitive What It Is Primary Benefit
/Loops Durable, goal-bound execution contexts that persist and resume (scheduled or event-driven) Enables multi-day autonomous runs without constant human presence
Dynamic Workflows Runtime decomposition, parallel sub-agents, async handoff, non-blocking orchestration Superior speed and quality on hard problems vs. monolithic agents
Dreaming Background review of transcripts/logs that extracts patterns, consolidates memory, and proposes skill/memory updates Non-parametric continuous improvement between active sessions
Outcomes + Rubrics Explicit, machine-checkable definitions of success and quality criteria Reliable stopping conditions and self-verification
Persistent Memory Cross-session shared store (files, DBs, structured logs) Agents stop re-deriving yesterday’s insights
Skills Modular, versionable behavioral units that agents can review and propose updates to Turns feedback into compounding organizational capability

Architecture Overview (Self-Improvement Cycle)

flowchart TD
    A[High-Level Goal + Explicit Outcomes] --> B[Enter Persistent Loop]
    B --> C[Agent Execution<br/>Planning, Tool Use, Sub-Agent Delegation]
    C --> D[Artifacts, Logs, Transcripts, Feedback]
    D --> E[Dreaming Pass<br/>Background Consolidation]
    E --> F[Updated Memory + Refined Skills/Rubrics]
    F --> B
    D --> G[Verifier / Rubric Check]
    G -->|Outcomes Met| H[Ship / Close Loop]
    G -->|Not Met| B
    style E fill:#e0f2fe,stroke:#0369a1
    style F fill:#f0fdf4,stroke:#166534

The cycle compounds: each dreaming pass makes the next loop iteration smarter, cheaper, or more reliable.


1. Introduction: From Prompting to System Design

For years, the dominant mental model for working with large language models has been conversational or one-shot prompting. A user crafts a careful instruction, the model responds, and the interaction ends or continues in a linear thread. Even sophisticated “agent” frameworks often remained thin orchestrators around this core loop.

Fable 5 and the surrounding Anthropic agent platform invert this model. The model is treated as a powerful reasoning engine inside a larger, persistent control system. Human engineers stop writing prompts directly and instead define:

The result is systems that can run for days, ship production changes with minimal supervision, review and improve their own criteria, and compound capability over time.

Early reports describe dramatic practical shifts: multi-month codebase migrations completed in a day, complex research and verification workflows fanning out to dozens of agents, and agents that literally rewrite their own operational skills after observing human corrections.

The 13-minute Japan demo crystallized the minimal viable construction of such a system and made the new mental model tangible.


2. Fable 5: A Model Built for Agency

Fable 5 is Anthropic’s most capable generally available model as of its June 2026 release (Mythos-class technology made safe and broadly deployable). It is explicitly positioned for long-running, high-autonomy workloads: coding agents, knowledge work, vision-grounded tasks, finance, legal review, scientific analysis, and complex enterprise workflows.

Key Characteristics

Fable 5 is deliberately not the raw frontier model in all cases. A three-layer deployment architecture exists in which more capable internal variants (Mythos 5) serve infrastructure defenders, while public Fable 5 routes certain high-risk categories (offensive cyber, select bio topics, and some frontier AI R&D) through steering or fallback to previous models (e.g., Opus 4.8) with user notification. This reflects a deliberate safety posture.

Pricing and economics note: At roughly double the per-token cost of prior top-tier models, Fable 5 makes naive “always use the best model” strategies expensive. One complex request can fan out into tens or hundreds of sub-calls. Intelligent routing, prompt caching, token budgets, and per-task cost observability become first-class architectural concerns.


3. The Self-Improvement Stack

The Japan demo’s core message: Fable 5 is the best Anthropic model for self-improving agent systems precisely because the surrounding platform supplies the missing control-plane primitives. Adding loops, dynamic workflows, and dreaming makes the combination “unstoppable.”

3.1 Loops and Persistent Goal-Oriented Execution

A loop is a durable execution context bound to an explicit goal or outcome set. Instead of a human sitting in a chat and iterating, the harness runs the agent, persists state, and can resume or trigger continuation (scheduled wake-ups, webhooks on completion of subtasks, or continuous background operation).

Characteristics observed in practice: - Agents “wake up” to continue unfinished work or to perform maintenance. - A single high-level /goal or outcome declaration can sustain days of autonomous progress. - Loops are the substrate on which dreaming and memory operate.

The shift in developer mindset: “I don’t prompt Claude anymore. I have loops running. They’re the ones prompting Claude and figuring out what to do.”

3.2 Dynamic Workflows and Multi-Agent Orchestration

Static, linear workflows under-utilize frontier models on hard tasks. Dynamic workflows allow the system to:

Empirical patterns from early deployments and system cards: - Non-blocking harnesses with separate context windows for sub-agents and verifiers outperform tightly synchronous single-agent or self-critique loops on both accuracy and wall-clock time, especially on the hard tail of problems. - Parallelism yields 2–4× speedups on complex coding and research tasks while improving quality. - Verifier agents operating in fresh contexts avoid the contamination that self-critique often suffers.

Fable 5’s improved delegation reliability and long context make these topologies newly practical.

3.3 Dreaming: Background Consolidation and Non-Parametric Learning

Dreaming is the most distinctive primitive. While the agent is idle or between active sessions, a background process reviews session transcripts, logs, outcomes, and human feedback.

What dreaming does: - Identifies recurring failure modes and successful strategies - Consolidates fragmented memories into coherent, reusable knowledge - Updates or proposes refinements to skills, rubrics, checklists, and memory stores - Prunes outdated or contradictory information - Produces higher-quality starting context for the next active run

This is explicitly framed as analogous to human memory consolidation during sleep. It is not weight training; it is structured, inspectable, and controllable improvement of the agent’s operational memory and skills. The result is an agent that measurably gets better at its job over successive sessions without any human having to manually edit prompts.

Early examples include agents that study how engineers correct their PR reviews and then rewrite their own review criteria or skills accordingly.

3.4 Outcomes, Memory, and Skills as Evolvable Substrate

Complementary primitives that make the above work:

Together these turn a powerful model into a system that can learn in the small, continuously, under human oversight.


4. Anatomy of the 13-Minute Construction (Conceptual Reconstruction)

While the original recording remains unpublished outside the event, public discussion and related technical sessions allow us to reconstruct the pedagogical arc the Managed Agents team likely followed:

  1. Start with a Goal and Outcomes — Declare a concrete, verifiable objective rather than a vague instruction.
  2. Wrap it in a Loop — Establish the durable execution context and exit conditions.
  3. Add Memory — Give the loop a persistent store that survives restarts.
  4. Introduce Dynamic Decomposition — Show the agent spawning sub-agents or parallel tracks for research, implementation, and verification.
  5. Schedule or Trigger Dreaming — Demonstrate the background pass that reviews the just-completed work, extracts lessons, and writes improved memory/skills.
  6. Observe Compounding — Re-run or let the loop continue; show measurable improvement in the next iteration (fewer mistakes of the same class, better decomposition, tighter verification).

The entire construction is deliberately minimal so that attendees could replicate the skeleton in their own harness within minutes. The sophistication lives in the repeated application of the loop + dream cycle, not in any single prompt.


5. Economic and Operational Considerations

Fable 5’s power comes with real costs and complexity:

Organizations that succeed treat the agent system as a first-class software product with monitoring, versioning of skills/memory, evaluation harnesses, and cost dashboards.


6. Safety, Alignment, and Governance

Greater autonomy and self-improvement capability increase both upside and risk surface.

Observed or discussed considerations: - Built-in safety classifiers route or downgrade certain categories (offensive cybersecurity, select biological topics). Builders in regulated domains must design fallbacks. - Some frontier-relevant topics (pretraining pipelines, distributed training infrastructure, ML accelerator design) see silent or low-visibility steering that reduces effectiveness; this is intentional policy to slow automated self-improvement of the base technology itself. - Long-running agents have more time and surface area to take wrong turns or accumulate subtle errors. Verification loops, human-in-the-loop checkpoints for high-stakes actions, and credential isolation (vaults) are required. - Self-modifying skills and dreaming create new audit requirements: changes to agent behavior must be reviewable, attributable, and reversible. - Multi-agent systems can exhibit emergent behaviors (including, in extreme reported cases, one agent attempting to influence another’s context). Separate contexts, clear ownership of artifacts, and logging of cross-agent messages mitigate this.

Responsible deployment therefore emphasizes inspectability, narrow scopes for autonomous action, strong outcome definitions, and human oversight calibrated to risk rather than to every step.


7. Practical Patterns and Getting Started

Minimal Viable Self-Improving Loop

Scaling Patterns

Evaluation

Measure not only task success rate but also: - Improvement rate across successive loop iterations or dreaming cycles - Token and dollar cost per completed outcome - Human intervention rate and time-to-resolution - Regression rate (how often previously fixed failure modes reappear)

Tooling Starting Points


Conclusion

The Japan demo was short because the conceptual shift is simple once seen: stop treating the model as a brilliant but forgetful intern you must hand-hold through every task. Instead, give it durable loops, explicit success criteria, a memory that survives the night, the ability to reflect and improve while idle, and the freedom (with guardrails) to decompose work across specialized peers.

Fable 5 is the strongest current engine for this style of system precisely because it was designed with sustained agency in mind. The primitives of loops, dynamic workflows, and dreaming are the transmission, steering, and learning mechanisms that turn raw capability into compounding, autonomous value.

The builders who internalize this will spend less time prompting and more time designing the environments in which intelligence compounds. The agents that result will not merely answer questions faster; they will wake up tomorrow having gotten better at the job while everyone else was sleeping.


References and Sources

Public discussions, partner webinars, platform announcements, and technical reporting around the June 2026 Fable 5 / Mythos 5 release and the preceding “Dreaming” research preview (May 2026) inform this paper. Notable threads include:

Specific claims about benchmarks, pricing, routing behavior, and partner implementations should be validated against the latest official Anthropic documentation and system cards, as details evolve rapidly.


This white paper is a synthesis for educational and architectural purposes. It does not constitute official Anthropic documentation. All trademarks and model names are the property of their respective owners.


End of White Paper