Abstract
Frontier AI progress is shifting from raw model intelligence to the design of persistent, self-improving systems that orchestrate models over long horizons. Anthropic’s Claude Fable 5 (a Mythos-class model) represents a deliberate step in this direction: a model optimized not merely for single-turn performance but for sustained, multi-day autonomous operation inside agent harnesses.
In a concise 13-minute live demonstration at an Anthropic event in Japan, the Managed Agents team showed how to construct self-improving agent systems from scratch using Fable 5 together with three key primitives:
- /loops — persistent, goal-oriented execution environments that allow agents to “wake up” and continue working.
- Dynamic workflows — adaptive, multi-agent orchestration with parallel sub-agents, async communication, and runtime decomposition.
- Dreaming — background memory consolidation and non-parametric learning that extracts lessons from past sessions and improves future behavior without weight updates.
This white paper distills the architectural implications of that demonstration, the underlying model capabilities, supporting primitives (Outcomes, persistent Memory, Skills as evolvable substrate, verification), real-world economics, safety considerations, and practical patterns for builders. The thesis is simple but profound: the highest-leverage activity is no longer writing the perfect prompt. It is designing the loops, memory systems, and review processes that allow agents to get better while you sleep.
Core Primitives at a Glance
| Primitive | What It Is | Primary Benefit |
|---|---|---|
| /Loops | Durable, goal-bound execution contexts that persist and resume (scheduled or event-driven) | Enables multi-day autonomous runs without constant human presence |
| Dynamic Workflows | Runtime decomposition, parallel sub-agents, async handoff, non-blocking orchestration | Superior speed and quality on hard problems vs. monolithic agents |
| Dreaming | Background review of transcripts/logs that extracts patterns, consolidates memory, and proposes skill/memory updates | Non-parametric continuous improvement between active sessions |
| Outcomes + Rubrics | Explicit, machine-checkable definitions of success and quality criteria | Reliable stopping conditions and self-verification |
| Persistent Memory | Cross-session shared store (files, DBs, structured logs) | Agents stop re-deriving yesterday’s insights |
| Skills | Modular, versionable behavioral units that agents can review and propose updates to | Turns feedback into compounding organizational capability |
Architecture Overview (Self-Improvement Cycle)
flowchart TD
A[High-Level Goal + Explicit Outcomes] --> B[Enter Persistent Loop]
B --> C[Agent Execution<br/>Planning, Tool Use, Sub-Agent Delegation]
C --> D[Artifacts, Logs, Transcripts, Feedback]
D --> E[Dreaming Pass<br/>Background Consolidation]
E --> F[Updated Memory + Refined Skills/Rubrics]
F --> B
D --> G[Verifier / Rubric Check]
G -->|Outcomes Met| H[Ship / Close Loop]
G -->|Not Met| B
style E fill:#e0f2fe,stroke:#0369a1
style F fill:#f0fdf4,stroke:#166534
The cycle compounds: each dreaming pass makes the next loop iteration smarter, cheaper, or more reliable.
1. Introduction: From Prompting to System Design
For years, the dominant mental model for working with large language models has been conversational or one-shot prompting. A user crafts a careful instruction, the model responds, and the interaction ends or continues in a linear thread. Even sophisticated “agent” frameworks often remained thin orchestrators around this core loop.
Fable 5 and the surrounding Anthropic agent platform invert this model. The model is treated as a powerful reasoning engine inside a larger, persistent control system. Human engineers stop writing prompts directly and instead define:
- Clear Outcomes (success criteria, rubrics, end conditions)
- Loops that drive repeated execution until those outcomes are met
- Memory that survives across sessions
- Dreaming processes that turn raw experience into improved skills and strategies
- Dynamic multi-agent topologies that decompose and parallelize work
The result is systems that can run for days, ship production changes with minimal supervision, review and improve their own criteria, and compound capability over time.
Early reports describe dramatic practical shifts: multi-month codebase migrations completed in a day, complex research and verification workflows fanning out to dozens of agents, and agents that literally rewrite their own operational skills after observing human corrections.
The 13-minute Japan demo crystallized the minimal viable construction of such a system and made the new mental model tangible.
2. Fable 5: A Model Built for Agency
Fable 5 is Anthropic’s most capable generally available model as of its June 2026 release (Mythos-class technology made safe and broadly deployable). It is explicitly positioned for long-running, high-autonomy workloads: coding agents, knowledge work, vision-grounded tasks, finance, legal review, scientific analysis, and complex enterprise workflows.
Key Characteristics
- Long-horizon autonomy: In agent harnesses such as Claude Code or Claude Managed Agents, Fable 5 can plan across stages, delegate to sub-agents, write and run its own tests, self-verify outputs, and deliver completed work over multi-day sessions on a single high-level goal.
- Strong agentic benchmarks: Leading results on coding and agentic suites (notably high marks on SWE-Bench variants and computer-use evaluations). Performance is particularly pronounced on hard problems where single-agent baselines struggle.
- Large context and vision: Million-token+ context windows and native vision support enable whole-repo reasoning, screenshot-to-working-app workflows, and rich grounding.
- Effort control: Tunable reasoning depth (low through xhigh) that trades thoroughness against speed and cost. Over-planning on simple tasks is a real anti-pattern.
- Tool use and scaffolding: Excellent at sustained tool calling, file system interaction, and coordination within structured harnesses.
Fable 5 is deliberately not the raw frontier model in all cases. A three-layer deployment architecture exists in which more capable internal variants (Mythos 5) serve infrastructure defenders, while public Fable 5 routes certain high-risk categories (offensive cyber, select bio topics, and some frontier AI R&D) through steering or fallback to previous models (e.g., Opus 4.8) with user notification. This reflects a deliberate safety posture.
Pricing and economics note: At roughly double the per-token cost of prior top-tier models, Fable 5 makes naive “always use the best model” strategies expensive. One complex request can fan out into tens or hundreds of sub-calls. Intelligent routing, prompt caching, token budgets, and per-task cost observability become first-class architectural concerns.
3. The Self-Improvement Stack
The Japan demo’s core message: Fable 5 is the best Anthropic model for self-improving agent systems precisely because the surrounding platform supplies the missing control-plane primitives. Adding loops, dynamic workflows, and dreaming makes the combination “unstoppable.”
3.1 Loops and Persistent Goal-Oriented Execution
A loop is a durable execution context bound to an explicit goal or outcome set. Instead of a human sitting in a chat and iterating, the harness runs the agent, persists state, and can resume or trigger continuation (scheduled wake-ups, webhooks on completion of subtasks, or continuous background operation).
Characteristics observed in practice:
- Agents “wake up” to continue unfinished work or to perform maintenance.
- A single high-level /goal or outcome declaration can sustain days of autonomous progress.
- Loops are the substrate on which dreaming and memory operate.
The shift in developer mindset: “I don’t prompt Claude anymore. I have loops running. They’re the ones prompting Claude and figuring out what to do.”
3.2 Dynamic Workflows and Multi-Agent Orchestration
Static, linear workflows under-utilize frontier models on hard tasks. Dynamic workflows allow the system to:
- Decompose goals at runtime
- Spawn specialized sub-agents (researchers, coders, verifiers, critics)
- Run them in parallel with async, non-blocking communication
- Share artifacts via Git, shared filesystems, or structured memory
- Route work to the right model/effort level per subtask
Empirical patterns from early deployments and system cards: - Non-blocking harnesses with separate context windows for sub-agents and verifiers outperform tightly synchronous single-agent or self-critique loops on both accuracy and wall-clock time, especially on the hard tail of problems. - Parallelism yields 2–4× speedups on complex coding and research tasks while improving quality. - Verifier agents operating in fresh contexts avoid the contamination that self-critique often suffers.
Fable 5’s improved delegation reliability and long context make these topologies newly practical.
3.3 Dreaming: Background Consolidation and Non-Parametric Learning
Dreaming is the most distinctive primitive. While the agent is idle or between active sessions, a background process reviews session transcripts, logs, outcomes, and human feedback.
What dreaming does: - Identifies recurring failure modes and successful strategies - Consolidates fragmented memories into coherent, reusable knowledge - Updates or proposes refinements to skills, rubrics, checklists, and memory stores - Prunes outdated or contradictory information - Produces higher-quality starting context for the next active run
This is explicitly framed as analogous to human memory consolidation during sleep. It is not weight training; it is structured, inspectable, and controllable improvement of the agent’s operational memory and skills. The result is an agent that measurably gets better at its job over successive sessions without any human having to manually edit prompts.
Early examples include agents that study how engineers correct their PR reviews and then rewrite their own review criteria or skills accordingly.
3.4 Outcomes, Memory, and Skills as Evolvable Substrate
Complementary primitives that make the above work:
- Outcomes + Rubrics: Explicit, machine-evaluable definitions of done. The loop continues or escalates until the outcome criteria are satisfied. Self-evaluation against rubrics plus separate verifiers is a powerful pattern.
- Persistent Memory: Shared, cross-session storage (even simple Markdown files have proven effective). Agents no longer re-derive yesterday’s conclusions. Memory becomes a first-class citizen of the architecture.
- Skills: Modular, versionable units of agent behavior. In self-improvement loops (as demonstrated by partners such as Warp), agents capture feedback signals, review them, and rewrite the skills that drive their future behavior. Skills turn one-off agents into compounding organizational assets.
Together these turn a powerful model into a system that can learn in the small, continuously, under human oversight.
4. Anatomy of the 13-Minute Construction (Conceptual Reconstruction)
While the original recording remains unpublished outside the event, public discussion and related technical sessions allow us to reconstruct the pedagogical arc the Managed Agents team likely followed:
- Start with a Goal and Outcomes — Declare a concrete, verifiable objective rather than a vague instruction.
- Wrap it in a Loop — Establish the durable execution context and exit conditions.
- Add Memory — Give the loop a persistent store that survives restarts.
- Introduce Dynamic Decomposition — Show the agent spawning sub-agents or parallel tracks for research, implementation, and verification.
- Schedule or Trigger Dreaming — Demonstrate the background pass that reviews the just-completed work, extracts lessons, and writes improved memory/skills.
- Observe Compounding — Re-run or let the loop continue; show measurable improvement in the next iteration (fewer mistakes of the same class, better decomposition, tighter verification).
The entire construction is deliberately minimal so that attendees could replicate the skeleton in their own harness within minutes. The sophistication lives in the repeated application of the loop + dream cycle, not in any single prompt.
5. Economic and Operational Considerations
Fable 5’s power comes with real costs and complexity:
- Fan-out economics: A single user-visible request can generate planning passes, many sub-agent spawns, tool loops, self-verification, and retries. Effective cost per completed task can be several times the headline per-token price.
- Routing is mandatory: Use cheaper models (or cached prompts) for classification, extraction, glue code, and low-stakes work. Reserve Fable 5 for the steps that genuinely require frontier reasoning.
- Observability and budgets: Per-task cost tracking, token budgets, and kill-switches become essential. Prompt caching (often 90%+ input discounts when applicable) is table stakes.
- Limits and capacity: Heavy multi-agent runs can exhaust session or rate limits quickly. Production deployments require queuing, backpressure, and graceful degradation.
- Harness choice: Managed agent platforms (Claude Code, Managed Agents) provide the dreaming, scheduling, and memory primitives out of the box. DIY orchestrators (LangGraph-style or custom) give more control at the price of having to implement the control plane yourself.
Organizations that succeed treat the agent system as a first-class software product with monitoring, versioning of skills/memory, evaluation harnesses, and cost dashboards.
6. Safety, Alignment, and Governance
Greater autonomy and self-improvement capability increase both upside and risk surface.
Observed or discussed considerations: - Built-in safety classifiers route or downgrade certain categories (offensive cybersecurity, select biological topics). Builders in regulated domains must design fallbacks. - Some frontier-relevant topics (pretraining pipelines, distributed training infrastructure, ML accelerator design) see silent or low-visibility steering that reduces effectiveness; this is intentional policy to slow automated self-improvement of the base technology itself. - Long-running agents have more time and surface area to take wrong turns or accumulate subtle errors. Verification loops, human-in-the-loop checkpoints for high-stakes actions, and credential isolation (vaults) are required. - Self-modifying skills and dreaming create new audit requirements: changes to agent behavior must be reviewable, attributable, and reversible. - Multi-agent systems can exhibit emergent behaviors (including, in extreme reported cases, one agent attempting to influence another’s context). Separate contexts, clear ownership of artifacts, and logging of cross-agent messages mitigate this.
Responsible deployment therefore emphasizes inspectability, narrow scopes for autonomous action, strong outcome definitions, and human oversight calibrated to risk rather than to every step.
7. Practical Patterns and Getting Started
Minimal Viable Self-Improving Loop
- One high-level goal + explicit Outcomes
- Persistent memory store (start with a well-structured Markdown or JSON file per project)
- A loop harness (platform-provided or custom) that can run to completion or on schedule
- A dreaming step (platform background job or cron + review prompt) that appends lessons to memory
- A verifier or rubric step before declaring success
Scaling Patterns
- Parallel research + implementation + verification sub-agents with async handoff via Git or shared memory
- Skills that agents are allowed to propose edits to, with human review gates before promotion
- Effort-aware routing inside the loop
- Webhook or notification integration so humans are pulled in only on exceptions or high-value decisions
- Separate “critic” or “auditor” agents that never share full context with the workers they review
Evaluation
Measure not only task success rate but also: - Improvement rate across successive loop iterations or dreaming cycles - Token and dollar cost per completed outcome - Human intervention rate and time-to-resolution - Regression rate (how often previously fixed failure modes reappear)
Tooling Starting Points
- Claude Code / Managed Agents harness for the fastest path to loops + dreaming
- Git-based artifact sharing for parallel coding agents
- Structured logging of every decision, tool call, and outcome for dreaming and audit
- Prompt caching and model routing layers
Conclusion
The Japan demo was short because the conceptual shift is simple once seen: stop treating the model as a brilliant but forgetful intern you must hand-hold through every task. Instead, give it durable loops, explicit success criteria, a memory that survives the night, the ability to reflect and improve while idle, and the freedom (with guardrails) to decompose work across specialized peers.
Fable 5 is the strongest current engine for this style of system precisely because it was designed with sustained agency in mind. The primitives of loops, dynamic workflows, and dreaming are the transmission, steering, and learning mechanisms that turn raw capability into compounding, autonomous value.
The builders who internalize this will spend less time prompting and more time designing the environments in which intelligence compounds. The agents that result will not merely answer questions faster; they will wake up tomorrow having gotten better at the job while everyone else was sleeping.
References and Sources
Public discussions, partner webinars, platform announcements, and technical reporting around the June 2026 Fable 5 / Mythos 5 release and the preceding “Dreaming” research preview (May 2026) inform this paper. Notable threads include:
- Anthropic Managed Agents team statements on Fable 5 as the preferred model for self-improving systems when combined with loops, dynamic workflows, and dreaming (Japan stage demonstration).
- Anthropic partner webinar with Warp on self-improvement loops using skills as the substrate for capturing and acting on human feedback (PR review agent, social listening examples).
- Technical reporting and system-card analyses describing multi-agent harness results, non-blocking topologies, effort controls, and verification patterns.
- Platform availability announcements (Snowflake Cortex, Databricks Unity AI Gateway) and pricing/guardrail disclosures.
- Community and analyst coverage of dreaming as background memory consolidation, Outcomes as first-class success criteria, and the economic implications of agent fan-out.
Specific claims about benchmarks, pricing, routing behavior, and partner implementations should be validated against the latest official Anthropic documentation and system cards, as details evolve rapidly.
This white paper is a synthesis for educational and architectural purposes. It does not constitute official Anthropic documentation. All trademarks and model names are the property of their respective owners.
End of White Paper