Agents can improve in two very different ways. One is to change the model itself through training or fine-tuning. The other is to change the context around the model: instructions, examples, memories, tools, and reusable procedures. For enterprise agents, the second path is often faster to iterate, easier to evaluate, and easier to reverse.
The practical problem is what happens after a good run. A reviewer corrects a classification, adds a missing verification step, changes the structure of an output, or teaches the agent how a local process actually works. Unless that correction becomes a reusable artifact, the next run starts from the same baseline and the organization keeps paying for the same learning.
In Facthory, we are addressing this with self-improving skills. A skill is a versioned package of instructions, examples, references, and output expectations that an agent can load for a matching task. Reviewed work can be harvested into a candidate skill version, evaluated, reviewed, and promoted for broader use.
The important part is the release model. The component that proposes a better skill is not the component that decides what is active in production. Runtime permissions are also resolved independently from skill content. This lets us improve task behavior without turning the optimization loop into a second authorization system.
This post covers the architecture we are implementing: how reviewed work becomes a skill candidate, how skills are scoped and versioned, how candidate versions are evaluated and promoted, and how runtime authorization remains separate from the learning loop.
The lifecycle starts after work has already been reviewed. We do not treat every successful trace as training data, and we do not update a shared skill as a hidden side effect of a conversation. Capture is explicit.
A selected trace is first minimized and redacted inside the tenant boundary. The optimizer receives only the material required to improve the procedure. Its output is a new draft version. That draft is evaluated on held-out work, then sent to a reviewer with the diff and evaluation evidence. An authorized publisher can move it to canary and, later, active. A failed candidate is rejected without changing the active version.

Diagram key: each outlined group is an ownership boundary: Reviewer / Publisher, Tenant Runtime, Optimizer and Evaluation, or Skill Registry. Solid arrows are normal handoffs, thick arrows are publication or promotion, and dotted arrows are rejection or rollback. Visual colors are applied by the website renderer rather than embedded in Mermaid.
There are two deliberate boundaries in this flow. The optimizer cannot write an active version, and passing an evaluation does not make a candidate active. Evaluation produces evidence for a release decision; it is not the release mechanism itself.
Most agent systems begin with instructions in a system prompt. That is a reasonable starting point, but a prompt string becomes difficult to operate once it carries business-critical procedure.
Production procedures need identity and lifecycle. We need to know which version ran, where it applies, who owns it, what changed, which evidence supported the change, and how to restore a previous version. We also need to distinguish a reusable procedure from the permissions required to execute it.
The open Agent Skills specification provides a useful packaging model: a SKILL.md file with metadata and instructions, plus optional scripts, references, and assets. Its progressive-disclosure model also keeps context costs manageable by loading the full skill only when it is relevant. Facthory follows the same general idea of a portable, inspectable skill artifact.
We add product-level lifecycle around that artifact: scope, version history, provenance, evaluation results, review decisions, canary state, active-version pointers, and rollback. We also make one implementation choice explicit: skill metadata can describe compatibility or requested capabilities, but it is not the source of runtime authorization. The Agent Skills specification includes an experimental allowed-tools field; in our runtime, any such metadata is still subordinate to identity, tenant policy, and the capability decision made for the current task.
| Artifact | Primary responsibility | Lifecycle |
|---|---|---|
| Prompt | Frames the current model interaction | Rendered for a task |
| Skill | Reusable procedure, examples, references, output contract | Versioned, scoped, evaluated, published |
| Policy | Determines allowed capabilities, limits, approvals, and scope | Governed independently from skill content |
| Tool | Executes a typed action against a system | Invoked only after runtime authorization |
The architecture is easier to reason about when the skill lifecycle and the agent runtime are separate systems that meet at one narrow interface: the runtime can resolve an active, scoped skill version. It cannot treat a draft as active, and the optimizer cannot bypass the registry to inject instructions into a production task.
Within the runtime, the behavior path and the authority path remain distinct. The skill resolver and context builder determine what procedure the model sees. The policy engine determines what the task is allowed to do. A requested tool call must satisfy both paths before it executes.

Diagram key: the left group is the skill lifecycle, the middle group is the agent runtime, and the two external nodes are the model plane and enterprise systems. The main path is optimization → evaluation → registry → skill resolution → model → runtime gate → tool gateway. Publisher and policy enter only at the control points they own.
This separation also gives us a clean failure model. A bad candidate can fail in evaluation without affecting runtime. A malformed active skill can be rolled back without changing policy. A model can request a tool that the skill expects, but the runtime gate can still deny the call because the current principal or tenant does not have that capability.
A skill library becomes much easier to operate once the data model is explicit. The important object is not only the skill name. It is the relationship between a skill, its immutable versions, the scopes where it is visible, the evidence attached to each candidate, and the publication pointer that selects the version a runtime can resolve.
We intentionally keep evaluation and review records separate from the skill body. That lets us re-run a candidate against a new evaluation set without modifying the candidate itself, and it preserves the evidence used for each publication decision.

Diagram key: crow's-foot notation shows multiplicity between skills, versions, scopes, evaluation runs, review decisions, and publication.
PUBLICATION → SKILL_VERSIONidentifies the immutable version selected for publication. Entity styling is left to the website renderer.
The PUBLICATION → SKILL_VERSION reference is the basis of rollback. Rollback does not rewrite the previous skill. It changes the publication pointer to a version that already exists and already has its own provenance and evaluation history.
A score is not a state transition. Candidate versions move through a small, explicit lifecycle, and the transition rules are part of the product contract.
The optimizer can create Draft. Evaluation evidence can make the draft eligible for In review. Only an authorized reviewer can move a candidate to Canary, and only an authorized publisher can promote a canary to Active. Rejected and archived versions remain addressable for audit and comparison.

Diagram key: gray states are stored but not live, yellow states are under review or limited canary exposure, green is the active production version, and red is rejected. Arrow labels describe the condition that permits a state transition; the label background is transparent so the transition remains visually continuous.
Rollback is intentionally represented as a publication operation rather than a special optimizer action. If version v17 is active and a canary exposes a regression, the publisher can restore v16 by changing the active pointer. The version graph does not need to be reconstructed from prompt history.
Publishing a skill only makes it eligible to shape a task. It does not pre-authorize every action the skill might recommend.
At the start of a task, the runtime resolves the active skill versions that match the tenant, scope, and mode. Separately, the policy engine computes the capability envelope for the current principal and task. The model receives the resolved instructions, but every requested tool action is checked against the capability envelope before execution.

Diagram key: solid sequence arrows are requests or actions; dashed arrows are responses, decisions, or returned evidence. The yellow
altframe is the authorization branch. Skill resolution occurs before model reasoning, but authorization is checked again at the point of action.
This distinction matters when skills are shared. An organization-level troubleshooting skill can describe a diagnostic procedure that uses several systems. A user who can read only one of those systems should still be able to use the skill, but the runtime must narrow execution to that user's actual grants or stop at an approval boundary.
The optimizer needs a metric, but a production release needs more than a single score. We evaluate candidate versions against held-out work that is separate from the examples used to generate the candidate. The evaluation record can include task quality, safety or policy checks, regression cases, latency, and cost.
A candidate should not gain publication rights because it improves an average. It may improve common cases while failing a required field, overfitting one site, increasing cost materially, or violating a control that appears only in a small safety slice. Those failures need to be visible before the candidate receives broader exposure.
This is why we think about agent improvement on two axes: how much the system can adapt and how much production control surrounds that adaptation. The quadrant below is conceptual; the positions are architectural, not benchmark measurements.

Diagram key: the x-axis represents the system's ability to update reusable behavior; the y-axis represents versioning, evaluation, scope, publication control, and rollback around those updates. Point locations are a conceptual architecture comparison, not empirical benchmark scores. Facthory is highlighted in primary blue.
The goal is not to maximize either axis independently. A completely static procedure can be well governed but expensive to maintain. An automatically rewritten playbook can adapt quickly but make it difficult to explain which instruction is active and why. We want the adaptation loop to remain fast while publication remains an explicit, attributable operation.
A self-improving personal agent can optimize for one user. An enterprise has a different problem: the same procedure may legitimately vary by site, role, business unit, classification level, or local process.
That changes how we handle promotion. A skill captured from one user's work remains private until it is deliberately shared. A site-level candidate does not become an organization default because it performed well locally. Promotion to a wider scope is a separate review decision.
Conflicts are also first-class. Two sites can disagree about a filing rule or an escalation threshold and both can be correct in their local context. Automatically merging those instructions would remove information that reviewers actually need to see. We therefore treat conflicting candidates as review items rather than attempting to converge them into a single global instruction set.
This fits the broader durable and multiplayer agent model in Facthory: long-running work preserves plans, evidence, checkpoints, decisions, and shared context so people and agents can continue from a common operational record. Skills add a reusable procedure layer to that same collaboration model; they do not replace the underlying scope and authority model.
The idea of improving an agent without retraining the underlying model is supported by several lines of research.
Voyager demonstrated an embodied agent that grows a reusable library of verified skills and retrieves them for later tasks. Reflexion showed that linguistic feedback stored in episodic memory can improve subsequent attempts without gradient updates. DSPy made prompt and pipeline optimization explicit by compiling language-model programs against metrics rather than relying only on manual prompt editing.
More recently, Agentic Context Engineering (ACE) treats context as an evolving playbook updated through generation, reflection, and curation. ACE specifically addresses failure modes such as brevity bias and context collapse, and reports improvements of 10.6% on agent benchmarks and 8.6% on finance tasks while reducing adaptation latency and rollout cost.
The open Agent Skills specification complements this research from a packaging perspective by defining a portable directory format for instructions, references, scripts, and assets, with progressive disclosure at runtime.

Diagram key: the timeline is chronological. Colors distinguish periods only; they do not encode quality or relative importance.
The research establishes that context and reusable procedures are powerful adaptation surfaces. The product problem begins after that: making those updates attributable, tenant-scoped, evaluable, publishable, and reversible when more than one person depends on them.
We do not want the safety properties of the skill loop to live only in documentation. The core constraints should map to components and tests that can fail during development.
Mermaid's requirement diagram is useful here because it separates the requirement from the component that satisfies or verifies it. The diagram below is intentionally small; these are the invariants that define the boundary of the system.

Diagram key: requirement nodes define the invariants; element nodes are the implementation components or control boundaries responsible for them. A
satisfiesrelationship identifies which component enforces each requirement. Styling is left to the website renderer.
These requirements are intentionally independent of model choice. A stronger model can improve candidate generation, but it should not change who may publish, how tenant isolation works, or how the runtime decides whether a tool call is allowed.
Several mechanisms would make the system appear more autonomous while making it harder to operate.
Direct optimizer writes to the active version. The optimizer writes drafts. Publication is a separate operation with separate rights.
Tool grants inside skill content. A skill can describe a procedure that expects a tool, but the runtime capability envelope remains authoritative.
Cross-tenant replay pools. Reviewed operational traces stay tenant-scoped. We do not need a global memory pool to improve a tenant-specific procedure.
Automatic conflict merge. When two scoped skills disagree, reviewers should see the disagreement and decide whether it represents an error or a legitimate local variant.
Weight updates as the default improvement path. For operational procedure, a versioned skill is easier to inspect, evaluate, scope, and roll back than a tenant-specific weight update.
Hidden capture. A successful task does not silently become a shared organizational procedure. Capture and promotion are explicit events.
This keeps the improvement surface narrow enough to reason about. The optimizer can search aggressively within a candidate. The rest of the system can still treat that candidate like any other production artifact: version it, test it, review it, release it gradually, observe it, and roll it back.
The immediate benefit is repeatability. A procedure that works does not have to remain in one person's chat history or in an unversioned prompt fragment. It can become a scoped artifact with an owner and evidence.
The longer-term benefit is a more useful form of organizational memory. Reviewed work can improve the procedures available to future agents and people, while preserving the differences between a personal technique, a site standard, and an organization-wide policy.
We are implementing the skill lifecycle as part of Facthory's durable and multiplayer agent architecture. The product surface follows directly from the model above: a skill library, version history, diff, evaluation evidence, scoped sharing, review, canary state, publication, archive, and rollback.
The objective is not to build an agent that rewrites itself unchecked. It is to make operational procedures measurable, versioned, and improvable while keeping publication and runtime authority explicit.