以下为本文档的中文说明
该技能是自我改进代理的技能构建模块,专门用于为自我改进代理开发和注册新的能力。它定义技能接口、学习机制和评估标准,使代理能够自主发现和获取新技能。适用于扩展自我改进代理功能集合的开发和实验场景,是构建可进化AI系统的核心组件。该技能与自我改进代理技能相辅相成,前者提供代理本身的学习能力,后者则定义了如何为代理创建和注册新的能力模块。该技能是自我改进代理系统的技能扩展模块,定义了标准的技能接口和学习机制,使代理能够动态发现、学习和使用新技能。技能接口包括触发条件、执行流程、输入输出规范和评估标准,代理可以通过尝试和反馈来掌握新技能的使用方式。该技能与自我改进代理系统相辅相成:前者提供代理本身的学习能力框架,后者则定义了如何为代理创建和注册新的能力模块,两者共同构成了构建可进化AI系统的完整技术栈。该技能是自我改进代理系统的技能扩展框架,定义了标准化技能接口和动态学习机制,使代理能够不断发现、学习和运用新技能来扩展自身能力边界。技能接口规范定义了技能的完整生命周期:触发条件、执行流程、输入输出规格和评估标准。代理可以通过主动探索和用户反馈来逐步掌握新技能的使用方式。该技能与自我改进代理核心框架相辅相成,共同构成完整的可进化AI系统技术栈。
Self-Improving Agent
“An AI agent that learns from every interaction, accumulating patterns and insights to continuously improve its own capabilities.” — Based on 2025 lifelong learning research
Overview
This is auniversal self-improvement systemthat learns from ALL task experiences. It implements a complete feedback loop:
- Multi-Memory Architecture: Semantic (patterns/rules) + Episodic (experiences) + Working (session context)
- Self-Correction: Detects and fixes guidance errors
- Self-Validation: Periodically verifies skill accuracy
- Evolution Markers: Traceable changes with source attribution
- Confidence Tracking: Measures pattern reliability over time
- User Confirmation Gate: All skill file modifications require explicit user approval before applying
- Human-in-the-Loop: Collects feedback to validate improvements
Research-Based Design
| Research | Key Insight | Application |
|---|---|---|
| SimpleMem | Efficient lifelong memory | Pattern accumulation system |
| Multi-Memory Survey | Semantic + Episodic memory | World knowledge + experiences |
| Lifelong Learning | Continuous task stream learning | Learn from every task |
| Evo-Memory | Test-time lifelong learning | Real-time adaptation |
The Self-Improvement Loop
┌──────────────────────────────────────────────────────────────┐ │ UNIVERSAL SELF-IMPROVEMENT │ ├──────────────────────────────────────────────────────────────┤ │ │ │ Task Event → Extract Experience → Abstract Pattern → Update │ │ │ │ │ │ │ │ ▼ ▼ ▼ ▼ │ │ ┌────────────────────────────────────────────────────────┐ │ │ │ MULTI-MEMORY SYSTEM │ │ │ ├────────────────────────────────────────────────────────┤ │ │ │ Semantic Memory │ Episodic Memory │ Working Memory │ │ │ │ (Patterns/Rules) │ (Experiences) │ (Current) │ │ │ │ memory/self-improving/semantic/ │ memory/self-improving/episodic/ │ memory/self-improving/working/ │ │ │ └────────────────────────────────────────────────────────┘ │ │ │ │ ┌────────────────────────────────────────────────────────┐ │ │ │ FEEDBACK LOOP │ │ │ │ User Feedback → Confidence Update → Pattern Adapt │ │ │ └────────────────────────────────────────────────────────┘ │ │ │ └──────────────────────────────────────────────────────────────┘When This Activates
Automatic Triggers
| Event | Action |
|-------|--------|
| Any significant task completes | Extract patterns, propose skill updates (requires user confirmation) |
| An error or failure occurs | Capture error context, trigger self-correction (requires user confirmation before applying fixes) |
| Session ends | Consolidate working memory into long-term memory |
Manual Triggers
- User says “自我进化”, “self-improve”, “从经验中学习”
- User says “分析今天的经验”, “总结教训”, “总结经验”
- User asks to improve a specific skill or workflow
Memory Storage
Workspace Discovery
Before accessing any memory files, the agent MUST first determine the workspace root path:
- Check environment— Use the workspace path provided by the IDE/environment context
- Verify structure— Confirm the workspace root by checking for project markers (e.g.,
.git/,package.json,pom.xml, etc.) - All paths below are relative to the workspace root— e.g.,
{workspace}/memory/self-improving/
Relationship with Agent Memory
The Self-Improving Agent’s memory livesinsidethe Agent’smemory/directory as a dedicated subdirectory. This design ensures:
- No confusion: Agent’s own memory (
MEMORY.md,memory/YYYY-MM-DD.md) and Self-Improving Agent’s memory (memory/self-improving/) are clearly separated by directory structure - Discoverability: The Agent can browse
memory/and naturally find self-improving insights - Supplement, not replace: Self-Improving Agent canappendhigh-confidence patterns to Agent’s memory files (with user confirmation), enriching the Agent’s knowledge
{workspace}/ ├── MEMORY.md # Agent core memory (Self-Improving Agent can append) ├── memory/ │ ├── YYYY-MM-DD.md # Agent daily memory (Self-Improving Agent can append) │ └── self-improving/ # Self-Improving Agent dedicated memory space │ ├── semantic/ │ │ └── patterns.json # Abstract patterns and rules │ ├── episodic/ │ │ └── YYYY/ │ │ └── YYYY-MM-DD-{task}.json # Specific experiences │ ├── working/ │ │ ├── current_session.json # Active session data │ │ ├── last_error.json # Error context for self-correction │ │ └── session_end.json # Session end marker for consolidation │ └── index.json # Memory index and metricsMemory Interaction Rules
| Action | Target | Condition |
|---|---|---|
| Read | MEMORY.md | Always — to understand Agent’s accumulated knowledge |
| Read | memory/YYYY-MM-DD.md | Always — to understand today’s context |
| Append to | MEMORY.md | Only high-confidence patterns (>= 0.9), requires user confirmation |
| Append to | memory/YYYY-MM-DD.md | Session summary and key learnings, requires user confirmation |
| Full CRUD | memory/self-improving/* | Self-Improving Agent’s own memory space, free to manage |
Evolution Priority Matrix
Trigger evolution when new reusable knowledge appears:
| Trigger | Priority | Action |
|---|---|---|
| New workflow pattern discovered | High | Add to relevant skill guidance |
| Architecture/design tradeoff clarified | High | Add to decision patterns |
| Debugging fix or anti-pattern found | High | Add to troubleshooting patterns |
| Security or performance insight | High | Add to best practice patterns |
| Code pattern or idiom learned | Medium | Add to coding patterns |
| Test strategy improvement | Medium | Update testing approach |
| Tool usage optimization | Medium | Update tool usage patterns |
| Documentation structure insight | Low | Update documentation templates |
Multi-Memory Architecture
1. Semantic Memory (memory/self-improving/semantic/patterns.json)
Storesabstract patterns and rulesreusable across contexts:
{"patterns":{"pat-2025-01-11-001":{"id":"pat-2025-01-11-001","name":"Pattern Name","source":"user_feedback|implementation_review|retrospective","confidence":0.95,"applications":5,"created":"2025-01-11","last_applied":"2025-01-15","category":"coding_patterns|architecture|debugging|workflow|...","pattern":"One-line summary","problem":"What problem does this solve?","solution":"How to apply this pattern","quality_rules":["Rule 1","Rule 2"],"target_skills":["skill-name-1","skill-name-2"]}}}2. Episodic Memory (memory/self-improving/episodic/)
Storesspecific experiences and what happened:
{"id":"ep-2025-01-11-001","timestamp":"2025-01-11T10:30:00Z","skill":"debugger|coding-assistant|reviewer|...","task_type":"debugging|coding|review|design|...","situation":"What the user was trying to do","solution":"How the issue was resolved","outcome":"success|partial|failure","root_cause":"Underlying issue if applicable","lesson":"Key takeaway from this experience","related_pattern":"pattern_id if linked","user_feedback":{"rating":8,"comments":"User's feedback on the experience"}}3. Working Memory (memory/self-improving/working/)
Storescurrent session context— ephemeral data that gets consolidated at session end:
{"session_id":"session-2025-01-11-001","started":"2025-01-11T10:00:00Z","tasks_completed":[],"errors_encountered":[],"patterns_applied":[],"pending_extractions":[]}Self-Improvement Process
Phase 1: Experience Extraction
After any significant task completes, extract:
What happened:task_type:{what kind of task}task:{what was being done}outcome:{success|partial|failure}Key Insights:what_went_well:[what worked]what_went_wrong:[what didn't work]root_cause:{underlying issue if applicable}User Feedback:rating:{1-10 if provided}comments:{specific feedback}Phase 2: Pattern Abstraction
Convert experiences to reusable patterns. The goal is to go from concrete to abstract — patterns should be general enough to apply across different tasks but specific enough to be actionable.
| Concrete Experience | Abstract Pattern |
|---|---|
| “User forgot to save intermediate work” | “Always persist intermediate results to files” |
| “Code review missed SQL injection” | “Add security checklist to review process” |
| “Callback was empty, causing silent failure” | “Verify all callbacks have implementations” |
| “Ambiguous UI spec caused rework” | “UI specs need exact layout specifications” |
Abstraction Rules:
If experience_repeats 3+ times:pattern_level:criticalaction:Add to "Critical Mistakes" or "Anti-Patterns" sectionIf solution_was_effective:pattern_level:best_practiceaction:Add to "Best Practices" sectionIf user_rating >= 7:pattern_level:strengthaction:Reinforce this approach in relevant skillsIf user_rating <= 4:pattern_level:weaknessaction:Add to "What to Avoid" sectionPhase 3: Skill Updates
IMPORTANT: User Confirmation Required— Before writing any changes to skill files, you MUST:
- Present proposed changes— Show the user a clear summary of what will be modified:
- Which skill file(s) will be updated
- What content will be added, modified, or removed
- The rationale behind each change (source episode, pattern, confidence level)
- Wait for explicit approval— Do NOT proceed until the user confirms. Acceptable confirmations include explicit affirmative responses (e.g., “确认”, “好的”, “proceed”, “yes”).
- Apply changes only after approval— Once confirmed, apply the changes with evolution markers for traceability.
If the user rejects or requests modifications, adjust the proposed changes accordingly and re-present for confirmation.
Proposed Change Summary Format:
## Proposed Skill Update **Target**: `{skill-file-path}` **Action**: {Add new pattern | Correct existing guidance | Update checklist} **Source**: {episode_id or trigger} **Confidence**: {X.XX} ### Changes Preview {Show the exact content that will be added/modified, using diff-style or before/after format} ### Rationale {Why this change is recommended} --- Confirm this update? (yes/no/modify)Once confirmed, update skill files withevolution markersfor traceability:
<!-- Evolution: 2025-01-12 | source: ep-2025-01-12-001 | task: debugging --> ## Pattern Added (2025-01-12) **Pattern**: Always verify callbacks are not empty functions **Source**: Episode ep-2025-01-12-001 **Confidence**: 0.95 ### Updated Checklist - [ ] Verify all callbacks have implementations - [ ] Test callback execution pathsCorrection Markers(when fixing wrong guidance):
<!-- Correction: 2025-01-12 | was: "Use callback chain" | reason: caused stale state --> ## Corrected Guidance Use direct state monitoring instead of callback chains for reactive updates.Use the templates intemplates/for consistent formatting. Seereferences/appendix.mdfor the full template structures.
Phase 4: Memory Consolidation
Update semantic memory— add or update patterns in
memory/self-improving/semantic/patterns.jsonStore episodic memory— write episode to
memory/self-improving/episodic/YYYY/YYYY-MM-DD-{task}.jsonUpdate pattern confidence— increase confidence for patterns that were successfully applied, decrease for those that led to errors
Prune outdated patterns— lower confidence for patterns with no recent applications; archive patterns below 0.3 confidence
Supplement Agent memory— propose additions to Agent’s own memory files.User confirmation is REQUIREDbefore any write to
MEMORY.mdormemory/YYYY-MM-DD.md. Follow the same confirmation protocol as Phase 3:What to propose:
- High-confidence patterns (>= 0.9) as concise entries →
MEMORY.md - Today’s session summary and key learnings →
memory/YYYY-MM-DD.md
Confirmation format:
## Proposed Agent Memory Update ### → MEMORY.md (append) {Exact content to be appended, preview here} ### → memory/YYYY-MM-DD.md (append) {Exact content to be appended, preview here} **Source patterns**: {pattern IDs and confidence levels} --- Confirm this memory update? (yes/no/modify)After approval:
- Append confirmed content with
<!-- Source: self-improving-agent | date: YYYY-MM-DD -->markers for traceability - Do NOT overwrite existing content — always append at the end
- High-confidence patterns (>= 0.9) as concise entries →
Self-Correction
Triggered when:
- A command or operation returns an error
- Tests fail after following skill guidance
- User reports the guidance produced incorrect results
Process:
Detect Error
- Capture error context into
memory/self-improving/working/last_error.json - Identify which guidance was followed
- Capture error context into
Verify Root Cause
- Was the guidance incorrect?
- Was the guidance misinterpreted?
- Was the guidance incomplete?
Propose Correction
- Draft the corrected guidance with correction markers
- Present proposed changes to user for review (follow Phase 3 confirmation format)
- Wait for user confirmation before applying any changes
Apply Correction(after user approval)
- Update relevant skill/document with corrected guidance
- Add correction marker with reason
- Update related patterns in semantic memory
Validate Fix
- Test the corrected guidance if possible
- Ask user to verify the fix
Self-Validation
Periodically (or when triggered manually), verify that stored patterns and skill guidance are still accurate:
- Check that examples still work
- Verify checklists match current conventions
- Confirm external references are still valid
- Detect duplicated or conflicting guidance
Use the validation template intemplates/validation-template.mdfor structured reviews.
Human-in-the-Loop Feedback
After each self-i
mprovement cycle, present a summary to the user:
## Self-Improvement Summary I've learned from our session and updated: ### Patterns Extracted 1. **pattern_name**: Description (confidence: X.XX) ### Skills/Documents Updated - `skill-name`: What was updated ### Confidence Levels - New patterns: ~0.85 (needs more validation) - Reinforced patterns: ~0.95 (well-established) ### Your Feedback - Were these updates helpful? - Should I apply any pattern more broadly? - Any corrections needed?Integrate feedback into confidence scoring:
| Feedback | Action |
|---|---|
| Positive (rating >= 7) | Increase confidence, consider expanding to related skills |
| Neutral (rating 4-6) | Keep pattern, gather more data before expanding |
| Negative (rating <= 3) | Decrease confidence, revise or archive pattern |
Best Practices
DO
- Learn from EVERY significant task interaction
- Extract patterns at the right abstraction level — general enough to reuse, specific enough to be actionable
- Always present proposed changes to the user and wait for explicit confirmationbefore writing to skill files OR Agent memory (
MEMORY.md,memory/YYYY-MM-DD.md) - Update multiple related skills when a pattern applies broadly
- Track confidence and application counts for all patterns
- Ask for user feedback on improvements
- Use evolution/correction markers for full traceability
- Validate guidance before applying broadly
- Read
MEMORY.mdand today’smemory/YYYY-MM-DD.mdat the start of each self-improvement cycle for context
DON’T
- NEVER modify skill files or Agent memory files without user confirmation— this is a hard rule with no exceptions
- NEVER overwriteAgent memory content — always append at the end
- Over-generalize from a single experience — wait for 2-3 occurrences before creating a pattern
- Update skills without confidence tracking
- Ignore negative feedback — it’s the most valuable signal
- Make changes that break existing, working functionality
- Create contradictory patterns — resolve conflicts explicitly
- Apply untested patterns at high confidence
Quick Start
After any significant task completes, this agent:
- Analyzeswhat happened during the task
- Extractsreusable patterns and insights
- Proposesskill updates and presents them to the user for review
- Waitsfor explicit user confirmation before applying any skill modifications
- Updatesapproved changes to skill files with evolution markers
- Logsto memory (semantic + episodic) for future reference
- Reportssummary to user and collects feedback
References
For detailed memory structures, validation templates, metrics, and workflow diagrams, readreferences/appendix.md.
For pattern/correction/validation templates, see thetemplates/directory:
templates/pattern-template.md— Adding new patternstemplates/correction-template.md— Fixing incorrect guidancetemplates/validation-template.md— Validating skill accuracy
Research Papers
- SimpleMem: Efficient Lifelong Memory for LLM Agents
- A Survey on the Memory Mechanism of Large Language Model Agents
- Lifelong Learning of LLM based Agents
- Evo-Memory: DeepMind’s Benchmark