Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

OfficeAce Memory Plus (OAMP)

A local enhanced memory system for OfficeAce, inspired by TencentDB Agent Memory's 4-layer architecture. Zero external service dependencies — pure SQLite + Python, ready to use out of the box.


Why Do You Need It?

OfficeAce's native memory system has the following limitations:

  • No auto-capture: Information is lost after conversations end
  • No distillation: Key facts are buried in raw conversation logs
  • Weak retrieval: MEMORY.md is plain text with no structured search
  • Poor CJK search: FTS5's default tokenizer doesn't support Chinese word segmentation

OAMP solves all of the above.

Architecture

L0 Raw Conversations ──auto-capture──→ JSONL
L1 Atomic Facts ──distillation──→ Structured records (episodic/instruction/semantic)
L2 Scenario Experience ──clustering──→ Organized by scenario (project/workflow/preference/...)
L3 User Persona ──long-term learning──→ Core preferences and behavioral patterns
Layer Storage Retrieval Decay Half-life
L0 conversations table Time-range query —
L1 atoms table + FTS5 BM25 full-text + LIKE (CJK fallback) 90 days
L2 scenarios table + FTS5 BM25 full-text + LIKE (CJK fallback) 180 days
L3 persona_features table KV exact query 365 days

Key Features

  • CJK Search Fallback: FTS5's unicode61 tokenizer doesn't support Chinese segmentation; OAMP auto-detects CJK queries and falls back to LIKE pattern matching
  • Time-Decay Scoring: Recent memories score higher; configurable half-life per layer
  • Layered Token Budget: Retrieval budget allocated by L3→L2→L1 priority
  • Distillation Pipeline: Conversations → Atomic Facts → Scenario Experience → User Persona, refined layer by layer
  • LLM-Assisted Extraction: Distillation prefers LLM when available; auto-falls back to rule-based extraction

Quick Start

Install

# Clone the repository
git clone https://github.com/yourname/officeace-memory-plus.git
cd officeace-memory-plus

# No extra dependencies needed — just Python stdlib + SQLite
# Optionally:
pip install -e .

Initialize

from scripts.memory_plus import MemoryPlus

# Create/open the memory store
store = MemoryPlus('~/.officeclaw/memory-tdai/memory_plus.db')
store._ensure_db()  # Initialize schema on first use

Basic Usage

# Capture conversation (L0)
store.capture_conversation(
    session_id='session-001',
    role='user',
    content='Generate an ecology PPT, body font no smaller than 18pt'
)

# Add atomic facts (L1)
store.add_atoms_batch([
    {
        'type': 'instruction',      # episodic | instruction | semantic
        'category': 'PPT',
        'subject': 'PPT font rule',
        'content': 'PPT body font must be at least 18pt',
        'confidence': 0.95,
        'importance': 0.9,
        'tags': ['PPT', 'font', 'rule']
    }
])

# Search (auto CJK fallback)
results = store.search('font', max_results=10)
# results = {
#   'results': [{'layer': 'L1', 'type': 'atoms', 'items': [...]}],
#   'total_chars': 42,
#   'layers_searched': ['L3', 'L2', 'L1'],
#   'query': 'font'
# }

# Set user persona (L3)
store.set_persona('writing_style', 'concise')
store.set_persona('ppt_font_min', '18pt')

# Get persona
style = store.get_persona('writing_style')  # → 'concise'

Distillation Pipeline

from scripts.distill import DistillPipeline

pipeline = DistillPipeline(store)
# Extract L1 atomic facts from L0 conversations
pipeline.conversations_to_atoms(session_id='session-001')
# Cluster L1 atoms into L2 scenario experience
pipeline.atoms_to_scenarios()
# Distill L2 scenarios into L3 user persona
pipeline.scenarios_to_persona()

Session Hook (Auto-Capture)

from hooks.session_hook import SessionHook

hook = SessionHook(store)
# Call at the end of each session
hook.on_session_end(session_id='session-001', messages=[...])
# Automatically: 1. Capture conversation → 2. Trigger distillation

Project Structure

officeace-memory-plus/
├── SKILL.md              # OfficeAce skill documentation
├── README.md             # This file
├── LICENSE               # MIT
├── .gitignore
├── scripts/
│   ├── memory_plus.py    # Core engine: 4-layer memory CRUD + retrieval
│   ├── distill.py        # Distillation pipeline: L0→L1→L2→L3
│   ├── setup.py          # Setup script: init DB + migrate MEMORY.md
│   ├── sql/
│   │   └── schema.sql    # Database schema (with FTS5 indexes)
│   ├── test_quick.py     # Quick smoke test
│   ├── test_distill.py   # Distillation pipeline test
│   └── test_full.py      # Full integration test
└── hooks/
    └── session_hook.py   # OfficeAce session hook

Migrate Existing Data

setup.py can automatically migrate OfficeAce's existing MEMORY.md into the new system:

python scripts/setup.py --migrate-memory-md ~/.officeclaw/memory/MEMORY.md

Comparison with TencentDB Agent Memory

Feature TencentDB Agent Memory OAMP
4-layer architecture ✅ ✅
Distillation pipeline ✅ ✅ (LLM + rule dual-mode)
FTS5 full-text search ✅ ✅ + CJK LIKE fallback
Vector search ✅ (TCVDB) ❌ (pure local)
Team memory ✅ (Team Memory Hub) ❌ (single user)
External dependencies Docker 3 services Zero
Cross-device sync ✅ ❌
CJK search ❌ (unicode61 limitation) ✅ (auto LIKE fallback)

License

MIT License — see LICENSE


OfficeAce Memory Plus (OAMP) — 中文说明

参照 TencentDB Agent Memory 四层架构,为 OfficeAce 设计的本地增强记忆系统。 零外部服务依赖,纯 SQLite + Python,开箱即用。

为什么需要它?

OfficeAce 原生记忆系统存在以下问题:

  • 无自动捕获:对话结束后信息丢失
  • 无蒸馏提取:关键事实淹没在原始对话中
  • 检索能力弱:MEMORY.md 纯文本,无结构化检索
  • 中文搜索差:FTS5 默认 tokenizer 不支持中文分词

OAMP 解决了以上所有问题。

架构

L0 原始对话 ──自动捕获──→ JSONL
L1 原子事实 ──蒸馏提取──→ 结构化记录 (episodic/instruction/semantic)
L2 场景经验 ──聚类归并──→ 按场景组织 (project/workflow/preference/...)
L3 用户画像 ──长期学习──→ 核心偏好与行为模式
层级 存储 检索方式 时间衰减半衰期
L0 conversations 表 时间范围查询 —
L1 atoms 表 + FTS5 BM25 全文 + LIKE(CJK回退) 90天
L2 scenarios 表 + FTS5 BM25 全文 + LIKE(CJK回退) 180天
L3 persona_features 表 KV 精确查询 365天

关键特性

  • CJK 搜索回退:FTS5 unicode61 不支持中文分词,自动降级为 LIKE 模糊匹配
  • 时间衰减评分:越近期的记忆权重越高,可配置半衰期
  • 分层 Token 预算:按 L3→L2→L1 优先级分配检索预算
  • 蒸馏流水线:对话→原子事实→场景经验→用户画像,逐层提炼
  • LLM 辅助提取:蒸馏时优先使用 LLM,不可用时自动降级为规则提取

快速开始

安装

git clone https://github.com/yourname/officeace-memory-plus.git
cd officeace-memory-plus
# 无需额外依赖,Python 标准库 + SQLite 即可

初始化

from scripts.memory_plus import MemoryPlus

store = MemoryPlus('~/.officeclaw/memory-tdai/memory_plus.db')
store._ensure_db()  # 首次使用时初始化表结构

基本使用

# 捕获对话 (L0)
store.capture_conversation(
    session_id='session-001',
    role='user',
    content='帮我生成一份生态学PPT,正文字体不小于18磅'
)

# 添加原子事实 (L1)
store.add_atoms_batch([
    {
        'type': 'instruction',
        'category': 'PPT',
        'subject': 'PPT字体规则',
        'content': 'PPT正文字体不得小于18磅',
        'confidence': 0.95,
        'importance': 0.9,
        'tags': ['PPT', '字体', '规则']
    }
])

# 搜索(自动 CJK 回退)
results = store.search('字体', max_results=10)

# 设置用户画像 (L3)
store.set_persona('writing_style', '简洁干练')
style = store.get_persona('writing_style')  # → '简洁干练'

蒸馏流水线

from scripts.distill import DistillPipeline

pipeline = DistillPipeline(store)
pipeline.conversations_to_atoms(session_id='session-001')
pipeline.atoms_to_scenarios()
pipeline.scenarios_to_persona()

会话钩子(自动捕获)

from hooks.session_hook import SessionHook

hook = SessionHook(store)
hook.on_session_end(session_id='session-001', messages=[...])
# 自动:1. 捕获对话 → 2. 触发蒸馏

项目结构

officeace-memory-plus/
├── SKILL.md              # OfficeAce 技能说明
├── README.md             # 本文件
├── LICENSE               # MIT
├── .gitignore
├── scripts/
│   ├── memory_plus.py    # 核心引擎:四层记忆 CRUD + 检索
│   ├── distill.py        # 蒸馏流水线:L0→L1→L2→L3
│   ├── setup.py          # 安装脚本:初始化 DB + 迁移 MEMORY.md
│   ├── sql/
│   │   └── schema.sql    # 数据库 schema(含 FTS5 索引)
│   ├── test_quick.py     # 快速冒烟测试
│   ├── test_distill.py   # 蒸馏流水线测试
│   └── test_full.py      # 完整集成测试
└── hooks/
    └── session_hook.py   # OfficeAce 会话钩子

迁移现有数据

python scripts/setup.py --migrate-memory-md ~/.officeclaw/memory/MEMORY.md

与 TencentDB Agent Memory 的对比

特性 TencentDB Agent Memory OAMP
四层架构 ✅ ✅
蒸馏流水线 ✅ ✅ (LLM+规则双模式)
FTS5 全文检索 ✅ ✅ + CJK LIKE 回退
向量检索 ✅ (TCVDB) ❌ (纯本地)
团队记忆 ✅ (Team Memory Hub) ❌ (单用户)
外部服务依赖 Docker 三服务 零依赖
跨端同步 ✅ ❌
CJK 搜索 ❌ (unicode61 限制) ✅ (自动 LIKE 回退)

许可证

MIT License - 详见 LICENSE

About

OfficeAce Memory Plus - Four-layer memory architecture for OfficeAce agents

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages