From 7c6a3dbe4a070a42b4231829e7d268dee3dad38e Mon Sep 17 00:00:00 2001 From: Joey Yakimowich-Payne Date: Fri, 13 Mar 2026 11:41:41 -0600 Subject: [PATCH] docs: add architecture and reference documentation Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode) Co-authored-by: Sisyphus --- ARCHITECTURE.md | 615 ++++++++++++++++++++++++++++++++++++++++++++++++ README.md | 7 + REFERENCES.md | 220 +++++++++++++++++ ROADMAP.md | 487 ++++++++++++++++++++++++++++++++++++++ SCHEMA.md | 443 ++++++++++++++++++++++++++++++++++ 5 files changed, 1772 insertions(+) create mode 100644 ARCHITECTURE.md create mode 100644 README.md create mode 100644 REFERENCES.md create mode 100644 ROADMAP.md create mode 100644 SCHEMA.md diff --git a/ARCHITECTURE.md b/ARCHITECTURE.md new file mode 100644 index 0000000..503c18a --- /dev/null +++ b/ARCHITECTURE.md @@ -0,0 +1,615 @@ +# Object-Addressed Memory Manager for OpenCode + +## Project Codename: Mnemosyne + +> A transparent proxy that implements demand-paged, object-addressed memory management +> for LLM context windows. Extends Pichay's demand paging with semantic objects, +> multi-fidelity compression, declared losses, queryable backing store, and +> goal-aware retrieval via a helper LLM. + +--- + +## 1. Problem Statement + +LLM coding agents (opencode, Claude Code) suffer from context window bloat: + +- **21.8% of input tokens are structural waste** (Pichay, 2026): unused tool schemas (11%), + stale tool results reprocessed at 84.4x amplification (8.7%), duplicated content (2.2%) +- **Context drift** silently degrades reasoning quality before hitting hard token limits +- **Binary eviction** (resident vs evicted) is too coarse -- a 200-byte tombstone can't answer + questions about 8KB of evicted code +- **No semantic awareness** -- eviction is by file path, not by conceptual relevance to the + current task + +## 2. Design Principles + +1. **The context window is L1 cache, not memory.** Everything lives in the backing store; + context is a curated working set. +2. **Eviction is cooperative.** The model participates in eviction decisions via cleanup tags + and phantom tools. It has incentive: cleaner context = better attention quality. +3. **Compression is authored, not algorithmic.** The model (or a helper LLM) writes summaries + with declared losses. It knows what matters. +4. **The backing store is queryable.** The model can ask questions of evicted content without + materializing it. Micro-faults replace full page-ins. +5. **Objects, not blocks.** The unit of memory is a semantic object (a design decision, a + debugging session, a file understanding) -- not a fixed-size page keyed by file path. +6. **Transparency.** The proxy is invisible to the client and the inference API. No changes + to opencode or the model required. + +--- + +## 3. System Architecture + +``` + opencode (client) + | + | HTTP (Messages API) + v + +---------------------+ + | MNEMOSYNE PROXY | + | | + | +---------------+ | + | | Context | | +------------------+ + | | Assembler |--|---->| Helper LLM | + | | | | | (Haiku / local) | + | +---------------+ | | | + | | Fidelity | | | - Summarization | + | | Manager | | | - Loss declaration| + | +---------------+ | | - Micro-fault QA | + | | Object | | | - Segmentation | + | | Segmenter | | +------------------+ + | +---------------+ | + | | Fault | | +------------------+ + | | Detector | | | Object Store | + | +---------------+ | | (PostgreSQL + | + | | Phantom Tool |--|---->| pgvector) | + | | Handler | | | | + | +---------------+ | | - Full content | + | | Cleanup Tag | | | - Multi-fidelity | + | | Parser | | | summaries | + | +---------------+ | | - Embeddings | + | | Pressure | | | - Relationships | + | | Monitor | | | - Fault history | + | +---------------+ | +------------------+ + +---------------------+ + | + | HTTP (Messages API, modified) + v + Inference API (Anthropic) +``` + +### Component Responsibilities + +| Component | Role | +|---|---| +| **Context Assembler** | Builds the modified message array for each API call. Selects which objects are resident at which fidelity. Injects phantom tool definitions. | +| **Fidelity Manager** | Tracks current fidelity level of each object. Degrades fidelity under pressure. Upgrades on access. Manages the L0-L3 fidelity ladder. | +| **Object Segmenter** | Splits the conversation stream into semantic objects. Runs after each turn. Uses embedding coherence + structural signals (tool boundaries, topic shifts). | +| **Fault Detector** | Detects page faults (model re-requests evicted content). Records fault history for pinning decisions. Detects micro-fault queries. | +| **Phantom Tool Handler** | Intercepts phantom tool calls from the model's streaming response before they reach the client. Handles `memory_release`, `memory_query`, `memory_restore`. | +| **Cleanup Tag Parser** | Parses structured directives from the model's text output: `drop`, `summarize`, `anchor`, `collapse`. Extended with `declare_losses`. | +| **Pressure Monitor** | Tracks token consumption per request. Determines pressure zone (Normal/Caution/Warning/Critical). Triggers fidelity degradation. | +| **Helper LLM** | Cheap model (Haiku, GPT-4o-mini, or local qwen2.5) that authors summaries, declares losses, answers micro-fault queries, and assists with object segmentation. | +| **Object Store** | PostgreSQL + pgvector database holding all semantic objects at all fidelity levels, with embeddings, metadata, relationships, and fault history. | + +--- + +## 4. Memory Hierarchy + +``` ++----------------------------------------------------------------+ +| L0: Full Content (in context window) | +| - Current working set of semantic objects | +| - Full text, no compression | +| - Capacity: ~60% of context window budget | ++----------------------------------------------------------------+ +| L1: Detailed Summary (in context window) | +| - Model-authored summary, ~30% of original size | +| - Preserves: file paths, function names, decisions, errors | +| - Declared losses: specific values, exact code, edge cases | +| - Capacity: ~20% of context window budget | ++----------------------------------------------------------------+ +| L2: Compact Summary (in context window) | +| - Model-authored headline, ~5% of original size | +| - Preserves: what was done, what was decided, key files | +| - Declared losses: implementation details, reasoning | +| - Capacity: ~15% of context window budget | ++----------------------------------------------------------------+ +| L3: Metadata Stub (in context window) | +| - One-line description + type + timestamp | +| - ~50-100 tokens per object | +| - Capacity: ~5% of context window budget | ++----------------------------------------------------------------+ +| L4: Evicted (not in context, in backing store only) | +| - Not present in context at all | +| - Queryable via memory_query phantom tool | +| - Restorable via memory_restore phantom tool | ++----------------------------------------------------------------+ +| BACKING STORE (PostgreSQL + pgvector) | +| - All objects at all fidelity levels, always | +| - Full content preserved indefinitely | +| - Embeddings for semantic search | +| - Cross-session persistence (future) | ++----------------------------------------------------------------+ +``` + +### Fidelity Transitions + +``` +Pressure rising (token count increasing): + L0 (full) --[summarize]--> L1 (detailed) --[compress]--> L2 (compact) --[stub]--> L3 --[evict]--> L4 + +Access / fault (model needs content): + L4 --[memory_restore]--> L0 (full page-in) + L4 --[memory_query]--> (helper LLM answers, object stays at L4) + L3 --[model references]--> L1 or L0 (upgrade on access) + L2 --[model references]--> L1 (upgrade on access) +``` + +### Pressure Zones (token thresholds, configurable) + +| Zone | Token % | Action | +|---|---|---| +| **Normal** | < 50% | Observe only. No fidelity changes. | +| **Caution** | 50-70% | Degrade oldest L0 objects to L1. | +| **Warning** | 70-85% | Degrade L0 to L1, L1 to L2, oldest L2 to L3. | +| **Critical** | 85-95% | Aggressive degradation. L2+ to L3. Evict L3 to L4. | +| **Emergency** | > 95% | Force-evict everything except last 2 user turns + system prompt. | + +--- + +## 5. Semantic Objects + +### 5.1 Object Types + +| Type | Description | Example | +|---|---|---| +| `conversation_phase` | A coherent stretch of dialogue about one topic | "Discussed auth architecture for 8 turns" | +| `design_decision` | An explicit decision with rationale | "Chose JWT over sessions because..." | +| `debugging_session` | A sequence of diagnose-hypothesize-fix-verify | "Tracked down the race condition in..." | +| `file_context` | A file read and understanding | "Read src/auth/middleware.ts (200 lines)" | +| `tool_result` | Output from a tool call (grep, bash, etc.) | "grep found 15 matches for handleAuth" | +| `plan` | A structured plan or todo list | "Implementation plan: 5 steps..." | +| `error_context` | An error and its diagnosis | "TypeError at line 42, caused by..." | +| `external_reference` | Docs, API reference, examples pulled from outside | "React docs on useEffect cleanup" | + +### 5.2 Object Segmentation Algorithm + +Segmentation runs after each model turn. Two strategies, selected by context: + +**Strategy A: Structural Segmentation (fast, rule-based)** +- Tool call boundaries are natural object boundaries +- Each `Read` result = `file_context` object +- Each `Bash`/`Grep` result = `tool_result` object +- User message + assistant response = potential `conversation_phase` boundary +- Heuristic: if topic similarity (embedding cosine) between consecutive turns drops below + threshold (0.7), start a new `conversation_phase` + +**Strategy B: Semantic Segmentation (slower, higher quality)** +- Based on xMemory's sparsity-semantics objective (arXiv:2602.02007) +- Embed all messages with a lightweight model (all-MiniLM-L6-v2, ~20ms) +- Cluster by coherence: maximize inter-object semantic diversity, minimize intra-object + redundancy +- Build hierarchy: messages -> episodes -> themes +- Use for long sessions (>50 turns) where structural boundaries are insufficient + +**Default: Strategy A for turns 1-50, Strategy B kicks in at 50+ turns.** + +### 5.3 Object Relationships + +Objects have typed relationships stored in the backing store: + +``` +parent_of: conversation_phase -> design_decision (decision made during phase) +caused_by: error_context -> file_context (error was in this file) +references: debugging_session -> file_context (files examined during debug) +supersedes: file_context(v2) -> file_context(v1) (file re-read after edit) +depends_on: plan -> design_decision (plan relies on this decision) +``` + +Relationships are used by the Context Assembler: when upgrading an object's fidelity, +also consider upgrading its `depends_on` and `references` relationships. + +--- + +## 6. Multi-Fidelity Compression + +### 6.1 Summary Generation + +When an object degrades from L0 to L1, the Helper LLM generates a summary: + +**Input to Helper LLM:** +``` +You are a context compression engine for a coding agent. Summarize the following +content while preserving maximum utility for future reference. + +CONTENT TYPE: {object.type} +CONTENT: +{object.content_full} + +INSTRUCTIONS: +1. Write a detailed summary (~30% of original length) +2. MUST preserve: file paths, function names, variable names, library names, + error messages, decision rationale, specific values that may be referenced later +3. List DECLARED LOSSES: specific information you omitted that someone might need. + Be precise -- "specific error codes" not "some details" +4. List CAN_ANSWER: categories of questions this summary can answer without + needing the original content + +OUTPUT FORMAT (JSON): +{ + "summary": "...", + "losses": ["exact error code for token expiry", "rate limit threshold values", ...], + "can_answer": ["auth approach used", "middleware chain order", "why JWT over sessions", ...], + "key_entities": ["src/auth/middleware.ts", "handleAuth()", "jsonwebtoken", ...] +} +``` + +**L1 -> L2 compression** uses the L1 summary as input (not L0), with instruction to +compress to ~5% of original. Additional losses are accumulated. + +**L2 -> L3 stub** is generated from L2: +``` +[debugging_session | 2026-03-13 14:30 | Fixed race condition in auth token refresh + by adding mutex lock in src/auth/refresh.ts | 12 related objects] +``` + +### 6.2 Declared Losses Schema + +```typescript +interface DeclaredLosses { + // What was dropped from this fidelity level + dropped: string[] + + // What questions this fidelity level CAN still answer + can_answer: string[] + + // Hint for when to fault (what would require the original) + fault_when: string[] + + // Key entities preserved (for relationship tracking) + key_entities: string[] +} +``` + +### 6.3 Loss Accumulation + +As objects degrade through fidelity levels, losses accumulate: + +``` +L0: (full content, no losses) +L1: losses = ["exact error codes", "line-by-line implementation"] +L2: losses = L1.losses + ["function signatures", "reasoning chain"] +L3: losses = L2.losses + ["what was decided", "which files involved"] + (at this point, only the one-line description remains) +``` + +The accumulated `fault_when` list tells the model exactly when it needs to fault: +"If you need exact error codes, specific function signatures, or the reasoning +behind the auth decision, restore this object." + +--- + +## 7. Queryable Backing Store + +### 7.1 The `memory_query` Phantom Tool + +Instead of restoring a full object (8KB+) to answer a simple question, the model +calls `memory_query`: + +```json +{ + "tool": "memory_query", + "input": { + "question": "What error code does the auth middleware return for expired tokens?", + "scope": "auth-related objects", + "max_tokens": 200 + } +} +``` + +**Proxy handling:** + +1. Proxy intercepts `memory_query` from the model's streaming response +2. Proxy queries the Object Store: + - Embed the question + - Find top-k relevant objects by cosine similarity (even evicted ones) + - Retrieve their L0 (full content) from the backing store +3. Proxy sends question + retrieved full content to the Helper LLM +4. Helper LLM returns a targeted answer (~50-200 tokens) +5. Proxy injects the answer as a synthetic tool result into the model's context +6. The evicted objects stay evicted -- no fidelity change + +**Token savings per micro-fault:** +- Traditional fault (Pichay): restore full page, ~4,000-8,000 tokens +- Micro-fault: inject targeted answer, ~50-200 tokens +- Savings: 95-99% per fault + +### 7.2 Semantic Search for Micro-Faults + +The backing store supports multiple retrieval strategies: + +```sql +-- Vector similarity (primary) +SELECT * FROM semantic_objects +WHERE session_id = $1 +ORDER BY embedding <-> $query_embedding +LIMIT 5; + +-- Hybrid: vector + keyword (for exact matches) +SELECT * FROM semantic_objects +WHERE session_id = $1 + AND ( + to_tsvector('english', content_full) @@ plainto_tsquery('english', $query) + OR embedding <-> $query_embedding < 0.3 + ) +ORDER BY embedding <-> $query_embedding +LIMIT 5; + +-- Metadata-filtered (for typed queries) +SELECT * FROM semantic_objects +WHERE session_id = $1 + AND object_type = 'error_context' + AND 'src/auth' = ANY(tags) +ORDER BY created_at DESC +LIMIT 3; +``` + +### 7.3 `memory_restore` — Full Page-In + +When the model needs full content (editing a file, reviewing exact code), it calls +`memory_restore` which does a traditional page-in: + +```json +{ + "tool": "memory_restore", + "input": { + "object_id": "obj_abc123", + "reason": "Need to edit the auth middleware" + } +} +``` + +This upgrades the object to L0, potentially triggering eviction of other objects +under pressure. + +### 7.4 `memory_release` -- Cooperative Eviction + +The model can voluntarily release objects it no longer needs: + +```json +{ + "tool": "memory_release", + "input": { + "object_ids": ["obj_abc123", "obj_def456"], + "reason": "Done with auth implementation, moving to tests" + } +} +``` + +This immediately degrades the objects to L3 (or L4 under pressure), freeing context +budget for new work. + +--- + +## 8. Goal-Aware Retrieval + +### 8.1 The Problem + +When the model starts a new sub-task (e.g., "now write tests for auth"), the objects +currently in context may be irrelevant (e.g., old debugging sessions for a different +module). Goal-aware retrieval proactively swaps context based on the current task. + +### 8.2 Detection: When Has the Goal Changed? + +The Pressure Monitor also tracks goal transitions by comparing: +- The current user message embedding vs the previous user message embedding +- If cosine similarity < 0.5 (topic shift), trigger goal-aware retrieval + +### 8.3 Goal-Aware Context Assembly + +On goal transition: + +1. **Helper LLM classifies the new goal** (~100ms): + ``` + Given this user message: "{message}" + What is the user's current goal? What context would be most relevant? + Return: { "goal": "...", "relevant_types": [...], "relevant_tags": [...] } + ``` + +2. **Query the Object Store** for relevant objects: + ```sql + SELECT * FROM semantic_objects + WHERE session_id = $1 + ORDER BY embedding <-> $goal_embedding + LIMIT 20; + ``` + +3. **Rank objects by relevance to new goal** (helper LLM or embedding similarity) + +4. **Assemble new context window**: + - Top-ranked objects at L0 or L1 (depending on budget) + - Previously active but now irrelevant objects degraded to L2 or L3 + - Always preserve: system prompt, last 2 user turns, any pinned objects + +5. **Inject into next API call** via `experimental.chat.messages.transform` + +### 8.4 Predictive Loading + +After goal classification, the helper LLM can predict what the model will need next: + +``` +Given goal "write tests for auth", the model will likely need: +- The auth middleware implementation (file_context for src/auth/middleware.ts) +- The existing test patterns (file_context for tests/*) +- The design decision about JWT (design_decision) +- NOT: the debugging session for the database migration +``` + +Pre-load predicted objects at L1, so they're available if the model needs them. + +--- + +## 9. Admission Control (Write Path) + +Not everything deserves to become a stored object. Based on A-MAC (arXiv:2603.04549): + +### 9.1 Admission Score + +``` +S(m) = w_T * TypePrior(m) + w_N * Novelty(m) + w_U * Utility(m) + w_R * Recency(m) +``` + +| Factor | Signal | Weight (learned) | +|---|---|---| +| **TypePrior** | `design_decision` > `error_context` > `file_context` > `tool_result` | ~0.35 | +| **Novelty** | Cosine distance to nearest existing object > 0.15 | ~0.25 | +| **Utility** | Helper LLM scores future relevance (0-1) | ~0.25 | +| **Recency** | Exponential decay from creation time | ~0.15 | + +**Threshold:** S(m) >= 0.4 to admit. Below threshold, content is kept only in the +client's unmodified history (Pichay's backing store) but not indexed in the Object Store. + +### 9.2 What Gets Rejected + +- Routine tool results with no lasting value (e.g., `ls` output, `git status`) +- Duplicate file reads where content hasn't changed +- Conversation turns that are purely procedural ("Sure, I'll do that") + +--- + +## 10. Entropy-Gated Faulting (L-RAG Integration) + +Based on L-RAG (arXiv:2601.06551): use the model's own uncertainty as a fault signal. + +### 10.1 Mechanism + +During the model's generation (streaming response), monitor token-level entropy: + +1. **Normal entropy** (H < 1.5): model is confident, no intervention +2. **Elevated entropy** (1.5 < H < 2.2): model may benefit from more context. + Check if any L2/L3 objects match the current generation topic. If so, + silently upgrade to L1. +3. **High entropy** (H > 2.2): model is struggling. Trigger a micro-fault -- + query the backing store with the current generation context, inject relevant + information. + +### 10.2 Complementarity with Declared Losses + +Entropy-gated faulting handles the case where the model doesn't know what it +doesn't know. Declared losses handle the case where it does. Together: + +- **Declared losses**: "I know I need exact error codes, let me fault" + -> model calls `memory_query` +- **Entropy signal**: model's generation becomes uncertain around error handling + -> proxy automatically upgrades relevant objects + +### 10.3 Implementation Complexity + +Entropy monitoring requires access to token logprobs in the streaming response. +The Anthropic API provides these via `stream_options.include_logprobs`. This is a +Phase 4e feature due to the complexity of real-time entropy calculation during +streaming. + +--- + +## 11. Integration Points + +### 11.1 With OpenCode (via oh-my-opencode) + +The proxy can integrate at two levels: + +**Level 1: Pure Proxy (Phase 1-3)** +- Standalone HTTP proxy between opencode and Anthropic API +- Zero changes to opencode or oh-my-opencode +- Configuration: set `ANTHROPIC_BASE_URL` to proxy address + +**Level 2: Plugin Integration (Phase 4+)** +- oh-my-opencode hook: `experimental.chat.messages.transform` for context assembly +- oh-my-opencode hook: `experimental.session.compacting` for custom compaction +- oh-my-opencode hook: `tool.execute.after` for object segmentation on tool results +- MCP server exposing `memory_query`, `memory_stats`, `memory_objects` tools + (so the user can inspect memory state) + +### 11.2 With Pichay (fork and extend) + +Start from `fsgeek/pichay` (commit `b56701a`): +- `proxy.py` -> extend with multi-fidelity eviction, object segmentation +- `probe.py` -> extend with object-level analytics +- Add: `helper_llm.py` for summary generation, micro-fault QA +- Add: `object_store.py` for PostgreSQL + pgvector integration +- Add: `segmenter.py` for semantic object detection +- Add: `fidelity.py` for multi-fidelity state machine + +### 11.3 With the Helper LLM + +The helper LLM is called via standard API (Anthropic for Haiku, or Ollama for local): + +| Task | Model | Expected Latency | Tokens In | Tokens Out | +|---|---|---|---|---| +| Summarize L0 -> L1 | Haiku | ~200ms | ~2000 | ~600 | +| Compress L1 -> L2 | Haiku | ~100ms | ~600 | ~100 | +| Micro-fault answer | Haiku | ~150ms | ~3000 | ~100 | +| Goal classification | Haiku | ~100ms | ~200 | ~50 | +| Object segmentation | local (MiniLM) | ~20ms | embedding only | N/A | +| Admission scoring | local (qwen2.5) | ~50ms | ~500 | ~10 | + +**Cost estimate per session (200 turns):** +- ~50 summarizations: 50 * ~3000 tokens = ~150K Haiku tokens (~$0.004) +- ~20 micro-faults: 20 * ~3000 tokens = ~60K Haiku tokens (~$0.002) +- ~10 goal classifications: ~2K Haiku tokens (~$0.00005) +- **Total helper cost: ~$0.006 per session** +- **Savings on main model**: 50-93% context reduction on Opus/Sonnet calls + +--- + +## 12. Failure Modes and Mitigations + +| Failure Mode | Consequence | Mitigation | +|---|---|---| +| **Helper LLM produces bad summary** | Model loses critical info, silent quality degradation | Validate via declared losses. Spot-check: can helper answer `can_answer` queries from summary? | +| **Object segmentation too coarse** | Related content split across objects, fidelity changes break coherence | Conservative defaults (prefer larger objects). Relationship tracking keeps related objects together. | +| **Object segmentation too fine** | Too many small objects, overhead dominates | Minimum object size (500 tokens). Merge adjacent objects of same type. | +| **Thrashing** | Objects repeatedly degraded and restored, wasting helper LLM calls | Fault-driven pinning (Pichay L2). After 1 fault, pin at current fidelity for N turns. | +| **Goal misclassification** | Wrong objects loaded for current task | Conservative: always keep last 2 turns at L0. Don't evict below L2 on goal change (can upgrade quickly). | +| **Backing store latency spike** | Micro-fault takes >500ms, model generation stalls | Timeout + fallback: if backing store slow, inject L2 summary instead of querying. | +| **Declared losses are incomplete** | Model doesn't know it's missing info, doesn't fault | Entropy-gated faulting (Phase 4e) as safety net. Also: periodic loss audit by helper LLM. | +| **Helper LLM unavailable** | No summaries, no micro-faults | Graceful degradation: fall back to Pichay-style binary eviction with tombstones. | + +--- + +## 13. Metrics and Evaluation + +### 13.1 Primary Metrics + +| Metric | Target | How to Measure | +|---|---|---| +| **Context reduction** | >80% vs baseline | (baseline tokens - actual tokens) / baseline tokens | +| **Fault rate** | <0.1% | faults / total evictions | +| **Micro-fault success rate** | >90% | micro-faults that avoided full page-in / total micro-faults | +| **Task quality** | No degradation | LLM-judged equivalence: full-context vs managed-context outputs | +| **Helper LLM overhead** | <5% of main model cost | helper cost / main model cost | +| **Latency overhead** | <300ms per turn average | (managed turn time - baseline turn time) | + +### 13.2 Evaluation Method + +1. **Offline replay**: Replay recorded opencode sessions through the proxy. + Compare managed output vs original output via LLM judge. +2. **A/B testing**: Run identical tasks with and without proxy. Measure + token usage, task completion, and code quality. +3. **Fault analysis**: Log every fidelity transition, fault, and micro-fault. + Identify patterns in what causes faults (guides admission control tuning). + +--- + +## 14. Technology Stack + +| Component | Technology | Rationale | +|---|---|---| +| **Proxy** | Python (asyncio + httpx) | Fork from Pichay (Python). Streaming support critical. | +| **Object Store** | PostgreSQL 16 + pgvector | Proven at scale by Letta. Hybrid vector + relational. | +| **Embeddings** | all-MiniLM-L6-v2 (ONNX, local) | Fast (~20ms), no API dependency, good enough for similarity. | +| **Helper LLM** | Anthropic Haiku (primary) / Ollama qwen2.5 (fallback) | Haiku: fast + cheap. Ollama: offline capable. | +| **Streaming parser** | Custom SSE parser | Must parse tool calls from streaming response before client sees them. | +| **Config** | TOML | Simple, human-readable. | +| **Testing** | pytest + recorded session replay | Replay real sessions for regression testing. | diff --git a/README.md b/README.md new file mode 100644 index 0000000..19e69fe --- /dev/null +++ b/README.md @@ -0,0 +1,7 @@ +# Mnemosyne + +Object-addressed context memory for LLM agents. Built on [Pichay](https://github.com/fsgeek/pichay)'s demand paging foundation. + +## Status + +Work in progress. See [ARCHITECTURE.md](ARCHITECTURE.md) and [ROADMAP.md](ROADMAP.md) for design details. diff --git a/REFERENCES.md b/REFERENCES.md new file mode 100644 index 0000000..1f8fd90 --- /dev/null +++ b/REFERENCES.md @@ -0,0 +1,220 @@ +# Research References + +All papers, repositories, and prior art that informed this design. + +--- + +## Core Papers + +### Pichay — Demand Paging for LLM Context Windows (PRIMARY) +- **Paper:** [The Missing Memory Hierarchy: Demand Paging for LLM Context Windows](https://arxiv.org/abs/2603.09023) +- **Author:** Tony Mason (UBC / Georgia Tech) +- **Date:** March 2026, accepted ACM SIGOPS +- **Repo:** https://github.com/fsgeek/pichay (tag: v0.1.0-paper, commit b56701a) +- **Archival:** https://doi.org/10.5281/zenodo.18930122 +- **Key findings:** 21.8% structural waste across 857 sessions / 4.45B tokens. + 93% context reduction in live deployment. 0.0254% fault rate over 1.4M evictions. + Cooperative eviction via phantom tools and cleanup tags. FIFO eviction with + pressure zones. Transparent HTTP proxy architecture. +- **Used in:** Phase 1 (fork baseline), Phase 2 (pressure zones, cleanup tags), + Phase 3 (phantom tools) + +### MemGPT / Letta — Virtual Memory for LLMs +- **Paper:** [MemGPT: Towards LLMs as Operating Systems](https://arxiv.org/abs/2310.08560) +- **Authors:** Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, + Ion Stoica, Joseph E. Gonzalez (UC Berkeley) +- **Date:** October 2023 (revised February 2024) +- **Repo:** https://github.com/letta-ai/letta (SHA: 4cb2f21c) +- **Key findings:** Three-tier memory hierarchy (core/recall/archival). Agent-initiated + paging via tool calls. PostgreSQL + pgvector for archival storage. Partial-evict + summarization (30% oldest messages). LLM-driven retrieval is surprisingly effective. +- **Used in:** Object Store design (SCHEMA.md), multi-fidelity concept, backing store + architecture (Phase 3) + +### xMemory — Hierarchical Structured Retrieval +- **Paper:** [Beyond RAG for Agent Memory: Retrieval by Decoupling and Aggregation](https://arxiv.org/abs/2602.02007) +- **Venue:** ICML 2026 +- **Key findings:** Standard RAG on agent memory fails due to correlated content. + Hierarchical retrieval (messages -> episodes -> semantics -> themes) prevents + redundant retrieval. Sparsity-semantics objective for segmentation. Top-down + retrieval reduces retrieved tokens while improving relevance. +- **Used in:** Phase 6 (xMemory hierarchy), Phase 4a (segmentation concept) + +### L-RAG — Entropy-Based Lazy Context Loading +- **Paper:** [L-RAG: Balancing Context and Retrieval with Entropy-Based Lazy Loading](https://arxiv.org/abs/2601.06551) +- **Date:** January 2026 +- **Key findings:** Token entropy reliably predicts model uncertainty (H=1.72 correct + vs H=2.20 errors, p<0.001). 26% retrieval reduction at balanced threshold. + Training-free. Works with any model. +- **Used in:** Phase 4e (entropy-gated faulting) + +### A-MAC — Adaptive Memory Admission Control +- **Paper:** [Adaptive Memory Admission Control for LLM Agents](https://arxiv.org/abs/2603.04549) +- **Authors:** Workday AI +- **Date:** March 2026 +- **Repo:** https://github.com/GuilinDev/Adaptive_Memory_Admission_Control_LLM_Agents +- **Key findings:** 5-factor admission scorer (Utility, Confidence, Novelty, Recency, + TypePrior). TypePrior is most influential factor. Uses local LLM (Ollama/qwen2.5) + for utility scoring. F1=0.583 on LoCoMo. 31% faster than LLM-native memory. +- **Used in:** Phase 4d (admission control) + +--- + +## Supporting Papers + +### Factory — Anchored Iterative Summarization +- **Source:** Factory's evaluation across 36,000 engineering sessions +- **Key findings:** Anchored summarization (persistent state with intent/changes/decisions/ + next_steps) outperforms rolling reconstruction. Scores: Factory 4.04 vs Anthropic 3.74 + vs OpenAI 3.43 on accuracy/completeness/continuity. +- **Used in:** Phase 2 (multi-fidelity compression design) + +### SWE-Pruner — Neural Context Pruning for Coding +- **Authors:** Wang et al., 2026 +- **Key findings:** 0.6B-parameter neural skimmer for task-aware pruning. 23-54% token + reduction on SWE-bench. Maintains solve rates. +- **Referenced for:** Alternative approach to context reduction (learned pruning vs + semantic objects) + +### ACON — Failure-Driven Compression Optimization +- **Paper:** arXiv, October 2025 +- **Key findings:** Unified history + observation compression. 26-54% peak context + reduction. Gradient-free, works with API models. Iteratively refines compression + prompt based on failure cases. +- **Referenced for:** Compression strategy comparison + +### Neural Paging — Learned Page Controller +- **Paper:** [Neural Paging: Learning Context Management Policies for Turing-Complete Agents](https://arxiv.org/abs/2603.02228) +- **Date:** February 2026 +- **Key findings:** Differentiable page controller. Semantic Belady's optimality. + Reduces O(N^2) to O(N*K^2) complexity. Theoretical framework. +- **Referenced for:** Future work (learned eviction policy) + +### CMV — DAG-Based Session History Trimming +- **Author:** Santoni, 2026 +- **Key findings:** DAG-based session history structure. Structurally lossless trimming. + Up to 86% reduction for tool-heavy sessions. +- **Referenced for:** Alternative structural approach + +### MemOS — Memory Operating System for AGI +- **Authors:** Li et al., 2025 +- **Key findings:** Full "Memory OS" with lifecycle control and persistent representations. +- **Referenced for:** Long-term architecture vision + +### SideQuest — KV Cache Eviction via Parallel Reasoning +- **Authors:** Kariyappa & Suh, 2026 +- **Key findings:** Fine-tuned parallel reasoning thread for KV cache eviction. + 56-65% peak memory reduction. Irreversible eviction. +- **Referenced for:** KV-cache-level optimization (complementary to our message-level approach) + +### Quest — Query-Aware KV Cache Sparsity +- **Paper:** [Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference](https://arxiv.org/abs/2406.10774) +- **Venue:** ICML 2024, MIT Han Lab +- **Repo:** https://github.com/mit-han-lab/Quest +- **Key findings:** 2.23x self-attention speedup, 7.03x inference latency reduction. + Query-aware page selection within KV cache. +- **Referenced for:** Within-model context selection (different layer than our system) + +### SpeContext — Speculative Context Sparsity +- **Paper:** [SpeContext: Enabling Efficient Long-context Reasoning](https://arxiv.org/abs/2512.00722) +- **Authors:** SJTU / Infinigence-AI, November 2025 +- **Key findings:** Small draft model predicts important KV cache tokens before main + model runs. Analogous to speculative decoding but for context selection. +- **Referenced for:** Helper model concept (similar philosophy at different layer) + +### SoK: Agentic RAG +- **Paper:** [SoK: Agentic RAG: Taxonomy, Architectures, Evaluation](https://arxiv.org/abs/2603.07379) +- **Date:** March 2026 +- **Key findings:** Definitive 2026 survey. Taxonomy of planning, retrieval, memory, + and tool coordination patterns. Identifies risks: compounding hallucination, + memory poisoning, retrieval misalignment. +- **Referenced for:** Taxonomy and risk awareness + +### Mem0 — Fact Extraction + Merge Pipeline +- **Paper:** [Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory](https://arxiv.org/abs/2504.19413) +- **Repo:** https://github.com/mem0ai/mem0 (49,561 stars) +- **Key findings:** 2-LLM-call pipeline (extract facts -> diff/merge with existing). + +26% accuracy over OpenAI Memory on LOCOMO. 91% faster, 90% fewer tokens. + 20+ vector store backends. +- **Referenced for:** Future cross-session memory (Phase 6+) + +--- + +## Key Repositories + +### Direct Dependencies +| Repo | What We Use | Phase | +|---|---|---| +| [fsgeek/pichay](https://github.com/fsgeek/pichay) | Fork as starting point for proxy | Phase 1 | +| [pgvector/pgvector](https://github.com/pgvector/pgvector) | PostgreSQL vector similarity | Phase 3+ | +| [sentence-transformers](https://github.com/UKPLab/sentence-transformers) | all-MiniLM-L6-v2 embeddings | Phase 3+ | + +### Reference Implementations +| Repo | What We Learn From | Stars | +|---|---|---| +| [letta-ai/letta](https://github.com/letta-ai/letta) | 3-tier memory architecture, archival search | 15k+ | +| [mem0ai/mem0](https://github.com/mem0ai/mem0) | Fact extraction pipeline, multi-backend vector store | 49k+ | +| [alibaizhanov/mengram](https://github.com/alibaizhanov/mengram) | 3-memory-type system (semantic/episodic/procedural) | 86 | +| [PavanVkAlapati/memory_orchestration](https://github.com/PavanVkAlapati/memory_orchestration) | Layered memory with Qdrant + Redis + MongoDB | - | +| [GuilinDev/Adaptive_Memory_Admission_Control_LLM_Agents](https://github.com/GuilinDev/Adaptive_Memory_Admission_Control_LLM_Agents) | A-MAC admission scoring | - | +| [vivek-tiwari-vt/agmem](https://github.com/vivek-tiwari-vt/agmem) | Git-like version control for agent memories | - | +| [lm-sys/RouteLLM](https://github.com/lm-sys/RouteLLM) | BERT classifier router for model selection | - | + +### MCP Servers (reference for Phase 5) +| Repo | What It Does | +|---|---| +| [adamrdrew/agent-memory-mcp](https://github.com/adamrdrew/agent-memory-mcp) | Hybrid BM25 + vector search, local embeddings, 12 memory categories | +| [Parswanadh/memory-mcp-server](https://github.com/Parswanadh/memory-mcp-server) | 3-tier hierarchical memory (working/short-term/long-term) | +| [vbcherepanov/claude-total-memory](https://github.com/vbcherepanov/claude-total-memory) | 4-tier search, 20 tools, ChromaDB + SQLite | +| [van-reflect/Reflect-Memory](https://github.com/van-reflect/Reflect-Memory) | Cross-agent memory, vendor-neutral | + +--- + +## OpenCode / Oh-My-OpenCode Integration Points + +### OpenCode Plugin Hooks (from sst/opencode) +| Hook | Location | Purpose for Mnemosyne | +|---|---|---| +| `experimental.chat.messages.transform` | `packages/opencode/src/session/prompt.ts:652` | Modify message array before LLM call (context assembly) | +| `experimental.session.compacting` | `packages/opencode/src/session/compaction.ts:169` | Custom compaction prompt/context | +| `experimental.chat.system.transform` | `packages/opencode/src/session/llm.ts:84` | Modify system prompt (inject memory instructions) | +| `tool.execute.before` | `packages/plugin/src/index.ts:184` | Intercept tool args before execution | +| `tool.execute.after` | `packages/plugin/src/index.ts:192` | Process tool results for object creation | +| `chat.params` | `packages/opencode/src/session/llm.ts:114` | Modify temperature, options | + +### Oh-My-OpenCode Hooks (from omc-sh/oh-my-opencode) +| Hook | Purpose for Mnemosyne | +|---|---| +| `context-window-monitor` | Existing hook -- can extend or replace | +| `preemptive-compaction` | Existing hook -- integrate with our pressure system | +| `tool-output-truncator` | Existing hook -- our fidelity system supersedes this | +| `compaction-context-injector` | Inject our memory state into compaction prompt | + +--- + +## Benchmark Datasets + +For evaluating memory quality: + +| Dataset | What It Tests | URL | +|---|---|---| +| LoCoMo | Long-conversation memory (QA over multi-session chat) | https://github.com/letta-ai/letta/tree/main/tests | +| PerLTQA | Personalized long-term QA | Referenced in xMemory paper | +| SWE-bench | Coding task completion (for measuring quality impact) | https://github.com/princeton-nlp/SWE-bench | +| Terminal-Bench | CLI agent task completion | Referenced in Letta Code evaluation | + +--- + +## Key Metrics from Literature + +| System | Context Reduction | Quality Impact | Cost | +|---|---|---|---| +| Pichay (baseline eviction) | 37% token, up to 93% extreme | 0.0254% fault rate | Zero (proxy only) | +| SWE-Pruner | 23-54% | Maintains solve rates | Training cost for 0.6B model | +| ACON | 26-54% peak | 95%+ task accuracy preserved | Multiple LLM calls for training | +| Factory summarization | High | 4.04/5 accuracy score | 1 LLM call per eviction | +| Cursor lazy MCP loading | 46.9% | No degradation | Zero (lazy loading) | +| Cline file deduplication | Variable | None (lossless) | Zero (dedup only) | +| Simple observation masking | ~50% | Matches LLM summarization | Zero | +| L-RAG entropy gating | 26% retrieval reduction | Marginal impact | Logprob monitoring | +| RouteLLM model routing | 85% cost reduction | 95% quality maintained | <10ms per route | diff --git a/ROADMAP.md b/ROADMAP.md new file mode 100644 index 0000000..e399419 --- /dev/null +++ b/ROADMAP.md @@ -0,0 +1,487 @@ +# Implementation Roadmap + +## Overview + +9 phases, incrementally building from a working Pichay fork to the full +object-addressed memory system. Each phase produces a testable, usable artifact. + +**Estimated total effort: 8-12 weeks for a solo developer.** + +--- + +## Phase 1: Pichay Baseline (Week 1) + +**Goal:** Get the existing Pichay proxy running between opencode and Anthropic API. +Validate structural waste reduction on real sessions. + +### Tasks + +- [ ] **1.1** Fork `fsgeek/pichay` at tag `v0.1.0-paper` (commit `b56701a`) +- [ ] **1.2** Set up development environment + - Python 3.11+, asyncio, httpx + - Local opencode instance + - Proxy configuration: `ANTHROPIC_BASE_URL=http://localhost:8080` +- [ ] **1.3** Run proxy in passthrough mode (no eviction) + - Verify: opencode works normally through proxy + - Log: request/response sizes, token counts, tool call inventory +- [ ] **1.4** Enable FIFO eviction (Pichay default settings) + - tau = 4 turns, s_min = 500 bytes + - Verify: tombstones appear for old tool results + - Measure: tokens saved, fault rate +- [ ] **1.5** Record 5+ real coding sessions through the proxy + - Use `probe.py` to generate session analytics + - Baseline metrics: waste %, amplification factor, fault rate +- [ ] **1.6** Write session replay infrastructure + - Record full message traces (request + response pairs) + - Replay tool for offline testing of later phases + +### Deliverable +Working proxy with FIFO eviction. Baseline metrics on real sessions. + +### Success Criteria +- Proxy is transparent (opencode works identically) +- Measurable token reduction (target: >15%) +- Fault rate < 0.1% +- 5+ recorded sessions for replay testing + +--- + +## Phase 2: Multi-Fidelity Eviction (Weeks 2-3) + +**Goal:** Replace binary eviction (resident/tombstone) with graduated fidelity levels. +Introduce the Helper LLM for summary generation. + +### Tasks + +- [ ] **2.1** Implement the Fidelity Manager state machine + - States: L0 (full), L1 (detailed summary), L2 (compact summary), L3 (stub), L4 (evicted) + - Transitions: degrade (pressure), upgrade (access/fault), pin (fault-driven) + - In-memory state per session (no DB yet) +- [ ] **2.2** Implement pressure zones + - Normal (<50%), Caution (50-70%), Warning (70-85%), Critical (85-95%), Emergency (>95%) + - Token counting per fidelity level + - Configurable thresholds via TOML config +- [ ] **2.3** Integrate Helper LLM (Anthropic Haiku) + - API client with retry, timeout, error handling + - Prompt template for L0 -> L1 summarization (detailed summary + declared losses) + - Prompt template for L1 -> L2 compression + - Response parsing: extract summary, losses, can_answer, key_entities +- [ ] **2.4** Implement fidelity degradation on pressure + - Walk objects oldest-first + - Generate summaries via Helper LLM on transition + - Replace content in message array with summary + loss declaration + - Format: `[Summary of {type}: {stub}]\n{summary}\n[Cannot answer: {losses}]` +- [ ] **2.5** Implement fidelity upgrade on access + - When model references an L1/L2/L3 object (detected by content overlap or tool call), + upgrade to L0 + - Restore full content from in-memory cache (client history is the backing store) +- [ ] **2.6** Extend fault detection for fidelity-aware faults + - L3 stub referenced -> upgrade to L1 (not necessarily L0) + - L2 compact referenced -> upgrade to L1 + - Only full page-in if model explicitly re-requests (tool call match) +- [ ] **2.7** Add declared losses to eviction tombstones + - Format: `[Paged out: {stub}. Lost: {losses}. Restore if you need: {fault_when}]` +- [ ] **2.8** Replay testing against Phase 1 baseline + - Compare: token reduction, fault rate, summary quality + - Manual review: are declared losses accurate? + +### Deliverable +Multi-fidelity proxy with Helper LLM summarization and declared losses. + +### Success Criteria +- Token reduction > 40% (up from Phase 1's ~20%) +- Fault rate < 0.05% (better than binary eviction -- summaries prevent unnecessary faults) +- Helper LLM cost < $0.01 per session +- Declared losses are accurate in >90% of spot checks + +--- + +## Phase 3: Queryable Backing Store + Micro-Faults (Weeks 4-5) + +**Goal:** Add PostgreSQL + pgvector as persistent backing store. Implement `memory_query` +phantom tool for micro-faults. This is the single highest-impact feature. + +### Tasks + +- [ ] **3.1** Set up PostgreSQL + pgvector + - Docker Compose for local dev + - Schema from SCHEMA.md (sessions, semantic_objects tables) + - Connection pooling (asyncpg) +- [ ] **3.2** Implement Object Store module + - CRUD operations for semantic_objects + - Embedding generation (all-MiniLM-L6-v2 via sentence-transformers, ONNX runtime) + - Store objects on creation (every tool result, message span) + - Full content always preserved in DB regardless of context fidelity +- [ ] **3.3** Implement semantic search + - Vector similarity search (pgvector cosine distance) + - Full-text search (PostgreSQL tsvector) + - Hybrid search (weighted combination) + - Metadata-filtered search (by type, tags, date range) +- [ ] **3.4** Implement `memory_query` phantom tool + - Add tool definition to phantom tool injection + - Intercept from streaming response + - Flow: intercept -> query Object Store -> send top-k results to Helper LLM -> inject answer + - Synthetic tool result format: `[Memory Query Result]\nQ: {question}\nA: {answer}\n[Source: {object stubs}]` +- [ ] **3.5** Implement `memory_restore` phantom tool + - Full page-in from backing store (traditional fault) + - Upgrade object to L0 + - Trigger pressure-based eviction of other objects if needed +- [ ] **3.6** Implement `memory_release` phantom tool + - Model voluntarily releases objects + - Immediate degradation to L3 or L4 + - Log cooperative eviction event +- [ ] **3.7** Implement phantom tool injection + - Add phantom tool definitions to the model's tool list in each request + - Parse phantom tool calls from streaming response BEFORE client receives them + - Handle phantom tool results transparently +- [ ] **3.8** Measure micro-fault effectiveness + - Track: micro-faults attempted, successful (model didn't need to restore after), + tokens saved per micro-fault + - Compare: micro-fault answer quality vs full restore (LLM-judged) + +### Deliverable +Full proxy with persistent backing store, semantic search, and micro-fault capability. + +### Success Criteria +- Token reduction > 60% +- Micro-fault success rate > 85% (model doesn't need full restore after micro-fault) +- Micro-fault latency < 500ms (embedding + DB query + Helper LLM) +- Token savings per micro-fault: >95% vs full restore +- DB query latency < 50ms for semantic search + +--- + +## Phase 4a: Object Segmentation (Week 6) + +**Goal:** Replace file-path-based page identity with semantic object detection. +Conversations are segmented into coherent objects with types and relationships. + +### Tasks + +- [ ] **4a.1** Implement structural segmentation (Strategy A) + - Rule-based: tool call boundaries -> tool_result / file_context objects + - User turn + assistant response = conversation span + - Topic shift detection: embedding cosine between consecutive spans < 0.7 -> new object + - Object type classification: rules based on tool name, content patterns + - Read tool -> file_context + - Bash/Grep tool -> tool_result + - Error in output -> error_context + - "I'll implement..." / plan language -> plan + - "Let's use X because Y" / decision language -> design_decision +- [ ] **4a.2** Implement object type classifier + - Helper LLM or simple keyword-based classifier + - Input: content span + preceding context + - Output: object_type + stub + tags + - Latency budget: <100ms (prefer rules, fallback to Helper LLM) +- [ ] **4a.3** Retroactive segmentation + - On session start: no segmentation (objects created per-tool-result) + - Every 10 turns: re-examine recent objects, merge small ones, split large ones + - Merge criterion: consecutive objects of same type with embedding similarity > 0.8 + - Split criterion: single object > 5000 tokens with internal topic shift +- [ ] **4a.4** Object deduplication + - file_context: dedup by source_key (file path). New read supersedes old. + - tool_result: no dedup (each is unique) + - conversation_phase: no dedup + - Add 'supersedes' relationship when replacing +- [ ] **4a.5** Store relationships + - Automatic: file_context referenced in debugging_session -> 'references' edge + - Automatic: design_decision made during conversation_phase -> 'parent_of' edge + - Detection: key_entities overlap between objects suggests relationship +- [ ] **4a.6** Test segmentation quality + - Replay recorded sessions + - Manual review: do objects correspond to intuitive "chunks" of work? + - Measure: average object size, type distribution, relationship density + +### Deliverable +Proxy segments conversations into typed semantic objects with relationships. + +### Success Criteria +- Objects correspond to intuitive conversation segments (manual review) +- Average object size: 500-3000 tokens (not too fine, not too coarse) +- Type classification accuracy > 85% +- Segmentation latency < 50ms per turn (structural) or < 200ms (semantic) + +--- + +## Phase 4b: Object Relationships + Co-Fidelity (Week 7) + +**Goal:** Use relationships between objects to make smarter fidelity decisions. +When one object is upgraded, related objects are considered for upgrade too. + +### Tasks + +- [ ] **4b.1** Implement relationship-aware fidelity management + - When upgrading object X to L0, check `depends_on` and `references` edges + - If related object Y is at L2+, upgrade to L1 (not L0 -- don't over-promote) + - Configurable: max relationship hops (default: 2), max co-upgrades (default: 3) +- [ ] **4b.2** Implement relationship-aware eviction + - When degrading object X, DON'T degrade objects that X `depends_on` if they're + actively referenced by other L0 objects + - Eviction priority: objects with no inbound `references` or `depends_on` edges + are evicted first (they're "leaf" objects) +- [ ] **4b.3** Relationship visualization (debug tool) + - CLI command: `mnemosyne graph --session ` + - Output: DOT graph of objects + relationships + fidelity levels + - For debugging: identify orphaned objects, over-connected clusters +- [ ] **4b.4** Test co-fidelity management + - Scenario: model starts working on auth -> auth objects at L0, related files at L1 + - Model switches to tests -> auth demoted to L1/L2, test objects promoted + - Model returns to auth -> auth restored, relationships pull in dependencies + +### Deliverable +Relationship-aware fidelity management. Objects that belong together stay together. + +### Success Criteria +- Co-fidelity reduces fault rate by >20% vs independent fidelity management +- No "orphaned dependency" faults (model needs X, but X's dependency Y is evicted) + +--- + +## Phase 4c: Goal-Aware Retrieval (Week 8) + +**Goal:** When the user's goal changes, proactively swap context. The Helper LLM +reads the new goal, queries the Object Store, and assembles a focused context window. + +### Tasks + +- [ ] **4c.1** Implement goal transition detection + - Embed each user message + - Compare to previous user message embedding (cosine similarity) + - Threshold: < 0.5 = major topic shift, trigger goal-aware retrieval + - Also detect explicit signals: "now let's work on...", "moving to...", "switching to..." +- [ ] **4c.2** Implement goal classification + - Helper LLM call (~100ms): + Input: current user message + last 2 turns of context + Output: { goal, relevant_types, relevant_tags, predicted_needs } + - Cache goal classification (don't re-classify if message is a follow-up) +- [ ] **4c.3** Implement context swap on goal change + - Query Object Store: find top-20 objects by similarity to new goal + - Rank by: embedding similarity * recency_weight * type_match_bonus + - Assemble new context: + - Always: system prompt, last 2 user turns, pinned objects + - From goal query: top objects at appropriate fidelity (budget permitting) + - Previously active but now irrelevant: degrade to L2/L3 (not L4 -- recent work) +- [ ] **4c.4** Implement predictive loading + - Helper LLM predicts what the model will need for this goal + - Pre-load predicted objects at L1 (ready for quick upgrade) + - Track prediction accuracy: did the model actually access predicted objects? +- [ ] **4c.5** Test goal-aware retrieval + - Scenario: multi-task session (auth -> tests -> docs -> bugfix) + - Measure: context relevance at each goal transition + - Compare: goal-aware vs naive (no swap) in token efficiency and fault rate + +### Deliverable +Proxy proactively loads relevant context on goal transitions. + +### Success Criteria +- Goal detection accuracy > 80% (catches real transitions, few false positives) +- Context relevance after swap > 70% (measured by: did the model use the loaded objects?) +- Predictive loading accuracy > 50% (better than random) +- No regressions in token efficiency or fault rate + +--- + +## Phase 4d: Admission Control (Week 9) + +**Goal:** Not everything deserves to be a stored object. Score incoming content +and reject low-value items to keep the Object Store clean. + +### Tasks + +- [ ] **4d.1** Implement admission scorer + - Four factors: TypePrior, Novelty, Utility, Recency + - TypePrior: static weights per object_type (design_decision=1.0, tool_result=0.4, etc.) + - Novelty: cosine distance to nearest existing object in session (>0.15 = novel) + - Utility: Helper LLM scores future relevance 0-1 (or local model via Ollama) + - Recency: exponential decay from turn number + - Configurable weights (default from A-MAC: T=0.35, N=0.25, U=0.25, R=0.15) +- [ ] **4d.2** Implement admission threshold + - Default: 0.4 + - Items below threshold: not stored in Object Store + - Still present in client's unmodified history (Pichay backing store) + - Log rejected items for threshold tuning +- [ ] **4d.3** Implement rejection patterns + - Always reject: `ls` output, `git status`, routine directory listings + - Always reject: duplicate file reads where content hash matches existing object + - Always reject: purely procedural assistant responses ("Sure, I'll do that") + - Configurable: rejection rules in TOML config +- [ ] **4d.4** Threshold tuning + - Replay recorded sessions with different thresholds + - Find threshold that minimizes (false rejections * fault_cost + storage * store_cost) + - Log admission_scores table for analysis + +### Deliverable +Admission gate filters low-value content from the Object Store. + +### Success Criteria +- >30% of tool results rejected (routine/duplicate content) +- No false rejections that cause faults later (measure: rejected items that model + would have needed, detected via fault-after-reject tracking) +- Object Store grows linearly with session complexity, not session length + +--- + +## Phase 4e: Entropy-Gated Faulting (Week 10) + +**Goal:** Use the model's token-level entropy during generation as an automatic +signal to inject more context. Complements declared losses. + +### Tasks + +- [ ] **4e.1** Implement logprob extraction from streaming response + - Anthropic API: `stream_options.include_logprobs` (if available) + - Parse token logprobs from SSE stream in real-time + - Calculate rolling entropy: H = -sum(p * log(p)) over top-k logprobs +- [ ] **4e.2** Implement entropy monitoring + - Rolling window: last 20 tokens + - Thresholds: normal (H < 1.5), elevated (1.5-2.2), high (H > 2.2) + - Debounce: don't trigger on single high-entropy token (require 3+ consecutive) +- [ ] **4e.3** Implement entropy-triggered context injection + - On elevated entropy: check if any L2/L3 objects match current generation topic + - Extract current generation context (last 50 tokens) + - Embed and search Object Store + - If match found: silently upgrade to L1 in NEXT request (can't modify current) + - On high entropy: more aggressive -- prepare micro-fault answer for likely question +- [ ] **4e.4** Evaluate entropy signal reliability + - Compare: entropy at points where model made errors vs correct generation + - Calibrate thresholds per model (Opus vs Sonnet vs Haiku have different baselines) + - Measure: false positive rate (elevated entropy but model was fine) + +### Deliverable +Proxy monitors generation entropy and proactively loads context when model is uncertain. + +### Success Criteria +- Entropy signal detects genuine uncertainty >70% of the time +- False positive rate < 30% (elevated entropy that didn't need intervention) +- Measurable quality improvement on tasks where entropy-gating activated +- Note: This is the most experimental phase. Success criteria may be revised. + +### Fallback +If logprobs are not reliably available from the API, this phase can be deferred. +The system works well without it -- declared losses + manual faulting cover most cases. + +--- + +## Phase 5: OpenCode Plugin Integration (Week 11) + +**Goal:** Package as an oh-my-opencode plugin with a companion MCP server for +user-facing memory inspection tools. + +### Tasks + +- [ ] **5.1** Create oh-my-opencode plugin package + - npm package: `opencode-mnemosyne` + - Hook: `experimental.chat.messages.transform` for context assembly + - Hook: `experimental.session.compacting` for custom compaction + - Hook: `tool.execute.after` for object creation on tool results + - Configuration via `oh-my-opencode.json` +- [ ] **5.2** Create MCP server for user-facing tools + - `memory_stats`: show current context pressure, object counts by type/fidelity + - `memory_objects`: list all objects with fidelity, type, stub + - `memory_inspect `: show object detail (all fidelity levels, losses, relationships) + - `memory_graph`: show object relationship graph + - `memory_config`: view/update runtime configuration +- [ ] **5.3** Documentation + - Installation guide + - Configuration reference + - Troubleshooting guide + - Architecture overview for contributors +- [ ] **5.4** Session dashboard (optional) + - Local web UI (served by proxy) showing: + - Real-time context pressure gauge + - Object timeline (creation, fidelity transitions, faults) + - Token savings over time + - Fault log + +### Deliverable +Installable plugin + MCP server. Users can inspect and configure memory behavior. + +--- + +## Phase 6: Semantic Segmentation + xMemory Hierarchy (Week 12) + +**Goal:** Replace rule-based segmentation with xMemory's hierarchical approach +for long sessions. Build the full messages -> episodes -> semantics -> themes hierarchy. + +### Tasks + +- [ ] **6.1** Implement xMemory-style sparsity-semantics segmentation + - Embed all messages in a session with all-MiniLM-L6-v2 + - Cluster by coherence using the sparsity-semantics objective: + maximize inter-cluster diversity, minimize intra-cluster redundancy + - Output: episodes (coherent sub-conversations) +- [ ] **6.2** Build hierarchy + - Level 0: individual messages/tool results (existing objects) + - Level 1: episodes (groups of related objects, from clustering) + - Level 2: semantics (abstract themes spanning multiple episodes) + - Level 3: themes (top-level categories for the entire session) +- [ ] **6.3** Hierarchical retrieval + - Top-down: query matches theme -> expand to semantics -> expand to episodes -> expand to objects + - Only expand when similarity score justifies it (reader uncertainty reduction) + - Prevents redundant retrieval (a key xMemory advantage over flat search) +- [ ] **6.4** Incremental hierarchy maintenance + - Don't rebuild from scratch every turn + - New objects: assign to nearest episode, update episode embedding + - Every 20 turns: re-cluster to catch topic drift + - Major goal change: full re-hierarchy +- [ ] **6.5** Hierarchy-aware fidelity management + - When an episode is at L2, all its objects are at L2 or lower + - Upgrading an episode promotes its most relevant objects to L1 + - Themes can have their own summaries (super-summaries of episode summaries) + +### Deliverable +Full hierarchical segmentation for long sessions (>50 turns). + +### Success Criteria +- Retrieval quality improves for sessions >100 turns (measured by LLM-judged relevance) +- Hierarchy reduces redundancy in retrieved context (measured by token overlap between results) +- Incremental maintenance is fast (<500ms per turn) + +--- + +## Dependencies Between Phases + +``` +Phase 1 (Pichay baseline) + | +Phase 2 (Multi-fidelity + Helper LLM) + | +Phase 3 (Backing store + micro-faults) + / \ + / \ +4a 4d (can run in parallel) +(segmentation) (admission control) + | +4b (relationships + co-fidelity) + | +4c (goal-aware retrieval) + | +4e (entropy-gated faulting) -- optional, experimental + | +Phase 5 (plugin integration) + | +Phase 6 (xMemory hierarchy) +``` + +Phases 4a-4e can be partially parallelized: +- 4a + 4d can be built simultaneously +- 4b depends on 4a +- 4c depends on 4a + 4b +- 4e is independent (depends only on Phase 3) +- Phase 5 can start after Phase 3 (plugin wrapping doesn't need 4a-4e) +- Phase 6 depends on 4a (needs basic segmentation first) + +--- + +## Risk Register + +| Risk | Likelihood | Impact | Mitigation | +|---|---|---|---| +| Anthropic API doesn't expose logprobs for streaming | Medium | Phase 4e blocked | Phase 4e is optional. System works without entropy gating. | +| Helper LLM summaries lose critical info | Medium | Quality degradation | Declared losses + micro-faults as safety net. Spot-check auditing. | +| Proxy adds too much latency | Low | User experience | Helper LLM calls are async (don't block response). Summarization happens post-response. | +| pgvector search too slow at scale | Low | Micro-fault latency | IVFFlat index. For extreme scale, switch to dedicated vector DB (Qdrant). | +| Object segmentation too noisy | Medium | Poor fidelity decisions | Conservative defaults (larger objects). Rule-based segmentation is robust. | +| Phantom tool parsing from streaming response is fragile | Medium | Proxy breaks | Extensive testing on recorded sessions. Fallback: don't parse, let tool call through. | +| Model doesn't use cooperative eviction (memory_release) | High | Reduced savings | Cooperative eviction is bonus. Pressure-based eviction works without model cooperation. | +| Cross-session memory introduces stale/wrong context | Medium | Wrong answers | Phase 6+ only. Confidence decay on persistent objects. | diff --git a/SCHEMA.md b/SCHEMA.md new file mode 100644 index 0000000..e53e2f2 --- /dev/null +++ b/SCHEMA.md @@ -0,0 +1,443 @@ +# Object Store Schema Design + +## Database: PostgreSQL 16 + pgvector + +--- + +## 1. Core Tables + +### 1.1 `sessions` + +Tracks proxy sessions (one per opencode session). + +```sql +CREATE TABLE sessions ( + id UUID PRIMARY KEY DEFAULT gen_random_uuid(), + external_id TEXT UNIQUE NOT NULL, -- opencode session ID + model TEXT NOT NULL, -- primary model (e.g., claude-opus-4) + created_at TIMESTAMPTZ NOT NULL DEFAULT now(), + last_active_at TIMESTAMPTZ NOT NULL DEFAULT now(), + total_turns INTEGER NOT NULL DEFAULT 0, + total_objects INTEGER NOT NULL DEFAULT 0, + total_faults INTEGER NOT NULL DEFAULT 0, + total_micro_faults INTEGER NOT NULL DEFAULT 0, + status TEXT NOT NULL DEFAULT 'active', -- active, completed, abandoned + config JSONB NOT NULL DEFAULT '{}' -- session-level config overrides +); + +CREATE INDEX idx_sessions_external ON sessions(external_id); +CREATE INDEX idx_sessions_active ON sessions(status) WHERE status = 'active'; +``` + +### 1.2 `semantic_objects` + +The central table. Every piece of context is a semantic object. + +```sql +CREATE EXTENSION IF NOT EXISTS vector; + +CREATE TABLE semantic_objects ( + id UUID PRIMARY KEY DEFAULT gen_random_uuid(), + session_id UUID NOT NULL REFERENCES sessions(id) ON DELETE CASCADE, + + -- Identity + object_type TEXT NOT NULL, -- conversation_phase, design_decision, debugging_session, + -- file_context, tool_result, plan, error_context, + -- external_reference + source_tool TEXT, -- tool that generated this (Read, Bash, Grep, etc.) + source_key TEXT, -- dedup key (e.g., file path for Read results) + + -- Multi-fidelity content + content_full TEXT NOT NULL, -- L0: complete original content + summary_detailed TEXT, -- L1: ~30% of original + summary_compact TEXT, -- L2: ~5% of original + stub TEXT NOT NULL, -- L3: one-line description (always present) + + -- Declared losses (per fidelity level) + losses_l1 JSONB DEFAULT '[]', -- what L1 dropped vs L0 + losses_l2 JSONB DEFAULT '[]', -- what L2 dropped vs L1 + can_answer_l1 JSONB DEFAULT '[]', -- what L1 can answer + can_answer_l2 JSONB DEFAULT '[]', -- what L2 can answer + fault_when JSONB DEFAULT '[]', -- when to fault (hints for model) + + -- Key entities extracted during summarization + key_entities JSONB DEFAULT '[]', -- file paths, function names, etc. + tags TEXT[] DEFAULT '{}', -- freeform tags for filtered queries + + -- State + current_fidelity INTEGER NOT NULL DEFAULT 0, -- 0=L0, 1=L1, 2=L2, 3=L3, 4=evicted + pinned BOOLEAN NOT NULL DEFAULT false, -- fault-driven pin + pin_reason TEXT, -- why pinned (fault hash, anchor tag, etc.) + + -- Metrics + created_at TIMESTAMPTZ NOT NULL DEFAULT now(), + last_accessed TIMESTAMPTZ NOT NULL DEFAULT now(), + access_count INTEGER NOT NULL DEFAULT 0, + fault_count INTEGER NOT NULL DEFAULT 0, -- times model faulted on this object + micro_fault_count INTEGER NOT NULL DEFAULT 0, -- times model queried without restore + + -- Size tracking (for pressure calculations) + tokens_l0 INTEGER NOT NULL DEFAULT 0, -- token count at each level + tokens_l1 INTEGER, + tokens_l2 INTEGER, + tokens_l3 INTEGER NOT NULL DEFAULT 0, + + -- Source message range (which messages this object was segmented from) + source_turn_start INTEGER, -- first user turn index + source_turn_end INTEGER, -- last user turn index + + -- Embedding for semantic search + embedding vector(384) NOT NULL -- all-MiniLM-L6-v2 = 384 dimensions +); + +-- Primary access patterns +CREATE INDEX idx_objects_session ON semantic_objects(session_id); +CREATE INDEX idx_objects_session_fidelity ON semantic_objects(session_id, current_fidelity); +CREATE INDEX idx_objects_session_type ON semantic_objects(session_id, object_type); +CREATE INDEX idx_objects_source_key ON semantic_objects(session_id, source_key) + WHERE source_key IS NOT NULL; +CREATE INDEX idx_objects_created ON semantic_objects(session_id, created_at); +CREATE INDEX idx_objects_last_accessed ON semantic_objects(session_id, last_accessed); +CREATE INDEX idx_objects_tags ON semantic_objects USING GIN(tags); + +-- Vector similarity search (IVFFlat for speed, switch to HNSW at scale) +CREATE INDEX idx_objects_embedding ON semantic_objects + USING ivfflat (embedding vector_cosine_ops) WITH (lists = 100); + +-- Full-text search on content +CREATE INDEX idx_objects_content_fts ON semantic_objects + USING GIN(to_tsvector('english', content_full)); +``` + +### 1.3 `object_relationships` + +Typed edges between semantic objects. + +```sql +CREATE TABLE object_relationships ( + id UUID PRIMARY KEY DEFAULT gen_random_uuid(), + session_id UUID NOT NULL REFERENCES sessions(id) ON DELETE CASCADE, + source_id UUID NOT NULL REFERENCES semantic_objects(id) ON DELETE CASCADE, + target_id UUID NOT NULL REFERENCES semantic_objects(id) ON DELETE CASCADE, + relationship TEXT NOT NULL, -- parent_of, caused_by, references, supersedes, depends_on + metadata JSONB DEFAULT '{}', + created_at TIMESTAMPTZ NOT NULL DEFAULT now(), + + UNIQUE(source_id, target_id, relationship) +); + +CREATE INDEX idx_rels_source ON object_relationships(source_id); +CREATE INDEX idx_rels_target ON object_relationships(target_id); +CREATE INDEX idx_rels_session ON object_relationships(session_id); +``` + +### 1.4 `fault_history` + +Records every fault for pinning decisions and analytics. + +```sql +CREATE TABLE fault_history ( + id UUID PRIMARY KEY DEFAULT gen_random_uuid(), + session_id UUID NOT NULL REFERENCES sessions(id) ON DELETE CASCADE, + object_id UUID NOT NULL REFERENCES semantic_objects(id) ON DELETE CASCADE, + fault_type TEXT NOT NULL, -- full_restore, micro_fault, entropy_triggered + turn_number INTEGER NOT NULL, + content_hash TEXT NOT NULL, -- hash of content at time of eviction + question TEXT, -- for micro_faults: the question asked + answer TEXT, -- for micro_faults: the answer returned + answer_tokens INTEGER, -- tokens in the micro-fault answer + avoided_tokens INTEGER, -- tokens saved vs full restore + latency_ms INTEGER, -- time to handle the fault + created_at TIMESTAMPTZ NOT NULL DEFAULT now() +); + +CREATE INDEX idx_faults_session ON fault_history(session_id); +CREATE INDEX idx_faults_object ON fault_history(object_id); +CREATE INDEX idx_faults_content_hash ON fault_history(content_hash); +``` + +### 1.5 `fidelity_transitions` + +Audit log of every fidelity change (for analytics and debugging). + +```sql +CREATE TABLE fidelity_transitions ( + id UUID PRIMARY KEY DEFAULT gen_random_uuid(), + session_id UUID NOT NULL REFERENCES sessions(id) ON DELETE CASCADE, + object_id UUID NOT NULL REFERENCES semantic_objects(id) ON DELETE CASCADE, + from_fidelity INTEGER NOT NULL, + to_fidelity INTEGER NOT NULL, + trigger TEXT NOT NULL, -- pressure, access, fault, cooperative, goal_change + turn_number INTEGER NOT NULL, + pressure_zone TEXT, -- normal, caution, warning, critical, emergency + token_count INTEGER, -- total tokens at time of transition + created_at TIMESTAMPTZ NOT NULL DEFAULT now() +); + +CREATE INDEX idx_transitions_session ON fidelity_transitions(session_id); +CREATE INDEX idx_transitions_object ON fidelity_transitions(object_id); +``` + +### 1.6 `admission_scores` + +Records admission decisions for tuning the scorer. + +```sql +CREATE TABLE admission_scores ( + id UUID PRIMARY KEY DEFAULT gen_random_uuid(), + session_id UUID NOT NULL REFERENCES sessions(id) ON DELETE CASCADE, + object_id UUID REFERENCES semantic_objects(id) ON DELETE SET NULL, -- NULL if rejected + admitted BOOLEAN NOT NULL, + score_total REAL NOT NULL, + score_type REAL NOT NULL, + score_novelty REAL NOT NULL, + score_utility REAL NOT NULL, + score_recency REAL NOT NULL, + threshold REAL NOT NULL, + content_preview TEXT, -- first 200 chars (for debugging rejected items) + created_at TIMESTAMPTZ NOT NULL DEFAULT now() +); + +CREATE INDEX idx_admission_session ON admission_scores(session_id); +``` + +--- + +## 2. Key Queries + +### 2.1 Context Assembly (every API call) + +```sql +-- Get all objects for this session, ordered by relevance for context assembly +SELECT + id, object_type, current_fidelity, pinned, + CASE current_fidelity + WHEN 0 THEN content_full + WHEN 1 THEN summary_detailed + WHEN 2 THEN summary_compact + WHEN 3 THEN stub + END AS context_content, + losses_l1, losses_l2, can_answer_l1, can_answer_l2, fault_when, + tokens_l0, tokens_l1, tokens_l2, tokens_l3, + key_entities, tags +FROM semantic_objects +WHERE session_id = $1 + AND current_fidelity < 4 -- not fully evicted +ORDER BY + pinned DESC, -- pinned objects first + current_fidelity ASC, -- higher fidelity first + last_accessed DESC -- most recently accessed first +; +``` + +### 2.2 Semantic Search (for micro-faults and goal-aware retrieval) + +```sql +-- Find objects most relevant to a query (including evicted ones) +SELECT + id, object_type, current_fidelity, + content_full, -- always return full content for micro-fault QA + stub, + key_entities, + 1 - (embedding <=> $query_embedding) AS similarity +FROM semantic_objects +WHERE session_id = $1 +ORDER BY embedding <=> $query_embedding +LIMIT $2; +``` + +### 2.3 Hybrid Search (vector + full-text) + +```sql +-- Combine semantic similarity with keyword matching +WITH vector_results AS ( + SELECT id, 1 - (embedding <=> $query_embedding) AS vec_score + FROM semantic_objects + WHERE session_id = $1 + ORDER BY embedding <=> $query_embedding + LIMIT 20 +), +text_results AS ( + SELECT id, ts_rank(to_tsvector('english', content_full), + plainto_tsquery('english', $query_text)) AS text_score + FROM semantic_objects + WHERE session_id = $1 + AND to_tsvector('english', content_full) @@ plainto_tsquery('english', $query_text) + LIMIT 20 +) +SELECT + COALESCE(v.id, t.id) AS id, + COALESCE(v.vec_score, 0) * 0.7 + COALESCE(t.text_score, 0) * 0.3 AS combined_score +FROM vector_results v +FULL OUTER JOIN text_results t ON v.id = t.id +ORDER BY combined_score DESC +LIMIT $2; +``` + +### 2.4 Pressure Calculation + +```sql +-- Calculate current token pressure for a session +SELECT + SUM(CASE current_fidelity + WHEN 0 THEN tokens_l0 + WHEN 1 THEN COALESCE(tokens_l1, 0) + WHEN 2 THEN COALESCE(tokens_l2, 0) + WHEN 3 THEN tokens_l3 + ELSE 0 + END) AS total_context_tokens, + COUNT(*) FILTER (WHERE current_fidelity = 0) AS objects_at_l0, + COUNT(*) FILTER (WHERE current_fidelity = 1) AS objects_at_l1, + COUNT(*) FILTER (WHERE current_fidelity = 2) AS objects_at_l2, + COUNT(*) FILTER (WHERE current_fidelity = 3) AS objects_at_l3, + COUNT(*) FILTER (WHERE current_fidelity = 4) AS objects_evicted, + COUNT(*) FILTER (WHERE pinned) AS objects_pinned +FROM semantic_objects +WHERE session_id = $1; +``` + +### 2.5 Deduplication Check (for file re-reads) + +```sql +-- Check if this file was already read (supersedes pattern) +SELECT id, content_full, current_fidelity +FROM semantic_objects +WHERE session_id = $1 + AND source_key = $2 -- e.g., file path + AND object_type = 'file_context' +ORDER BY created_at DESC +LIMIT 1; +``` + +### 2.6 Related Objects (for co-fidelity management) + +```sql +-- When upgrading an object, find related objects that should also upgrade +WITH RECURSIVE related AS ( + SELECT target_id AS id, relationship, 1 AS depth + FROM object_relationships + WHERE source_id = $1 + AND relationship IN ('depends_on', 'references', 'caused_by') + + UNION ALL + + SELECT r.target_id, r.relationship, rel.depth + 1 + FROM object_relationships r + JOIN related rel ON r.source_id = rel.id + WHERE rel.depth < 2 -- max 2 hops +) +SELECT DISTINCT so.* +FROM related r +JOIN semantic_objects so ON so.id = r.id +WHERE so.current_fidelity > 1; -- only objects that could benefit from upgrade +``` + +--- + +## 3. Data Flow Examples + +### 3.1 Model Reads a File + +``` +1. Model calls Read(src/auth/middleware.ts) +2. Proxy intercepts response +3. Segmenter creates semantic_object: + - object_type: 'file_context' + - source_tool: 'Read' + - source_key: 'src/auth/middleware.ts' + - content_full: (file contents) + - stub: "Read src/auth/middleware.ts (200 lines, auth middleware)" + - embedding: embed(content_full) + - tokens_l0: count_tokens(content_full) +4. Admission scorer: S(m) = 0.72 (above threshold) -> admitted +5. Dedup check: no existing object with source_key='src/auth/middleware.ts' -> INSERT + (If exists: create new object, add 'supersedes' relationship to old one, + degrade old one to L3) +``` + +### 3.2 Pressure Triggers Fidelity Degradation + +``` +1. Pressure Monitor: total_context_tokens = 78,000 (78% of 100K budget) + -> Zone: WARNING +2. Find oldest L0 objects not accessed in last 3 turns: + SELECT id FROM semantic_objects + WHERE session_id = $1 AND current_fidelity = 0 + AND last_accessed < (now() - interval '3 turns') + ORDER BY last_accessed ASC; +3. For each candidate: + a. Call Helper LLM to generate L1 summary + declared losses + b. UPDATE semantic_objects SET + summary_detailed = $summary, + losses_l1 = $losses, + can_answer_l1 = $can_answer, + current_fidelity = 1, + tokens_l1 = count_tokens($summary) + WHERE id = $candidate_id; + c. INSERT INTO fidelity_transitions (trigger='pressure', ...) +4. Recalculate pressure. If still in WARNING, degrade L1 -> L2. +``` + +### 3.3 Micro-Fault + +``` +1. Model generates: memory_query("What error code for expired tokens?") +2. Proxy intercepts phantom tool call +3. Proxy queries backing store: + SELECT id, content_full + FROM semantic_objects + WHERE session_id = $1 + ORDER BY embedding <=> embed("error code expired tokens") + LIMIT 3; +4. Proxy sends to Helper LLM: + "Answer this question using ONLY the provided context: + Q: What error code does auth middleware return for expired tokens? + Context: {top-3 objects' full content}" +5. Helper returns: "The auth middleware returns HTTP 401 with error code + 'TOKEN_EXPIRED' and body { error: 'token_expired', message: '...' }" +6. Proxy injects as synthetic tool result (NOT the full object content) +7. INSERT INTO fault_history (fault_type='micro_fault', + answer_tokens=45, avoided_tokens=3200, ...) +8. Object stays at current fidelity (no upgrade) +``` + +--- + +## 4. Migration Path + +### Phase 1 (no DB needed) +Pichay's proxy uses in-memory state + client's message history as backing store. +No PostgreSQL required. + +### Phase 2 (SQLite prototype) +Add SQLite with sqlite-vec for local development: +- Single file, no server +- Same schema, adapted types (TEXT instead of vector, custom cosine function) + +### Phase 3+ (PostgreSQL) +Full schema as defined above. Migration from SQLite via: +```sql +-- Export from SQLite, import to PostgreSQL +-- pgloader or custom Python migration script +``` + +### Future: Cross-Session Memory (L5) +```sql +-- Additional table for cross-session persistent objects +CREATE TABLE persistent_objects ( + id UUID PRIMARY KEY DEFAULT gen_random_uuid(), + user_id TEXT NOT NULL, -- across all sessions for this user + source_object_id UUID, -- which session object it came from + source_session_id UUID, + object_type TEXT NOT NULL, + content TEXT NOT NULL, -- curated persistent version + embedding vector(384) NOT NULL, + confidence REAL NOT NULL DEFAULT 1.0, -- decays if not reinforced + last_reinforced TIMESTAMPTZ NOT NULL DEFAULT now(), + created_at TIMESTAMPTZ NOT NULL DEFAULT now() +); + +CREATE INDEX idx_persistent_user ON persistent_objects(user_id); +CREATE INDEX idx_persistent_embedding ON persistent_objects + USING ivfflat (embedding vector_cosine_ops) WITH (lists = 50); +```