Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions cmd/odek/introspect.go
Original file line number Diff line number Diff line change
Expand Up @@ -72,6 +72,7 @@ func buildConfigView(resolved config.ResolvedConfig) map[string]any {
"facts_limit_env": resolved.Memory.FactsLimitEnv,
"extract_on_end": boolPtr(resolved.Memory.ExtractOnEnd),
"consolidate_on_end": boolPtr(resolved.Memory.ConsolidateOnEnd),
"consolidate_at_cap_pct": resolved.Memory.ConsolidateAtCapPct,
"min_turns_for_extraction": resolved.Memory.MinTurnsForExtraction,
},
"skills": map[string]any{
Expand Down
2 changes: 2 additions & 0 deletions docs/CONFIG.md
Original file line number Diff line number Diff line change
Expand Up @@ -520,6 +520,7 @@ The `memory` section controls the persistent memory system (see [docs/MEMORY.md]
"add_threshold": 0.3,
"auto_approve_episodes": false,
"episode_dedup_threshold": 0.92,
"consolidate_at_cap_pct": 80,
"max_episodes": 500,
"episode_ttl_days": 0,
"embedding": {
Expand All @@ -543,6 +544,7 @@ The `memory` section controls the persistent memory system (see [docs/MEMORY.md]
| `buffer_enabled` | true | Enable the turn-level buffer |
| `merge_on_write` | true | Use go-vector RP similarity to auto-merge related entries (fast, no LLM — uses simple string merge) |
| `consolidate_on_end` | true | At session end, run an LLM consolidation pass over `user.md` and `env.md` in a background goroutine. This is the quality complement to `merge_on_write`: merge-on-write handles obvious duplicates immediately (no LLM), while consolidation handles near-duplicates and paraphrases at session end with full LLM quality. Requires `llm_consolidate: true`. **Note:** facts in the borderline similarity band (0.3–0.7 cosine) are now always added immediately and only merged by this consolidation pass — if you set `consolidate_on_end: false`, near-duplicate facts will accumulate rather than being merged. |
| `consolidate_at_cap_pct` | 80 | Percentage of `facts_limit_user`/`facts_limit_env` above which a background consolidation pass fires for that fact file on `AddFact` — long-lived serve/REPL sessions never reach the session-end trigger, so entries otherwise fossilize near the cap. Snapshot/preview is taken flock-free; the flock is held only for the verified apply swap, never across the LLM call. A per-target in-flight guard prevents stacked duplicate LLM calls. Explicit `0` disables the cap trigger. Best-effort: errors are logged, never surfaced to callers. Requires `llm_consolidate: true`. |
| `extract_on_end` | true | At session end (≥3 turns), extract a narrative episode summary via LLM for later recall |
| `min_turns_for_extraction` | 3 | Minimum conversation turns before end-of-session extraction runs |
| `extract_facts` | **false** | **Opt-in.** At session end (≥3 turns), auto-extract a few **durable** facts (stable user preferences, project invariants) into `user.md`/`env.md`. Off by default — see the security note below. Independent of `extract_on_end`; to disable *all* end-of-session LLM extraction set `llm_extract: false`. |
Expand Down
5 changes: 3 additions & 2 deletions docs/MEMORY.md
Original file line number Diff line number Diff line change
Expand Up @@ -79,7 +79,7 @@ Episode extraction runs **asynchronously** — it does not block the agent loop.

## Automatic Cap Maintenance (LLM-driven eviction)

The agent maintains its own fact files. When a fact file fills up, the **agent itself** evicts older entries — there is no background rewriter and no silent consolidation. Every eviction is an explicit, auditable `remove`/`replace` call inside the normal agent loop.
The agent maintains its own fact files. When a fact file fills up, the **agent itself** evicts older entries — every eviction is an explicit, auditable `remove`/`replace` call inside the normal agent loop. The one automatic quality pass is cap-triggered consolidation (below): it merges redundant entries via LLM but never silently deletes — the applied snapshot is re-scanned and verified against the file state at snapshot time.

Three affordances make this work:

Expand All @@ -103,7 +103,7 @@ Additionally, the system-prompt memory block appends a one-line warning when a f
⚠ env fact file 97% full — evict stale entries via memory remove before your next add.
```

Caps are configured via `facts_limit_user` / `facts_limit_env` (defaults: 4,000 / 8,000 chars, counted as bytes — consistent with cap accounting). Deliberately **not** built: automatic consolidation on write (a mid-flow LLM rewrite risks dropping load-bearing details and amplifies provider throttling) and date-based auto-eviction (age alone says nothing about value — a days-old pointer to untracked work can be the only durable record of it).
Caps are configured via `facts_limit_user` / `facts_limit_env` (defaults: 4,000 / 8,000 chars, counted as bytes — consistent with cap accounting). When a fact file crosses `consolidate_at_cap_pct` (default 80%, `0` disables) of its cap, a background LLM consolidation pass merges redundant entries — the merge output is re-scanned and applied only if the file has not changed since the snapshot, so concurrent writes conflict instead of being overwritten. Deliberately **not** built: date-based auto-eviction (age alone says nothing about value — a days-old pointer to untracked work can be the only durable record of it).

## Merge-on-Write (go-vector Integration)

Expand Down Expand Up @@ -217,6 +217,7 @@ Key properties:
- **Size cap**: defaults to 100 MB with `retention_decay` eviction; pinned atoms are never evicted.
- **Tool surface**: `memory` tool actions `add_atom`, `search_atoms`, `forget_atom`, `pin_atom`, `list_quarantine`, `confirm_pending_review`, `reject_pending_review`, and `list_pending_review`.
- **CLI surface**: `odek memory extended forget|promote|pin|quarantine|compact|stats|consolidate|nudges|pending|confirm|reject`.
- **Pending user-model corrections**: unconfirmed `pending_review` entries age out after `memory.extended.user_state_pending_max_age_days` (default 14; explicit `0` disables expiry). Only unconfirmed pending entries are pruned — confirmed facts are never touched.

**Proactive nudges** (opt-in): when `memory.extended.proactive_nudges_enabled` is `true` (default `false`), Extended Memory can synthesize short, user-facing nudges from trusted atoms — open questions, stale goals, blockers, and drift. Delivery is capped by `nudge_max_per_day` with a per-kind cooldown (`nudge_cooldown_hours`); goals only become "stale" after `nudge_stale_goal_days`. The Telegram bot pushes at most one nudge after a completed turn (in the background, prefixed with 💡). `odek memory extended nudges` prints a preview of up to 2 nudges without consuming the daily cap.

Expand Down
15 changes: 15 additions & 0 deletions docs/PLANNING.md
Original file line number Diff line number Diff line change
Expand Up @@ -529,6 +529,21 @@ semantics, payload minimality, `ExtractPlan`), `cmd/odek/serve_plan_test.go`
fingerprint classified `local_write` or higher (or whose last outcome was
denied/blocked) escalates to “stop retrying that class” and names the next
pending step as `next_non_mutating`. Still a hint — never an auto-exec.
Hints **escalate** instead of resetting: the interval between fires doubles
(3 → 6 → 12 → … identical calls per fingerprint), and repeat fires are
marked with an `again` suffix/detail so the model can tell a fresh hint
from an escalated repeat (`tool_recovery` detail reads `repeated identical
call (Nx<tool>, again)`).
- **Bounded background-job polling.** `bg_status` / `bg_output` fingerprints
get a 3× raised stall threshold — 9 identical polls with no state change —
before the same corrective hint fires with a poll-specific message (the job
may be stuck or the id dead). Legitimate polling with changing results
never trips it.
- **Remaining-steps prefix on side calls.** The compaction-digest and
progress-summary side-call payloads are prepended with the remaining plan
steps **including titles** (`Remaining plan steps: …`), not bare ids —
these payloads may be the only surviving context after a trim, so ids alone
tell the summarizer nothing. Stall hints keep the title-free id format.
- **Blocked-step streak.** Three consecutive transitions to `blocked` inject
a decompose-or-`create` hint and emit `plan_blocked` (`steps`, `blocked`,
`version` only). The streak resets on `create` or a `done` / `in_progress`
Expand Down
1 change: 1 addition & 0 deletions docs/SECURITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -195,6 +195,7 @@ When a classification is set to `prompt`, an approver pauses the agent until the
- **Friction mode** engages after 3 approvals of the same class in 60 s. On TTY **and the bundled Web UI** the next prompt requires typing the literal word `approve` (no single-letter shortcut) and a 1.5 s pause before accepting input. Telegram hides the Trust shortcut and warns; a button `approve` still works (no typed word, no pause). REST typed `confirm` is opt-in (`dangerous.rest_approval_friction`).
- TTY prompts are serialized process-wide (one mutex, one shared approval log), so concurrent tool calls cannot print overlapping prompts, and the friction counter and trust cache persist across prompts and across `shell`/`parallel_shell` tool instances. A cancelled turn context closes the TTY so `ReadString` cannot wedge the process after Ctrl-C.
- **Non-interactive defaults to read-only.** When no TTY is available (headless/CI/piped input), prompted operations fall back to the `non_interactive` action, whose built-in default is `"read_only"`: read-only inspection proceeds — `safe`-classified shell commands (`ls`, `cat`, `tree`) and native read tools over ordinary paths — while writes, execution, egress, and reads of sensitive locations (anything at `system_write` or above) are denied. `"deny"` (block everything prompted, including reads) and `"allow"` remain available; an explicitly configured *invalid* value fails closed to `"deny"` with a load-time warning. The read_only default exists because containment via inability is not safe-and-useful: a headless agent that cannot even `ls` gets its operator to flip `non_interactive` to `allow`, which removes every protection — `read_only` is the setting that survives contact with a deadline.
- **Test binaries fail closed.** Inside a `go test` binary, the TTYApprover never opens the real controlling terminal — a test process without an explicit fixture TTY path (`/dev/tty` or empty) is denied outright instead of silently approving, covering both the test-binary case and the zero-value approver whose legacy path was fail-open. Approvers with an explicit fixture TTYPath are unaffected.

**Batch approval card.** `classifyToolCall` (in the loop) classifies every command inside `parallel_shell`, every path inside `batch_patch`, and the `browser` tool (action + URL → `network_egress`); MCP tools (detected by the `<server>__<tool>` naming convention) classify as `unknown`. The card shows full command/path text instead of truncating, and blanket `SetTrustAll` is refused for any iteration that still contains an unclassifiable tool — those must pass their own internal gates. Session-trusted risk classes are honored uniformly across `write_file`, `patch`, and `batch_patch`.

Expand Down
8 changes: 6 additions & 2 deletions docs/SUBAGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -124,7 +124,9 @@ The `delegate_tasks` tool is available in CLI, REPL, Web UI, Telegram, and headl
```

The `summary` is the headline channel: the child's final answer, capped at
2048 runes. When `summary_truncated` is set, the parent-side render appends
2048 runes. Truncation is **tail-keep** — the END of the answer is preserved
(verdicts, next actions, key decisions), and a leading `…` marks the cut.
When `summary_truncated` is set, the parent-side render appends
`headline truncated (2048 of N runes shown) — fetch artifacts via
artifact_read or re-run with a narrower goal` (the `artifact_read` half is
omitted in processes that do not have the tool, i.e. mid-tree parents).
Expand Down Expand Up @@ -502,7 +504,9 @@ Parent synthesizes: "Created 3 files:

## Result artifacts

The result contract is two-channel: the headline (≤ 2048 runes) carries status;
The result contract is two-channel: the headline (≤ 2048 runes, tail-kept —
the END of the child answer survives so verdicts and next actions reach the
parent; a leading `…` marks the cut) carries status;
the bulk rides files. The child is told this at request time: deliverables
larger than a headline go as FLAT files into its per-task staging directory
(`.odek-artifacts/<task_id>/` inside the workspace — nested directories are
Expand Down
3 changes: 3 additions & 0 deletions internal/config/loader.go
Original file line number Diff line number Diff line change
Expand Up @@ -2873,6 +2873,9 @@ func resolveMemory(cfg *memory.MemoryConfig) memory.MemoryConfig {
if cfg.ConsolidateOnEnd != nil {
def.ConsolidateOnEnd = cfg.ConsolidateOnEnd
}
if cfg.ConsolidateAtCapPct != nil {
def.ConsolidateAtCapPct = cfg.ConsolidateAtCapPct
}
if cfg.LLMSearch != nil {
def.LLMSearch = cfg.LLMSearch
}
Expand Down
188 changes: 188 additions & 0 deletions internal/memory/cap_consolidation_test.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,188 @@
package memory

import (
"context"
"errors"
"strings"
"sync"
"sync/atomic"
"testing"
"time"
)

// ── Cap-aware auto-consolidation (B4 step 2) ────────────────────────────
//
// Session-end consolidation never fires for long-lived serve/REPL
// sessions, so facts fossilize near the cap. Crossing a configurable
// percentage of the cap on AddFact must fire ONE background consolidation
// per crossing (in-flight guard), and 0 must disable the trigger.

// blockingConsolidLLM lets tests hold a consolidation in flight.
type blockingConsolidLLM struct {
calls atomic.Int64
release chan struct{}
closeOne sync.Once
mock *mockLLM
}

func (b *blockingConsolidLLM) Unblock() { b.closeOne.Do(func() { close(b.release) }) }

func (b *blockingConsolidLLM) SimpleCall(ctx context.Context, system, user string) (string, error) {
b.calls.Add(1)
// Respond only for the consolidation prompt shape; extraction prompts
// (if any) proxy to the mock.
if strings.Contains(user, "Consolidate the following memory entries") {
<-b.release
return `["merged fact"]`, nil
}
if b.mock != nil {
return b.mock.SimpleCall(ctx, system, user)
}
return `["x"]`, nil
}

func capFillingFacts(prefix string, n, size int) []string {
out := make([]string, 0, n)
for i := 0; i < n; i++ {
out = append(out, prefix+" "+strings.Repeat("x", size)+time.Now().Format("150405.000000000"))
}
return out
}

func TestAddFact_CapCrossingFiresConsolidation(t *testing.T) {
dir := t.TempDir()
llm := &blockingConsolidLLM{release: make(chan struct{})}
cfg := DefaultMemoryConfig()
cfg.MergeOnWrite = boolPtr(false)
cfg.LLMConsolidate = boolPtr(true)
cfg.ConsolidateOnEnd = boolPtr(false) // isolate the cap trigger
pct := 50 // tiny dir caps → cross easily
cfg.ConsolidateAtCapPct = &pct
mm := NewMemoryManager(dir, llm, cfg)
t.Cleanup(func() {
llm.Unblock() // unblock any late-fired consolidation first
mm.WaitForBackground(30 * time.Second)
})

// Fill past the trigger (env cap 8000; 50% = 4000): 6 × ~830 chars ≈ 5000.
for _, f := range capFillingFacts("unique fact alpha", 6, 800) {
if err := mm.AddFact("env", f); err != nil {
t.Fatalf("AddFact: %v", err)
}
}

deadline := time.Now().Add(5 * time.Second)
for time.Now().Before(deadline) {
if llm.calls.Load() > 0 {
break
}
time.Sleep(5 * time.Millisecond)
}
if llm.calls.Load() == 0 {
t.Fatal("crossing the cap threshold never fired a background consolidation")
}
llm.Unblock()
mm.WaitForBackground(30 * time.Second)
}

func TestAddFact_CapTriggerDisabledAtZero(t *testing.T) {
dir := t.TempDir()
llm := &blockingConsolidLLM{release: make(chan struct{})}
cfg := DefaultMemoryConfig()
cfg.MergeOnWrite = boolPtr(false)
cfg.LLMConsolidate = boolPtr(true)
cfg.ConsolidateOnEnd = boolPtr(false)
pct := 0 // disabled
cfg.ConsolidateAtCapPct = &pct
mm := NewMemoryManager(dir, llm, cfg)
t.Cleanup(func() { mm.WaitForBackground(30 * time.Second) })

for _, f := range capFillingFacts("unique fact beta", 6, 800) {
if err := mm.AddFact("env", f); err != nil {
t.Fatalf("AddFact: %v", err)
}
}
time.Sleep(100 * time.Millisecond)
if got := llm.calls.Load(); got != 0 {
t.Fatalf("disabled trigger fired %d consolidation(s), want 0", got)
}
}

func TestAddFact_InFlightGuardSingleFire(t *testing.T) {
dir := t.TempDir()
llm := &blockingConsolidLLM{release: make(chan struct{})}
cfg := DefaultMemoryConfig()
cfg.MergeOnWrite = boolPtr(false)
cfg.LLMConsolidate = boolPtr(true)
cfg.ConsolidateOnEnd = boolPtr(false)
pct := 50
cfg.ConsolidateAtCapPct = &pct
mm := NewMemoryManager(dir, llm, cfg)
t.Cleanup(func() {
llm.Unblock()
mm.WaitForBackground(30 * time.Second)
})

// Cross the threshold (env cap 8000; 50% = 4000) and start an in-flight
// consolidation. Two entries: Consolidate no-ops on single-entry dirs.
if err := mm.AddFact("env", "unique fact gamma "+strings.Repeat("x", 2400)); err != nil {
t.Fatalf("seed add: %v", err)
}
if err := mm.AddFact("env", "unique fact zeta "+strings.Repeat("y", 2400)); err != nil {
t.Fatalf("seed add 2: %v", err)
}
deadline := time.Now().Add(5 * time.Second)
for time.Now().Before(deadline) && llm.calls.Load() == 0 {
time.Sleep(5 * time.Millisecond)
}
if llm.calls.Load() == 0 {
t.Fatal("consolidation did not start (release channel would deadlock test)")
}

// More adds while the consolidation is in flight must NOT stack calls.
for _, f := range capFillingFacts("unique fact delta", 3, 600) {
if err := mm.AddFact("env", f); err != nil {
t.Fatalf("AddFact: %v", err)
}
}
if got := llm.calls.Load(); got != 1 {
t.Fatalf("consolidation calls while in flight = %d, want 1 (in-flight guard missing)", got)
}
llm.Unblock()
mm.WaitForBackground(30 * time.Second)
}

// errLLM always fails — the cap trigger must be best-effort: AddFact still
// succeeds and errors never surface to the caller.
type errLLM struct{ calls atomic.Int64 }

func (e *errLLM) SimpleCall(ctx context.Context, system, user string) (string, error) {
e.calls.Add(1)
return "", errors.New("llm down")
}

func TestAddFact_CapTriggerBestEffort(t *testing.T) {
dir := t.TempDir()
llm := &errLLM{}
cfg := DefaultMemoryConfig()
cfg.MergeOnWrite = boolPtr(false)
cfg.LLMConsolidate = boolPtr(true)
cfg.ConsolidateOnEnd = boolPtr(false)
pct := 50
cfg.ConsolidateAtCapPct = &pct
mm := NewMemoryManager(dir, llm, cfg)
t.Cleanup(func() { mm.WaitForBackground(30 * time.Second) })

for _, f := range capFillingFacts("unique fact epsilon", 6, 800) {
if err := mm.AddFact("env", f); err != nil {
t.Fatalf("AddFact must not surface consolidation errors: %v", err)
}
}
deadline := time.Now().Add(5 * time.Second)
for time.Now().Before(deadline) && llm.calls.Load() == 0 {
time.Sleep(5 * time.Millisecond)
}
if llm.calls.Load() == 0 {
t.Fatal("expected the trigger to attempt consolidation despite LLM failure")
}
}
Loading
Loading