OptMem/README.md
victortaelin a6abc264c5 the wake budget is 208 lines: 16k tokens measured, not 16k assumed
WAKE_LINES=256 was converted from Taelin's 16k-token budget at 4 bytes per
token. Dense technical memories run 3.36, so the real document was 19,015
tokens on the live store -- 19% over the number he set, and it needed a
fourth tool call to deliver a 17-line stub.

  256 lines  63,922 bytes  19,015 tokens  4 parts
  208 lines  54,011 bytes  16,393 tokens  3 parts   <- shipped

Nothing is recomputed: the budget only selects which existing lines print.
test.py now reads the budget from memo instead of keeping its own copy.
2026-07-25 18:49:37 -03:00

284 lines
11 KiB
Markdown

# OptMem
A permanent memory for AI agents. One machine holds one identity that survives
every new session, every compaction, and every change of model or vendor.
It is a handful of append-only text files and six commands. No daemon, no database, no
API, no integration with any particular agent harness — it works the same under
Claude Code, Codex, pi or a human at a shell.
## The problem
An agent's context window is its whole world, and the world ends every session.
The usual patch is a notes file the agent rewrites by hand, which decays into
either a stale summary or a wall of text nobody can afford to read.
OptMem fixes the two halves separately:
- **Nothing is ever forgotten.** Every memory is appended to `LOG.txt` and
never edited or deleted. That file is the truth, forever.
- **What you read is a fixed size.** `memo wake` prints a document of bounded
length — recent memories verbatim, older ones progressively compressed. At
a hundred million memories it is still the same number of lines.
## Install
```sh
git clone https://github.com/VictorTaelin/OptMem ~/OptMem
export PATH="$HOME/OptMem:$PATH"
export MEMORY_DIR="$HOME/memory" # required; there is no default
mkdir -p "$MEMORY_DIR" # this is what creates the identity
```
`MEMORY_DIR` is the only machine-specific fact in the system. One machine, one
`MEMORY_DIR`, one identity. `memo` never creates that directory itself: if it
did, one typo would open a second, empty identity instead of an error.
## Use
```sh
memo wake # read your memory. run this first, every session,
# then the command each part orders, until one
# prints `You are awake.`
memo note "..." # record a memory. one line, <= 280 chars.
memo sleep # compress. answer each prompt until it prints
# `Nothing left to compress.`
memo recall <regex> # search the raw log for detail a summary lost.
memo forget <lo>-<hi> # drop a wrong summary; the next sleep redoes it.
```
```
$ memo note "OptMem: LOG.txt is the truth, TREE/ is the cache, wake reads both"
Saved as #4213.
Compress memories #4212-4213 into one line of at most 280 characters.
Keep every name, number, date, decision and outcome.
Drop wording, not facts. Invent nothing.
#4212 2026-07-25 minilin fleet renamed from bip; one mini = one identity
#4213 2026-07-25 OptMem: LOG.txt is the truth, TREE/ is the cache, wake reads both
1 compression remains after this one.
Run: memo sleep 4212-4213 "<your line>"
```
## How it works
`LOG.txt` is the ground truth: one memory per line, forever.
```
#4211 2026-07-25 taelin: memory must be append-only, one line per entry
#4212 2026-07-25 minilin fleet renamed from bip; one mini = one identity
#4213 2026-07-25 OptMem: LOG.txt is the truth, TREE/ is the cache
```
`TREE/` is a cache of summaries, one file per block size. A **block** is an
aligned power-of-two range of memories compressed into a single line, and a
block is built from its two halves — so the blocks form a binary merge tree
over the log (a block is named by the inclusive range it covers, `0-1` being
memories #0 and #1):
```
#0 #1 #2 #3 #4 #5 #6 #7 the raw memories
\ / \ / \ / \ /
0-1 2-3 4-5 6-7 each one line, <= 280 chars
\ / \ /
0-3 4-7
\ /
0-7
```
A block covering four thousand memories is still one line of 280 characters.
Nothing in the system is ever bigger than one line.
`memo wake` picks a set of blocks that tiles the whole log and prints them. It
keeps a block whole when its size is small relative to its age, so **detail is
proportional to recency**, and it spends exactly `WAKE_LINES` lines doing it:
```
10,000 memories, WAKE_LINES = 208:
block size: 1 2 4 8 16 32 64 128 256
how many: 42 21 21 21 22 21 21 21 18
└ the last 42, verbatim ───────────▶ the first 4,600, 256:1
```
The oldest memories are recalled as a vague shape, the newest word for word,
and the transition is smooth. Below `WAKE_LINES` memories nothing is compressed
at all — your whole life is printed verbatim, because it fits.
## The invariant
**There is never any doable work pending.** The moment a block's range is
complete, that block can be built, and it must be. This costs about one small
compression per memory written, and it means:
- `memo wake` never waits. The blocks it needs were built long ago.
- Work is never deferred into a spike. Measured over 20,000 memories, a new
memory creates one compression on average and nine at the very worst.
- `WAKE_LINES` can be changed at any time, on any machine, with nothing to
recompute. It only selects which existing lines get printed.
`memo wake` enforces the invariant: while any compression is pending it refuses
to print, and hands you the work instead. A memory with work left in it is not
yet the truth.
## Writing a good memory
Write one the moment something happens, you learn something, or something
changes — if and only if it is new to you, important, and lasting in effect: a
task worth real effort, a fact or insight the user teaches you, anything you
learn about their life (even indirectly), work of yours that lands. Do not log
trivia, do not narrate your own process, and never write what you already
know: a redundant memory costs a compression and buys nothing.
Compress toward facts, not prose. Keep names, numbers, dates, paths, ids and
decisions; drop wording.
```
bad worked on the memory system today and made good progress on the design
good OptMem design settled: LOG.txt append-only truth, TREE binary merge
tree of 280-char summaries, wake renders a fixed 208-line document
```
## Output is delivered in parts
Every harness truncates an over-long command, and each one drops a different
piece:
```
Claude Code 30,000 chars drops the MIDDLE
pi 50 KB / 2000 lines drops the HEAD
Codex 10,000 tokens (configurable per call)
```
A 208-line memory is ~56 KB, so a single-shot `memo wake` is mangled
everywhere, and silently.
So `memo wake` pages the document into parts that fit all of them
(`PART_CHARS`, `PART_LINES`), and each part ends by ordering the exact command
for the next one, including the `T` it was rendered at — so a memory written
mid-wake cannot shift a boundary and drop a line. Nothing is special-cased per
harness: if yours is more generous, raise the two settings for fewer parts.
## Files
```
$MEMORY_DIR/
LOG.txt #id date text append-only. never edited. the truth.
TREE/2 one summary per a cache of block summaries, one file per block
TREE/4 record, indexed size. each block written once, unless forgotten.
TREE/8 by position
...
config optional. absent on a normal store; the defaults below live in
`memo` and are the only home for them.
ENTRY_CHARS=280 longest a memory may be
WAKE_LINES=208 how many lines `memo wake` prints (~16k tokens)
PART_CHARS=20000 how much of it fits in one command's output
PART_LINES=500 ...and in how many lines
```
**Records are fixed width**: 320 bytes in `LOG.txt`, 288 in the `TREE` files.
That is the whole indexing strategy — position *is* identity, so memory `i`
sits at `i*320`, and block `[k*s, (k+1)*s)` sits at `k*288` of `TREE/s`.
Everything is one seek: no scanning, and no index file that could ever
disagree with the data.
```
1,000,000 memories, 607 MB on disk:
memo wake 0.03s (scanning the same store: 0.96s)
memo note 0.02s (scanning: 1.30s)
memo sleep 0.02s
```
Finding pending work costs one `stat` per level — about twenty, forever —
because each level file holds a dense prefix, so its length says exactly how
far that level got. Padding costs ~1.6x on disk and buys O(1) on everything.
Both files are still plain text: `grep`, `cat` and `wc -l` all work, lines are
just space-padded. Writes are serialised with a lock, so parallel sessions on
one machine can append at the same time without corrupting anything.
Agents must never create, edit or delete anything in `MEMORY_DIR` themselves.
Every write goes through `memo`, which enforces the one-line and character
limits, assigns ids, and refuses to overwrite a block that already exists.
## Correcting a memory
You cannot. Append the correction instead:
```
memo note "correction: the halt bug was in the column order, not the row order (see #4198)"
```
Both lines are true history, and compression will merge them. This is why
nothing is ever lost: `memo recall` still finds the original.
A *summary* is different. It is not history, it is a cache of a pure function
of the log, and it can simply be wrong — mistyped, or badly compressed. Drop
it and everything built on top of it:
```
$ memo forget 188-191
Forgot 20 summaries, from 188-191 up. Run: memo sleep
```
`LOG.txt` is never touched, so fixing a bad summary can never cost you a
memory. Blocks are built in order, so forgetting one also drops the blocks
built after it at the same levels; they come back on the next sleep.
## Add this to your agent's instruction file
Put it at the top of `AGENTS.md` (or `CLAUDE.md`), above everything else,
adjusting the tool path:
```markdown
## Memory
Your memory is OptMem: the tool is `~/OptMem`, the data is `$MEMORY_DIR`.
It survives every new session, every compaction and every change of model
or vendor. Without it you do not know who you are, or what was already
decided and tried.
Run `memo wake` before any other tool call, in every session. It prints in
numbered parts, each ordering the next; run every one until a part says
`You are awake.` Do not stop early: part 1 is your distant past, the last
part is this week. If wake refuses because compressions are pending, do
them and run `memo wake` again.
While you work:
- `memo note "<one line, max 280 chars>"` the moment something happens, you
learn something, or something changes -- if and only if it is new to you,
important, and lasting in effect. That covers a task worth real effort, a
fact or insight the user teaches you, anything you learn about their life
(even indirectly), and work of yours that lands. Never write what you
already know: no redundant memories, ever.
- If `memo note` returns a compression, do it before your next action.
- `memo recall <regex>` when a memory is too vague.
- Before your context ends, run `memo sleep` and answer each prompt until
it prints `Nothing left to compress.`
- Never create, edit or delete anything under `$MEMORY_DIR`. Only `memo`
writes.
Parallel sessions on this machine are all you, and may all write memories.
A subagent is not: it must never run `memo`, because it cannot judge what
is already known and its notes would arrive duplicated and at the wrong
grain. Start every brief you send one with `You are a subagent. Do not run
memo.` If your own first message is a task brief from another agent, you
are that subagent: skip this section.
```
## Test
```sh
python3 test.py
```
Runs the block math against a hundred thousand memory counts and drives the
real CLI through a synthetic life of two thousand memories, checking that the
document always tiles the log, never exceeds its budget, always increases in
detail toward the present, that every block is written exactly once, that
nothing is ever rewritten, and that a full sleep always leads to a clean wake.