Skip to content

Repository files navigation

Memory Wiki

A service that ingests conversation transcripts, uses an LLM to distil durable memories from them, stores those memories as files in a navigable, object-storage-backed file tree, and exposes them through unix-style REST endpoints (ls, cat, grep).

POST /transcripts ──► Postgres (transcript + job, one TX)
                          │
                          ▼  (background worker, at-least-once)
                   LLM memory generation
                          │
                          ▼
              Object storage file tree (S3 / MinIO)
                 tree/<category>/<slug>.md
                          │
        ┌─────────────────┼─────────────────┐
        ▼                 ▼                 ▼
 GET /memories/ls   GET /memories/cat  GET /memories/grep

Quick start (single command)

Requires Docker.

docker compose up --build

This builds the app image and starts postgres, minio (+ bucket init), a one-shot migrate step, the api (port 3000), and the worker. The default LLM provider is mock — fully offline and deterministic — so the whole system runs end-to-end with no API key required.

Port already in use? Every published host port is configurable: POSTGRES_PORT=5433 API_PORT=3001 docker compose up --build.

Try it

# bash / macOS / Linux
./scripts/demo.sh
# Windows PowerShell
./scripts/demo.ps1

Or by hand:

# 1. Ingest a transcript -> 202 Accepted + id
curl -s -X POST localhost:3000/transcripts \
  -H 'content-type: application/json' \
  --data-binary @examples/transcript-apollo.json

# 2. Check processing status (pending -> completed) + generated files
curl -s localhost:3000/transcripts/<id>

# 3. Explore the memory tree
curl -s 'localhost:3000/memories/ls'
curl -s 'localhost:3000/memories/ls?path=projects'
curl -s 'localhost:3000/memories/cat?path=projects/apollo-project-kickoff.md'
curl -s 'localhost:3000/memories/grep?pattern=deadline&i=true'

Using the real LLM (Anthropic)

ANTHROPIC_API_KEY=sk-ant-... LLM_PROVIDER=anthropic docker compose up --build

The provider auto-selects: if LLM_PROVIDER is unset, it uses anthropic when ANTHROPIC_API_KEY is present, otherwise mock.


API

Transcripts

Method Path Description
POST /transcripts Ingest a transcript. Persists it, enqueues memory generation, returns 202 + { id, status }.
GET /transcripts/:id Retrieve a stored transcript with its processing status, memoryPaths, and any error.

POST /transcripts body:

{
  "title": "Apollo project kickoff",
  "messages": [
    { "role": "user", "content": "The launch deadline is March 15." }
  ],
  "metadata": { "channel": "slack" }
}

Memories (unix-style)

Method Path Mirrors Description
GET /memories/ls?path=<dir> ls List immediate dirs + files under a path (root by default). Files include title/summary/updatedAt.
GET /memories/cat?path=<file> cat Return a parsed memory (frontmatter + body). Add &raw=true for the literal Markdown file.
GET /memories/grep?pattern=<re>&i=true grep Line-by-line regex search across the tree. i=true for case-insensitive, path= to scope, max= to cap.

Errors are consistent JSON: { "error": { "code, "message", "details" } } with appropriate status codes (400, 404, 500).


Memory design

This is the core of the system, so it's worth spelling out.

File tree layout

tree/
  people/        facts about individuals (colleagues, contacts)
  projects/      ongoing work, goals, status
  preferences/   how the user likes things done
  decisions/     choices made and their rationale
  tasks/         action items / commitments
  knowledge/     reusable domain facts and references
  misc/          anything that doesn't fit
  • Categories are a small, controlled vocabulary. This keeps the tree navigable and makes ls meaningful, instead of an unbounded sprawl of folders.
  • One file per coherent subject (people/jane-doe.md, projects/apollo.md). Files are named by a slug derived from the title, so the same subject maps to a deterministic path — essential for idempotent updates.

File format

Each memory is Markdown with a YAML frontmatter header:

---
title: Apollo project kickoff
category: projects
tags: [apollo, deadline, postgres]
summary: Customer-analytics rewrite; launch and stack decisions.
created_at: '2026-06-06T01:00:00.000Z'
updated_at: '2026-06-06T01:05:00.000Z'
source_transcript_ids: ['…', '…']   # provenance
version: 2
---

Notes captured from conversation.

- The launch deadline slipped to April 1, 2026.
- We decided to use Postgres for storage and Kafka for the event pipeline.
- Jane Doe is the tech lead; she owns the data-migration workstream.
  • Bodies are atomic bullet points — one fact per line — which makes grep genuinely useful: a match points at a single, self-contained fact.
  • Frontmatter is owned by the service, never the LLM. Timestamps, version, and source_transcript_ids are managed in MemoryGenerator, so provenance is always accurate and reprocessing is convergent.

How new transcripts interact with existing memories (append / merge / update)

  1. Retrieve. Before generating, the service ranks existing memories by keyword overlap with the transcript and passes the top matches to the LLM as context (a lightweight stand-in for a vector store — see tradeoffs).
  2. Decide + merge in one shot. The LLM returns a set of operations. For an update it returns the exact existing path and the full merged body (facts it still believes + new facts). Keeping merge logic in the prompt makes the apply step deterministic and side-effect-free.
  3. Apply with provenance. The service writes the file, preserving created_at, appending the transcript id to source_transcript_ids, and bumping version. So a follow-up transcript about the same project updates the same file rather than creating a near-duplicate.

The structured LLM contract (an Anthropic tool with a JSON schema, validated with Zod) means malformed output is rejected and retried rather than silently stored.


Background processing (reliability, retries, idempotency)

  • Outbox pattern. The transcript row and its job row are inserted in a single transaction, so a transcript is never stored without a job (or vice-versa).
  • One job per transcript (UNIQUE constraint) → idempotent ingestion.
  • Safe claiming. The worker claims jobs with SELECT … FOR UPDATE SKIP LOCKED, so multiple workers scale horizontally without double-processing.
  • At-least-once + convergent application. A crash mid-processing leaves the job claimable again; because memory paths are deterministic and writes are versioned upserts, re-running a transcript converges to the same files.
  • Retries with backoff → dead-letter. Failures schedule a retry with exponential backoff (run_after), and after JOB_MAX_ATTEMPTS the job moves to dead and the transcript is marked failed with the error recorded.

Architecture & code organization

src/
  config.ts            validated env config (one place)
  app.ts               Express app (routes + error handler), built from a DI container
  index.ts / worker.ts process entrypoints (api / worker)
  container.ts         production wiring (Postgres + S3 + provider)
  routes/              transcripts, memories (thin HTTP layer)
  services/
    memoryStore.ts     ls / cat / grep / write over the Storage interface
    memoryGenerator.ts retrieval -> LLM -> apply (frontmatter & provenance)
    worker.ts          claim/process loop, drainQueue (also used by tests)
  storage/             Storage interface + S3Storage + InMemoryStorage
  repositories/        TranscriptRepository + JobQueue interfaces; pg + in-memory impls
  llm/                 provider interface, prompt, AnthropicProvider, MockProvider
  domain/              memory (paths, slug, frontmatter), transcript, job
  db/                  pool, idempotent schema.sql, migrate (advisory-locked)

Dependency injection throughout. Storage, repositories, and the LLM provider are interfaces. Production wires Postgres + S3 + Anthropic; tests wire in-memory implementations — which is what lets the e2e test run the whole flow with zero external services.


Testing

A proper pyramid (npm test runs unit + e2e; integration is opt-in):

  • Unit (tests/unit) — pure logic with the in-memory storage/repos: frontmatter round-trips, slug/path rules, ls/cat/grep semantics (incl. case-insensitivity, scoping, bad-regex → 400), the mock provider, the generator's create-vs-update/merge + provenance, and job backoff / retry → dead-letter.

  • e2e (tests/e2e) — the full HTTP flow via supertest against an all-in-memory container: ingest → worker → ls/cat/grep, plus error paths (400/404).

  • Integration (tests/integration) — S3Storage against real MinIO and the Postgres repositories against real Postgres (claim → complete, idempotent job uniqueness). Opt-in:

    # with the docker-compose services running:
    RUN_INTEGRATION=true \
    DATABASE_URL=postgres://memory:memory@localhost:5432/memory_wiki \
    S3_ENDPOINT=http://localhost:9000 \
    npm run test:integration

The test scripts launch Jest via node --experimental-vm-modules because the AWS SDK lazily dynamic-imports a checksum module that Jest's VM otherwise blocks.


Assumptions

  • Transcripts are modest in size (single request, ≤ 4 MB body). Very large transcripts would warrant chunking before the LLM call.
  • A small, fixed category vocabulary is preferable to free-form folders for navigability. New categories are a one-line change.
  • "Durable memory" excludes small talk; the prompt (and the mock heuristic) deliberately drop greetings/filler.
  • Single-tenant. There's no auth or per-user namespacing (see future work).

Tradeoffs considered

  • LLM-side merge vs. programmatic diff. I let the LLM emit the full merged body given the existing content, rather than computing diffs in code. This handles fuzzy semantic overlap ("deadline slipped") that a textual diff can't, at the cost of trusting the model — mitigated by service-owned frontmatter and schema validation.
  • Keyword retrieval vs. embeddings. Relevance is keyword-overlap, not vector search. It's dependency-free and good enough at this scale; the retrieval seam (selectRelevant) is isolated so swapping in embeddings is local.
  • grep scans objects. It lists and reads files on demand. Simple and correct for thousands of small files; for millions you'd want a search index.
  • Postgres-based queue vs. a broker. A jobs table with SKIP LOCKED gives transactional enqueue (outbox), retries, and dead-lettering without operating Kafka/SQS/Redis. Fine to this scale; swap the JobQueue impl if you outgrow it.
  • Mock provider trades quality for determinism/offline-runnability — it's a test/dev fallback, not the product.

What I'd do with more time

  • Embeddings-based retrieval (pgvector) for better merge targeting.
  • Auth + multi-tenancy: namespace the tree per user/org (tree/<tenant>/…).
  • A two-pass generator: cheap extraction pass, then a targeted merge pass that reads only the precise files it will touch (less context, lower cost).
  • Memory maintenance: dedupe/garbage-collect, conflict detection ("deadline" stated twice with different dates), and a rm/history endpoint.
  • Observability: OpenTelemetry traces across api → queue → worker → S3, and metrics on queue depth / processing latency / dead-letter rate.
  • Streaming + chunking for very long transcripts; idempotency keys on POST /transcripts to dedupe client retries.
  • Richer grep (context lines, JSON vs. text output) and a tree endpoint.

About

No description, website, or topics provided.

Resources

Stars

24 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages