A service that ingests conversation transcripts, uses an LLM to distil durable
memories from them, stores those memories as files in a navigable, object-storage-backed
file tree, and exposes them through unix-style REST endpoints (ls, cat, grep).
POST /transcripts ──► Postgres (transcript + job, one TX)
│
▼ (background worker, at-least-once)
LLM memory generation
│
▼
Object storage file tree (S3 / MinIO)
tree/<category>/<slug>.md
│
┌─────────────────┼─────────────────┐
▼ ▼ ▼
GET /memories/ls GET /memories/cat GET /memories/grep
Requires Docker.
docker compose up --buildThis builds the app image and starts postgres, minio (+ bucket init),
a one-shot migrate step, the api (port 3000), and the worker.
The default LLM provider is mock — fully offline and deterministic — so the
whole system runs end-to-end with no API key required.
Port already in use? Every published host port is configurable:
POSTGRES_PORT=5433 API_PORT=3001 docker compose up --build.
# bash / macOS / Linux
./scripts/demo.sh# Windows PowerShell
./scripts/demo.ps1Or by hand:
# 1. Ingest a transcript -> 202 Accepted + id
curl -s -X POST localhost:3000/transcripts \
-H 'content-type: application/json' \
--data-binary @examples/transcript-apollo.json
# 2. Check processing status (pending -> completed) + generated files
curl -s localhost:3000/transcripts/<id>
# 3. Explore the memory tree
curl -s 'localhost:3000/memories/ls'
curl -s 'localhost:3000/memories/ls?path=projects'
curl -s 'localhost:3000/memories/cat?path=projects/apollo-project-kickoff.md'
curl -s 'localhost:3000/memories/grep?pattern=deadline&i=true'ANTHROPIC_API_KEY=sk-ant-... LLM_PROVIDER=anthropic docker compose up --buildThe provider auto-selects: if LLM_PROVIDER is unset, it uses anthropic when
ANTHROPIC_API_KEY is present, otherwise mock.
| Method | Path | Description |
|---|---|---|
POST |
/transcripts |
Ingest a transcript. Persists it, enqueues memory generation, returns 202 + { id, status }. |
GET |
/transcripts/:id |
Retrieve a stored transcript with its processing status, memoryPaths, and any error. |
POST /transcripts body:
{
"title": "Apollo project kickoff",
"messages": [
{ "role": "user", "content": "The launch deadline is March 15." }
],
"metadata": { "channel": "slack" }
}| Method | Path | Mirrors | Description |
|---|---|---|---|
GET |
/memories/ls?path=<dir> |
ls |
List immediate dirs + files under a path (root by default). Files include title/summary/updatedAt. |
GET |
/memories/cat?path=<file> |
cat |
Return a parsed memory (frontmatter + body). Add &raw=true for the literal Markdown file. |
GET |
/memories/grep?pattern=<re>&i=true |
grep |
Line-by-line regex search across the tree. i=true for case-insensitive, path= to scope, max= to cap. |
Errors are consistent JSON: { "error": { "code, "message", "details" } } with
appropriate status codes (400, 404, 500).
This is the core of the system, so it's worth spelling out.
tree/
people/ facts about individuals (colleagues, contacts)
projects/ ongoing work, goals, status
preferences/ how the user likes things done
decisions/ choices made and their rationale
tasks/ action items / commitments
knowledge/ reusable domain facts and references
misc/ anything that doesn't fit
- Categories are a small, controlled vocabulary. This keeps the tree
navigable and makes
lsmeaningful, instead of an unbounded sprawl of folders. - One file per coherent subject (
people/jane-doe.md,projects/apollo.md). Files are named by a slug derived from the title, so the same subject maps to a deterministic path — essential for idempotent updates.
Each memory is Markdown with a YAML frontmatter header:
---
title: Apollo project kickoff
category: projects
tags: [apollo, deadline, postgres]
summary: Customer-analytics rewrite; launch and stack decisions.
created_at: '2026-06-06T01:00:00.000Z'
updated_at: '2026-06-06T01:05:00.000Z'
source_transcript_ids: ['…', '…'] # provenance
version: 2
---
Notes captured from conversation.
- The launch deadline slipped to April 1, 2026.
- We decided to use Postgres for storage and Kafka for the event pipeline.
- Jane Doe is the tech lead; she owns the data-migration workstream.- Bodies are atomic bullet points — one fact per line — which makes
grepgenuinely useful: a match points at a single, self-contained fact. - Frontmatter is owned by the service, never the LLM. Timestamps,
version, andsource_transcript_idsare managed inMemoryGenerator, so provenance is always accurate and reprocessing is convergent.
- Retrieve. Before generating, the service ranks existing memories by keyword overlap with the transcript and passes the top matches to the LLM as context (a lightweight stand-in for a vector store — see tradeoffs).
- Decide + merge in one shot. The LLM returns a set of operations. For an update it returns the exact existing path and the full merged body (facts it still believes + new facts). Keeping merge logic in the prompt makes the apply step deterministic and side-effect-free.
- Apply with provenance. The service writes the file, preserving
created_at, appending the transcript id tosource_transcript_ids, and bumpingversion. So a follow-up transcript about the same project updates the same file rather than creating a near-duplicate.
The structured LLM contract (an Anthropic tool with a JSON schema, validated with Zod) means malformed output is rejected and retried rather than silently stored.
- Outbox pattern. The transcript row and its job row are inserted in a single transaction, so a transcript is never stored without a job (or vice-versa).
- One job per transcript (
UNIQUEconstraint) → idempotent ingestion. - Safe claiming. The worker claims jobs with
SELECT … FOR UPDATE SKIP LOCKED, so multiple workers scale horizontally without double-processing. - At-least-once + convergent application. A crash mid-processing leaves the job claimable again; because memory paths are deterministic and writes are versioned upserts, re-running a transcript converges to the same files.
- Retries with backoff → dead-letter. Failures schedule a retry with
exponential backoff (
run_after), and afterJOB_MAX_ATTEMPTSthe job moves todeadand the transcript is markedfailedwith the error recorded.
src/
config.ts validated env config (one place)
app.ts Express app (routes + error handler), built from a DI container
index.ts / worker.ts process entrypoints (api / worker)
container.ts production wiring (Postgres + S3 + provider)
routes/ transcripts, memories (thin HTTP layer)
services/
memoryStore.ts ls / cat / grep / write over the Storage interface
memoryGenerator.ts retrieval -> LLM -> apply (frontmatter & provenance)
worker.ts claim/process loop, drainQueue (also used by tests)
storage/ Storage interface + S3Storage + InMemoryStorage
repositories/ TranscriptRepository + JobQueue interfaces; pg + in-memory impls
llm/ provider interface, prompt, AnthropicProvider, MockProvider
domain/ memory (paths, slug, frontmatter), transcript, job
db/ pool, idempotent schema.sql, migrate (advisory-locked)
Dependency injection throughout. Storage, repositories, and the LLM provider are interfaces. Production wires Postgres + S3 + Anthropic; tests wire in-memory implementations — which is what lets the e2e test run the whole flow with zero external services.
A proper pyramid (npm test runs unit + e2e; integration is opt-in):
-
Unit (
tests/unit) — pure logic with the in-memory storage/repos: frontmatter round-trips, slug/path rules,ls/cat/grepsemantics (incl. case-insensitivity, scoping, bad-regex → 400), the mock provider, the generator's create-vs-update/merge + provenance, and job backoff / retry → dead-letter. -
e2e (
tests/e2e) — the full HTTP flow viasupertestagainst an all-in-memory container: ingest → worker →ls/cat/grep, plus error paths (400/404). -
Integration (
tests/integration) —S3Storageagainst real MinIO and the Postgres repositories against real Postgres (claim → complete, idempotent job uniqueness). Opt-in:# with the docker-compose services running: RUN_INTEGRATION=true \ DATABASE_URL=postgres://memory:memory@localhost:5432/memory_wiki \ S3_ENDPOINT=http://localhost:9000 \ npm run test:integration
The test scripts launch Jest via
node --experimental-vm-modulesbecause the AWS SDK lazily dynamic-imports a checksum module that Jest's VM otherwise blocks.
- Transcripts are modest in size (single request,
≤ 4 MBbody). Very large transcripts would warrant chunking before the LLM call. - A small, fixed category vocabulary is preferable to free-form folders for navigability. New categories are a one-line change.
- "Durable memory" excludes small talk; the prompt (and the mock heuristic) deliberately drop greetings/filler.
- Single-tenant. There's no auth or per-user namespacing (see future work).
- LLM-side merge vs. programmatic diff. I let the LLM emit the full merged body given the existing content, rather than computing diffs in code. This handles fuzzy semantic overlap ("deadline slipped") that a textual diff can't, at the cost of trusting the model — mitigated by service-owned frontmatter and schema validation.
- Keyword retrieval vs. embeddings. Relevance is keyword-overlap, not vector
search. It's dependency-free and good enough at this scale; the retrieval seam
(
selectRelevant) is isolated so swapping in embeddings is local. grepscans objects. It lists and reads files on demand. Simple and correct for thousands of small files; for millions you'd want a search index.- Postgres-based queue vs. a broker. A jobs table with
SKIP LOCKEDgives transactional enqueue (outbox), retries, and dead-lettering without operating Kafka/SQS/Redis. Fine to this scale; swap theJobQueueimpl if you outgrow it. - Mock provider trades quality for determinism/offline-runnability — it's a test/dev fallback, not the product.
- Embeddings-based retrieval (pgvector) for better merge targeting.
- Auth + multi-tenancy: namespace the tree per user/org (
tree/<tenant>/…). - A two-pass generator: cheap extraction pass, then a targeted merge pass that reads only the precise files it will touch (less context, lower cost).
- Memory maintenance: dedupe/garbage-collect, conflict detection
("deadline" stated twice with different dates), and a
rm/history endpoint. - Observability: OpenTelemetry traces across api → queue → worker → S3, and metrics on queue depth / processing latency / dead-letter rate.
- Streaming + chunking for very long transcripts; idempotency keys on
POST /transcriptsto dedupe client retries. - Richer
grep(context lines, JSON vs. text output) and atreeendpoint.