Pseudolife-MCP

MCP serverDocs & knowledge

Persistent memory for MCP-compatible agents: memory bank, fact cortex, dreams, graph.

Unavailable. This server has no hosted endpoint yet, so ahel can't serve it.

Add to setup to save this item as a reference. ahel cannot run it, and signing in will not install it.

Getting started

  1. Save this item in Your setup as a reference.
  2. Read the source or reference documentation for its setup requirements. Saving it here does not connect it to your AI.
  3. Check this page for availability before trying to install it through ahel.

From the project's README

As published by pseudogiant-xr/pseudolife-mcp in README.md.

简体中文 · 日本語 · 한국어 · Português (BR) · Español

Persistent long-term memory for Claude Code, Codex, and other MCP clients.

An MCP server that gives coding agents a long-term memory that persists across sessions — surviving context compactions and fresh tasks. Your coding agent is the intelligence; this server is its memory on disk.

What you get:

  • Associative memory with honest forgetting — a flat similarity store ranked by hybrid dense-plus-lexical retrieval, with conflict detection that admits potential updates while preserving earlier source notes; whole-note replacement is explicit. (The measured verdict: a preregistered ablation campaign found the previous 8-band continuum tied a flat store on every gate, so the simpler structure ships; the continuum remains one config line away.)
  • Canonical facts, not vibes — one current value per entity.attribute slot (or a member set, for slots that hold many concurrent values); corrections supersede rather than silently overwrite, and the full version history survives.
  • Dreams — a bundled local extractor, or any OpenAI-compatible endpoint (a Claude model on your Max plan, a GPT-5.6 model on a ChatGPT plan, LM Studio, Ollama, vLLM), consolidates the memory stream into facts and a knowledge graph while you're not looking.
  • Lessons from its own work — successes, dead-ends, and your corrections become do/avoid guidance surfaced at the start of every session.
  • A web console to watch it think — the Cortex Console above, plus cited world facts, session episodes, and document RAG.

Measured, with receipts — the full 500-question LongMemEval sweep, all six question types, and every number ships with its committed run artifact:

LongMemEval oracle, 500 questionsnaive RAGcommit-gated cascade
accuracy, all six question types0.6880.690
context tokens per question~1210~883
knowledge-update slice (78 of the 500)0.8590.936 (retired — see below)

Equal accuracy to naive RAG across the whole benchmark on ~73% of the context, and better calibrated about what it does not know: on BEAM-100K's abstention questions the fact spine scores 0.950 against naive RAG's 0.775, unchanged under two independent judges. Read that as calibration, not recall — in the budget-matched five-arm run of 2026-09-02 (rag 0.725 there; one replicate, local judge) an arm served no memory at all scores 1.000 on the same questions, because refusing is the right answer there and an empty context always refuses. The fact spine loses where an answer has to be aggregated across sessions. The second claim to survive a judge swap is a win rather than a wash: re-run on 2026-09-04 with the hybrid arm budget-matched to the control at 6 turns, the same 500 questions give hybrid 0.730 against naive RAG's 0.690 under the local judge and 0.736 against 0.694 under claude-opus-5 — paired +0.040 / +0.042, p 0.015 / 0.013 — bought with more context, ~1229 tokens against the control's ~1124, not less, and carried mostly by temporal-reasoning questions. Graded by a local, byte-reproducible judge (the cross-judge check names its second judge) — compare within rows, never against GPT-judged leaderboards.

Retired 2026-08-25 (#188): the 0.936 knowledge-update headline. It was measured on the 2026-07-30 bench stack (Qwen3.6-27B answerer and judge). Re-running the same 78 questions after the 2026-08-17 migration to Qwen3.8-27B puts the cascade at 0.846, below the naive-RAG control — which lands on 0.859 on both stacks. The cascade serves the fact-spine answer unless that channel says "I don't know", so it measures the answerer's abstention behaviour as much as the memory: 32/78 abstentions at 46/46 commit precision on the old stack, 22/78 at 0.839 on the new one. The 500-question table above is on the older judge and has not been re-judged, so read its cascade row as an upper bound.

Full tables, the per-type breakdown, both stacks side by side, and every artifact: Benchmarks.

Quickstart

Install and register the lite tier. No Docker, no database to set up, no container runtime:

pip install "pseudolife-mcp[lite]"
claude mcp add --scope user pseudolife-memory -- pseudolife-mcp

Codex instead of Claude Code — same shape:

pip install "pseudolife-mcp[lite]"
codex mcp add pseudolife-memory --env PSEUDOLIFE_WRITER_ID=codex -- pseudolife-mcp

For Codex, finish setup before starting a fresh task. In the existing [mcp_servers.pseudolife-memory] table in ~/.codex/config.toml, add startup_timeout_sec = 240, tool_timeout_sec = 240, and required = true. The shim can wait up to 180 seconds for a cold daemon; Codex's default startup budget is 10 seconds. required makes missing memory visible at startup and waits for its initial catalog. These are starting budgets, not a promise that a first model download fits. The tool budget leaves time for the shim's 180-second deadline to report a failure before the host cancels it; prewarm with pseudolife-mcp serve in a terminal if needed.

The MCP handshake delivers compact recall/capture/reflection instructions. For the complete standing guidance, copy the bundled memory block into your project AGENTS.md or ~/.codex/AGENTS.md. For session briefings and per-turn reminders, follow Codex hooks and verification. Use one MCP registration and one hook source; an installed plugin may already provide either. After the daemon is running, execute pseudolife-mcp doctor from the same environment as the registered command. It checks the handshake and annotations without calling bank tools.

Then in either coding agent: "remember that my staging box is haze-02" → the agent calls memory_store; next session, "which box is staging?" → memory_search finds it. Browse everything at the Cortex Console: http://127.0.0.1:8765/ui/.

The first session auto-starts the daemon, which provisions an embedded PostgreSQL 18 (pgvector included, via pg0-embedded) under a stable per-user data dir and downloads the embedding model (~1.2 GB, one-time). It is a real Postgres bank, not a cut-down one: pseudolife-mcp backup writes a standard owner-free pg_dump archive (plus a state archive, 7-day rotation) that restores into any PostgreSQL 18 target regardless of role — the Docker tier included — so outgrowing lite is a dump/restore, not a migration project (backups). For a tier- and Postgres-version-independent copy, pseudolife-mcp export / import move the whole bank as portable JSONL (logical export / import). Windows needs an ASCII-only data path (PSEUDOLIFE_MCP_DATA_DIR).

What lite gives you, and the one thing it doesn't

lite (pip)durable (Docker)
Associative store, hybrid search, supersession, version historyyesyes
Cortex facts, knowledge graph, lessons, world facts, episodesyesyes
Cortex Console, document RAG, pseudolife-mcp backupyesyes
Dream consolidation filling the cortex on its ownno extractor shipsyes — bundled local CPU sidecar
External volumes, health-checked services, deploy/rollback toolingnoyes

The gap, stated plainly. Lite ships no extractor, so the dream pass still runs, prunes, and acknowledges its input batch, but writes no canonical facts: on this path memory_fact_set is the only cortex writer. Everything else above works. Nothing about this is silent — curl http://127.0.0.1:8765/health reports "extractor": "none", and the stdio shim says the same on stderr at session start.

Any OpenAI-compatible endpoint closes it. The daemon inherits the environment it starts from, so two variables are the whole fix — with a local Ollama:

export PSEUDOLIFE_DREAM_BASE_URL=http://localhost:11434/v1
export PSEUDOLIFE_DREAM_MODEL=qwen2.5:7b
pseudolife-mcp serve
$env:PSEUDOLIFE_DREAM_BASE_URL = "http://localhost:11434/v1"
$env:PSEUDOLIFE_DREAM_MODEL    = "qwen2.5:7b"
pseudolife-mcp serve

/health then reports "extractor": "configured". One gotcha: a daemon that is already running keeps the environment it started with, and the shim reattaches to it rather than spawning a new one — stop the old daemon first. A hosted endpoint works too, and costs you the zero-egress property: memory text leaves the machine. Extractor tiers, quality, and the trade-offs: Dreaming.

Durable tier — Docker (recommended for a long-lived bank)

Everything above plus the bundled extractor, external volumes, health-checked services, and backup/rollback tooling. Requires Docker and at least one MCP-capable coding agent — Claude Code, Codex, and Gemini CLI are wired end-to-end; anything else gets paste-ready config (provider matrix). One command from clone to first memory:

git clone https://github.com/Pseudogiant-xr/Pseudolife-MCP.git
cd Pseudolife-MCP
ops/install.sh          # Linux / macOS
ops\install.ps1         # Windows (pwsh 7+)
# Codex: add --client codex / -Client codex
# Codex defaults to automatic hook-source detection and asks once for approval.
# Unattended hook approval: --codex-hook-trust yes / -CodexHookTrust yes
# Instructions only: --codex-hooks skip --instructions append
# PowerShell equivalent: -CodexHooks skip -Instructions append
# Both:  add --client both  / -Client both
# Gemini: add --client gemini — or several: --client claude,codex,gemini
# Other MCP agents (Cursor, Windsurf, Zed, ...): --client generic

The installer asks which agents to wire (multi-select, with a capability matrix showing exactly what each one gets — session briefing, per-turn discipline, standing file), runs the preflight (one exact fix line per missing prerequisite), then asks which dream extractor should consolidate memories —

  • sidecar — the bundled local CPU model; no Claude plan needed, works for everyone, and keeps every memory on the box (~11.8 GB image);
  • sonnet-only — the lightest install: a Claude model via a CLI shim (claude-opus-5 by default; the mode name is historical. Needs a logged-in Max-plan claude CLI); the sidecar image is never built or pulled (~11.8 GB lighter; dreams pause while the shim is down);
  • sonnet-fallback — the Claude shim primary, the bundled sidecar as automatic fallback (Max-plan CLI plus the ~11.8 GB image);
  • codex-only / codex-fallback — the same two shapes on an OpenAI subscription: a GPT-5.6 model (Sol / Terra / Luna) via the Codex CLI shim on a signed-in ChatGPT plan (extraction quality unmeasured — see the dreaming guide) —

then brings the stack up, installs the selected clients' session hooks (where the client has a hook system), registers the MCP transport (the stdio shim by default, with a per-provider writer id; direct HTTP via --transport http), and health-checks the daemon — finishing with a per-agent ladder of what got wired and what that agent's platform cannot support. Codex setup offers one choice to enable automatic memory briefings, reminders, and session cleanup, use standing instructions only, or skip. Automatic setup reuses an enabled PseudoLife plugin or installs the three lifecycle hooks, backs up configuration, approves only their exact current definitions, and verifies execution. If verification fails, the same approval allows the standing memory block as a fallback; setup reports the remaining repair step. Hook-less providers (Gemini CLI and generic agents) are offered the standing block. --instructions append always writes the block from examples/CLAUDE.memory.md into ~/.claude/CLAUDE.md / ~/.codex/AGENTS.md / ~/.gemini/GEMINI.md (useful for subagent visibility even with hooks). Idempotent — re-run any time; --extractor <mode> switches extractor setups. Non-interactive example: ops/install.sh --extractor sidecar --client codex --codex-hook-trust yes. Without explicit hook approval, unattended setup does not grant trust; use --codex-hooks skip --instructions append for instructions only. Explicit --instructions skip prevents fallback edits. Linux (Docker Engine): your user must be in the docker group — sudo usermod -aG docker $USER, then log out/in (the preflight checks this).

Image sizes, the Windows WSL2 memory cap, and what the installer automates: the containerized install below.

ops/preflight.sh --client codex    # or ops\preflight.ps1 -Client codex
docker volume create pseudolife-mcp-bank
docker volume create pseudolife-mcp-state
docker compose -f ops/docker-compose.yml up -d --build   # first build, once

# ...or pull the prebuilt images instead of building (releases >= 0.14.0):
docker compose -f ops/docker-compose.yml -f ops/docker-compose.ghcr.yml pull pseudolife-pg pseudolife-daemon
docker compose -f ops/docker-compose.yml -f ops/docker-compose.ghcr.yml up -d

# Verify, then wire the transport into one or both clients.
curl http://127.0.0.1:8765/health

# Stdio shim (the installer's default — per-session episode identity).
# PSEUDOLIFE_MCP_NO_SPAWN=1 makes the shim wait for the container instead
# of spawning a host fallback that can shadow the Docker bank after a
# reboot; set it on Docker-tier registrations like these.
pip install pseudolife-mcp
claude mcp add --scope user pseudolife-memory --env PSEUDOLIFE_MCP_NO_SPAWN=1 -- pseudolife-mcp
codex mcp add pseudolife-memory --env PSEUDOLIFE_MCP_NO_SPAWN=1 -- pseudolife-mcp

# ...or direct HTTP (no pip package needed; fine for single-session setups):
claude mcp add --transport http --scope user pseudolife-memory http://127.0.0.1:8765/mcp
codex mcp add pseudolife-memory --url http://127.0.0.1:8765/mcp

# Reinforce the protocol-level memory loop with a global standing instruction:
cat examples/CLAUDE.memory.md >> ~/.claude/CLAUDE.md
cat examples/CLAUDE.memory.md >> ~/.codex/AGENTS.md
# (PowerShell: Add-Content "$env:USERPROFILE\.claude\CLAUDE.md" (Get-Content examples\CLAUDE.memory.md -Raw))

Optional knobs live in ops/.env (cp ops/.env.example ops/.env — the install/update scripts scaffold it too; every value is commented, a missing file runs entirely on defaults).

What this is

A memory engine exposed over MCP. There's no chat UI and no LLM doing the thinking — your coding agent is the intelligence; these are tools it calls to store and recall what matters. (Models are bundled as plumbing: baked embedding weights for retrieval, and the optional CPU extractor sidecar that consolidates memories into facts while you sleep.)

Where it sits among the common approaches to agent memory — each column is a fair tool for what it's for; this table is about what question each one answers, not who's wrong:

notes file (CLAUDE.md)auto-journaling pluginplain vector storePseudolife-MCP
Survives sessions and compactionsyesyesyesyes
"What is X now?" has one current answerif you curate itno — replays what happenedno — every stored version competes at recallyes — slot-keyed cortex
A canonical-fact correction replaces the old valueyou edit the fileappended beside itold and new both retrievable, unranked by recency of truthcortex supersedes, with full version history kept
Facts know their age and go stalenononodated, freshness-decayed, quarantined when stale
Distils do/avoid lessons from its own outcomesnononoyes
Benchmark numbers ship with their raw run artifacts—typically notypically noevery published number, test-enforced

Auto-journaling records what the agent did; Pseudolife curates what it learned. Both are useful — they answer different questions. Named alternatives — Mem0, Zep/Graphiti, Letta, Cognee, memU, Memori — and the cases where one of them is the better pick: Comparison.

It layers several complementary stores: the associative store (a flat embedding store ranked by cosine similarity fused with a BM25 lexical pool (on by default), with conflict-aware admission and explicit source-note replacement; an 8-tier banded layout is available as an opt-in preset); the cortex (slot-keyed canonical facts — one current value per entity.attribute, or a member set for set-valued slots — with provenance tiers and contender parking instead of silent overwrites); a typed knowledge graph over those facts with a closed relation vocabulary and on-read inference; the world cortex (durable cited facts about external reality, age-decayed trust); procedural lessons learned from the agent's own work; and a ChromaDB reference bank for document RAG. The canonical layers in depth: the memory model; the graph and multi-hop recall: retrieval.

State lives in Postgres (the durable source of truth) behind a single long-lived daemon; every session attaches through a thin stdio shim (installer default — per-session identity) or directly over HTTP (single-session setups). The result: Claude can pick up where it left off, correct itself when facts change, and reason over relationships — without you re-explaining context each session.

Documentation

This README is the front door — install, wiring, and the basic loop. The deep material lives in the user guide:

PageWhat's in it
ConfigurationEnv vars, tuned defaults, toolset tiers, stdio shim, LAN sharing, data layout, backups, schema history
ProvidersCapability matrix per coding agent, memory instruction layers, AGENTS.md standard, Codex hook setup and verification, writer ids
RetrievalReranker, BM25 hybrid, abstention floors, ranking-trace debugging, memory_recall, the knowledge graph
DreamingExtractor tiers, the bundled sidecar, upgrading the extractor, Sonnet-fallback, cadence, deep dream, consolidation
Episodes & sessionsDaemon-owned session episodes, the briefing hook, nested sub-episodes, tags
The memory modelCortex slots, provenance contenders, world cortex, lessons, temporal/HLC stamps
BenchmarksLongMemEval results; why extraction quality dominates
ComparisonMem0, Zep/Graphiti, Letta, Cognee, memU, Memori — the axes, and when to use something else
Security postureMemory poisoning (ASI06): every shipped mitigation, and what is not defended

Plus evals/README.md (full benchmark methodology) and CONTRIBUTING.

Tools exposed

The surface was consolidated 2026-07-02 (55 → 32 tools; now 37 with memory_toolset, the set-slot pair and coordination): lifecycle families became verb-dispatched tools (memory_dream, memory_forget, memory_graph_review), and dump/introspection views moved to the Cortex Console (REST) — the manifest is agent context every session, so it stays lean.

Shortened here. Read the whole README on GitHub.

Signals

GitHub stars
5
Forks
1
Last commit
Sep 2026
Advanced
Delivery
pseudolife-mcp MCP server → your ahel connector (mcp.ahel.ai) → your AI.
Item type
mcp-server
Key
io-github-pseudogiant-xr-pseudolife-mcp
Source
github.com/pseudogiant-xr/pseudolife-mcp