Agent State and Memory
SkillDocs & knowledgeDurable agent state, memory, and replay for autonomous research workflows. Use when an agent must resume, audit, or compare multi-step runs.
Available today. Use it from your connected AI after setup.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the Agent State and Memory skill
What this skill tells your AI
The instructions your AI receives, as published by ml4t/skills in advanced-ai/agent-state-memory/SKILL.md and read by ahel’s review.
Agent memory should be an explicit state object that can be serialized, replayed, and inspected. Chat history is not enough because it loses typed evidence, tool provenance, and quality gate status.
The Problem
Long-running research agents make dozens of tool calls and intermediate judgments. If the only memory is the prompt transcript, the run cannot be resumed safely, audited by a reviewer, or compared against an ablation. The agent may also reuse stale claims because evidence has no timestamp or source boundary.
The Pattern
WRONG
messages = []
messages.append({"role": "user", "content": task})
while True:
answer = llm(messages)
messages.append(answer)
if "done" in answer["content"].lower():
break
CORRECT
from dataclasses import asdict, dataclass, field
from pathlib import Path
from typing import Any
import json
@dataclass
class AgentState:
task: str
evidence: list[dict[str, Any]] = field(default_factory=list)
tool_trace: list[dict[str, Any]] = field(default_factory=list)
decisions: list[dict[str, Any]] = field(default_factory=list)
quality_gates: dict[str, bool] = field(default_factory=dict)
def checkpoint(self, path: Path) -> None:
path.write_text(json.dumps(asdict(self), indent=2), encoding="utf-8")
@classmethod
def load(cls, path: Path) -> "AgentState":
return cls(**json.loads(path.read_text(encoding="utf-8")))
state = AgentState(task="Evaluate whether the ensemble improves holdout Sharpe")
state.tool_trace.append(
{"tool": "query_registry", "status": "ok", "observed_at": current_utc_iso()}
)
state.quality_gates["read_relevant_skills"] = True
state.checkpoint(Path("runs/ensemble_eval/state.json"))
Memory Layers
- Run state - current task, evidence, decisions, open questions, gates
- Tool trace - every external observation with arguments and status
- Evidence memory - source, timestamp, freshness, and extracted claims
- Long-term memory - only stable project facts; never live market facts without dates
- Replay artifact - enough inputs and outputs to re-render the run without live APIs
Guardrails
- Transcript-only replay - reviewers cannot verify which tools produced which facts
- Undated evidence - market and API observations need explicit timestamps
- Mutable memory overwrite - append decisions; do not silently edit prior reasoning
- Cross-run contamination - reset task state between independent experiments
Checklist
- State serializes to a stable JSON artifact
- Evidence includes source, timestamp, and freshness notes
- Tool trace can be replayed or inspected without the LLM
- Quality gates are explicit booleans or statuses
- Long-term memory excludes ephemeral market observations
Signals
- GitHub stars
- 20
- Forks
- 11
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
ml4t-agent-state-memory- Source
- github.com/ml4t/skills