Long-term memory
for autonomous agents.

A local-first, dual-process cognitive memory engine that gives AI agents persistent recall without blocking their real-time execution loop. Built with SQLite, vector search, and native Model Context Protocol (MCP) support.

ENGINE ARCHITECTURE • Dual-Process Cognitive Memory
SQLite + Vector LTM
System 1 • Immediate Buffer < 5ms Ingestion
  • Attention Gate: Discards low-salience noise before it pollutes the context window.
  • Episode Buffer: Synchronous, low-latency appending to local storage.
  • Miller's Law Guard: Strictly bounds active working memory to 7 ± 2 slots.
  • Zero Inference Lag: Does not block model generation or tool execution loops.
System 2 • Async Reflection Background Daemon
  • Semantic Palace Graph: Extracts relational entities and rooms when agent is idle.
  • Conflict Resolution: Automatically resolves contradictory facts without duplicates.
  • Spaced Repetition: Reinforces recalled nodes; decays stale, unaccessed facts.
  • Obsidian Vault Mirror: Continuous bidirectional sync to human-readable Markdown.

The SMRITI Stack

Modular developer tooling to capture, inspect, and curate agent memory across your entire workflow.

SMRITI Desktop

macOS Application

The visual control deck. A standalone native macOS GUI to inspect real-time memory formation, trigger System 2 consolidation, and manage Obsidian sync.

Download .dmg (v1.0.0) →

AIMeter

macOS Cost & Token Proxy

The spending monitor. A lightweight local proxy and native macOS menu bar widget tracking real-time API tokens and costs across Cursor, Claude Code, and custom agent swarms.

Download AIMeter.dmg (v0.3.9) →

SMRITI Core

Python SDK

The foundation. A high-performance LTM engine with multi-factor retrieval, salience filtering, and memory consolidation APIs for custom agent architectures.

SMRITI Daemon

Background Service

The pulse. An always-on, low-overhead background daemon that manages memory reflection, conflict resolution, and spaced-repetition decay when the agent is idle.

MCP Server

Model Context Protocol

The bridge. A native MCP server exposing 19 memory tools to developer environments like Claude Code, Cursor, Gemini Antigravity, and custom MCP clients.

Obsidian Sync

Vault Integration

The mirror. Automatically syncs and translates the agent's Semantic Palace Graph into clean markdown files inside an Obsidian vault for human curation.

Graph Explorer

D3.js Visualizer

The window. A built-in, responsive D3.js visualization dashboard with Prometheus metrics monitoring to analyze memory strength and recall latency.

Specifications & Tooling

Open-source specifications and developer utilities built for the AI agent memory ecosystem.

Project
Description
Quick Install / Use
Link
macOS App

Smriti Desktop

Native macOS application with real-time memory visualizer, live daemon logs, Miller's Law slot monitor, and Obsidian vault sync controller.

macOS App

AIMeter

The ultra-fast, local-first LLM API cost & token tracker for macOS. Real-time menu bar indicator, zero-latency proxy, and Claude Code watcher with local SQLite storage.

AMP Logo Specification

Agent Memory Protocol (AMP)

An open, community-driven specification defining a standard, vendor-agnostic interface for persistent memory in AI agents. Shares identical schemas across REST and MCP.

pip install amp-server
CLI Tool

git-story

An interactive Git history and codebase architecture evolution slide-deck generator. Summarizes structural changes using LLMs and compiles them into portable HTML presentations.

pip install git-story
AI Security

mcp-guard

A lightweight local firewall and permission gateway proxy. Sits between AI editors (Cursor, Claude Code) and MCP servers, prompting for confirmation before executing dangerous commands.

npm install -g mcp-guard
Observability

mcp-lens

A local visual workspace supervisor and topology mapper for MCP servers. Aggregates and tests active MCP servers configured across your environment from a single dashboard.

npx mcp-lens

Dual-Process Architecture

SMRITI splits memory operations to match human cognitive processes, separating real-time execution from deep reflection.

01

Attention Gate

Every incoming observation passes through a local salience filter. Irrelevant noise is discarded immediately, protecting the agent's context window from pollution.

02

System 1 (Immediate Ingestion)

Salient memories are instantly appended to the local **Episode Buffer** in under 5ms. The agent continues its execution loop without waiting for slow LLM processing or vector indexing.

03

System 2 (Async Consolidation)

When the agent is idle, background processes consolidate the raw episodes: extracting entities, building a **Semantic Palace Graph**, resolving contradictions, and applying temporal decay to weak memories.

Raw Input Attention System 1 (LTM) Episode Buffer System 2 (LTM) Palace Graph

Engineered for Production Agents

Built to solve context bloat and hallucination limits of standard vector databases.

01

Miller's Law Guard

Standard RAG floods context windows, causing LLM distraction. SMRITI bounds active working memory to **7 ± 2 slots**, ensuring only the most relevant, high-salience context is injected.

02

Private Rooms

Protect user privacy natively. Create localized semantic rooms. Memories tagged as `private` are isolated locally, preventing sensitive data from syncing to shared team storage.

03

Model Context Protocol

Ship memory out-of-the-box. SMRITI operates as a native MCP server, instantly integrating with Claude Code, Cursor, Gemini Antigravity, and custom agent systems.

04

Obsidian Sync

Make your agent's mind human-readable. SMRITI automatically maps and syncs its semantic palace graph into a local Obsidian vault, allowing you to curate, edit, and visualize memories.

05

Automatic Conflict Resolution

When new facts contradict old ones, System 2 background reflection resolves the conflict, updating outdated nodes and preserving historical context without duplicate entries.

06

Spaced Repetition & Decay

Memories have a dynamic **strength** value. Frequently recalled items are reinforced, while unused, low-salience details decay naturally over time, preventing database bloat.

Quickstart & Integration

Deploy as an MCP server for AI coding agents, or embed directly via Python SDK.

Terminal
# Method A: Install using the one-line bash installer
bash <(curl -s https://raw.githubusercontent.com/smriti-memcore/smriti-memcore/main/install_smriti_mcp.sh)

# Method B: Install via PyPI and run CLI setup
pip3 install smriti-memcore
smriti_install
SMRITI automatically registers itself in your Claude, Gemini, or custom MCP config files.
python_sdk_example.py
from smriti import SMRITI, SmritiConfig

# 1. Initialize SMRITI LTM Engine
config = SmritiConfig(
    storage_path="./agent_memory",
    llm_model="gpt-4o",
    openai_api_key="your-key-here"
)
memory = SMRITI(config=config)

# 2. Ingest observations (System 1 - Instant)
memory.encode("User prefers using PyTorch for neural networks.")
memory.encode("User is allergic to shellfish.", context="medical", private=True)

# 3. Query relevant context with dynamic retrieval
results = memory.recall("What framework does the user prefer?")
for mem in results:
    print(f"[{mem.strength:.2f}] {mem.content}")

# 4. Trigger System 2 Background Consolidation
memory.consolidate()
memory.save()
langchain_smriti.py
from langchain.agents import AgentExecutor, create_openai_tools_agent
from smriti.integrations.langchain import SmritiLangChainMemory

# Create SMRITI wrapper for LangChain
smriti_memory = SmritiLangChainMemory(
    storage_path="./langchain_ltm",
    api_key="your-key-here"
)

# Integrates directly into your LangChain Agent Executor
agent_executor = AgentExecutor(
    agent=agent,
    tools=tools,
    memory=smriti_memory, # SMRITI handles working & long-term memory
    verbose=True
)

Research & Blog

Featured dispatches on GPU memory dynamics, LLM inference determinism, and cognitive agent architectures.

Loading featured dispatches...