hindsight
Agent memory server with retain, recall and reflect operations
Hindsight stores agent memories in banks and extracts facts, entities and timestamps from retained text using an LLM. Recall runs semantic, BM25, graph and temporal retrieval in parallel, then reranks the merged results. Background jobs consolidate facts into observations and mental models. It exposes a REST API, Python, Node.js and Go clients, a CLI, and a per-bank MCP endpoint.
Strengths
- Four parallel retrieval strategies merged with reciprocal rank fusion and cross-encoder reranking
- Works with 25+ LLM providers, including local ollama, lmstudio and llamacpp
- Built-in MCP endpoint per bank, plus 60+ listed integrations
- Embedded mode runs in-process via pip with a bundled pg0 database
Weaknesses
- Every retain call requires an LLM, adding cost and latency
- Accuracy claims are the vendor's own benchmarks; independent reproduction is partial
- Managed Cloud and Enterprise tiers exist; feature differences from self-hosted are unclear
- Minimum RAM and GPU needs are not stated in the README
- Docker
- Needs LLM provider (hosted or local), PostgreSQL (embedded pg0 by default), Oracle AI Database (optional)
- Models: openai, anthropic, gemini, groq, bedrock
- port 9999