Hindsight, the open-source agent memory system from Vectorize.io, hit #1 on GitHub Trending last week with 38,000 stars. On September 25, it became more important — and if you have been running it inside Hermes Agent, it also became slightly broken. Nous Research moved Hindsight out of the Hermes core codebase and into its own standalone plugin, and that change quietly altered how updates work. Developers who installed Hindsight between September 14 and September 21 are running a version with a known event-loop bug. The fix takes one command.
What Just Changed
Hindsight used to live inside the Hermes repository at hermes-agent/plugins/memory/hindsight/. As of September 23, that directory is gone. The plugin now lives in the Hindsight repository under hindsight-integrations/hermes/, maintained by the Vectorize team rather than Nous Research. The official migration announcement is clear: “Memory providers are moving out of the Hermes tree into their maintainers’ own repositories, published through the plugin catalog.”
The migration is automatic for most users — running hermes update or starting an agent installs the plugin and preserves all existing configurations. Memory bank IDs, API keys, settings, and tool names remain untouched. However, there is one exception: if you have security.allow_lazy_installs: false in your config, you need to run the install manually.
The structural reason this happened is worth understanding. Memory providers move fast. The Hindsight team ships bug fixes and features on their own schedule, and being bundled inside Hermes meant waiting for Nous Research’s release cycle every time. Moving to the plugin catalog decouples the two, so a memory bug gets a same-day patch instead of queuing behind a framework release. Moreover, it makes version pinning possible — teams can freeze their memory provider version independently of the framework. Other memory providers will follow the same path.
Three Things to Do Right Now
1. Fix the embedded mode bug if you installed in mid-September. hindsight-embed 0.10.0, which covers installs from September 14 through September 21, has an event-loop defect that breaks local_embedded mode entirely. Calls hang and timeout silently. Update to 0.10.1:
hermes plugins update hindsight
2. Change how you update Hindsight going forward. hermes update no longer touches the plugin. From now on, the command you want is:
hermes plugins update hindsight
3. Check your recall behavior. The plugin changed the default recall response to return observations only — consolidated, deduplicated beliefs — rather than all memory types. If your agent relies on getting raw world facts and experience records back from recall, restore the previous behavior by adding to ~/.hermes/hindsight/config.json:
# In ~/.hermes/hindsight/config.json
"recall_types": "observation,world,experience"
For new installations, the flow is now two commands as described in the Hermes Agent memory provider docs:
hermes plugins install hindsight
hermes memory setup
Why Hindsight Agent Memory Leads on Benchmarks
The benchmark numbers explain the 38,000 GitHub stars. On LongMemEval — the standard evaluation for multi-session agent memory — Hindsight using Gemini-3 Pro scores 91.4 percent. That is the first time any agent memory system has crossed the 90 percent threshold on that evaluation. The nearest established competitor, Supermemory, scores 84.6 percent. Zep lands at 71.2 percent.
The number that should give pause is Mem0 at 49 percent. The full-context baseline — no memory system at all, just cramming everything into the context window — scores 60.2 percent. A poorly designed memory layer actively makes agents worse. That gap is why architecture matters here, and it is why the other agent memory systems have been competing hard on this benchmark.
| System | LongMemEval Score | License | Self-hosted |
|---|---|---|---|
| Hindsight | 91.4% | MIT | Yes |
| Supermemory | 84.6% | Proprietary | No |
| Letta | 83.2% | Apache 2.0 | Yes |
| Zep | 71.2% | Proprietary | Partial |
| Mem0 | 49% | Apache 2.0 | Yes |
| No memory (full context) | 60.2% | — | — |
How the Architecture Works
Most agent memory systems treat memory as a retrieval layer — you store embeddings and run similarity search. Hindsight takes a different approach. It organizes memories into four distinct networks: world facts, agent experiences, observations (consolidated beliefs with evidence tracking), and mental models (synthesized understanding across many memories). Each type answers a different kind of query, and each gets written to and read from independently.
The recall subsystem, called TEMPR, runs four parallel search strategies simultaneously: semantic similarity, keyword matching via BM25, graph traversal for entity relationships, and temporal filtering. Results are fused through reciprocal rank fusion and cross-encoder reranking. Furthermore, this is why Hindsight handles multi-hop reasoning and temporal queries — “what did the user prefer two months ago about deployment environments” — where pure vector search fails.
The reflection operation goes further still. It reasons over the full memory bank using configurable disposition traits — how skeptical the agent should be about new claims, how much weight to give recent versus older information. Additionally, you can configure the agent’s mission and directives, shaping how it forms beliefs, not just what it stores.
What This Signals About Agent Infrastructure
This plugin migration is the right architectural call. When Hermes Agent launched, bundling memory providers reduced setup friction. That model does not scale. Agent infrastructure is maturing into distinct layers: the framework, the memory system, the tooling surface. Each moves at its own pace and is maintained by different teams with different priorities.
The pattern mirrors how npm separated from Node.js, or how Rust’s crate ecosystem works: the runtime ships the core, and ecosystem components live in their own repositories on their own release schedules. Hindsight is the first memory provider to make this move in Hermes. It will not be the last. Therefore, if you are building agents with any framework that bundles memory tightly, expect this separation to happen there too — and build your configs to accommodate it when it does.
If you are building a production agent and have not evaluated your memory layer recently, the migration guide and benchmark numbers are worth an hour of your time. The 91.4 percent LongMemEval score matters — but so does having a maintainer who can ship a hotfix the same day a bug lands in production.













