
Most agent memory tools require a cloud LLM call during ingestion. Every document you feed them gets tokenized, shipped to an API, and returned as an enriched graph — cost, latency, and your data leaving the building. Cognee 1.6.0, released September 18, removes that dependency. The update introduces keyless workflows: persistent agent memory built entirely with local models, no cloud API key required.
What Keyless Mode Actually Does
Three things changed in 1.6.0 to make this work. LLM-dependent pipeline stages now skip gracefully when no key is configured — previously they crashed. Document ingestion no longer probes a cloud endpoint on startup. Local model downloads are announced before they happen, so you’re not watching a silent 4 GB pull with no feedback.
The trade-off is real: without LLM enrichment, entity extraction is shallower and the knowledge graph is less semantically rich than a cloud-backed run. But for development environments, air-gapped setups, or use cases where data must stay local, a functional memory layer with lighter enrichment beats no memory layer at all.
import cognee
import asyncio
async def main():
# No API key needed for this step
await cognee.add("Authentication uses JWT tokens. Expiry is 24 hours.")
await cognee.cognify() # LLM stages skip cleanly if no key configured
results = await cognee.search("How does auth work?")
for r in results:
print(r)
asyncio.run(main())
Install with pip install cognee. Default backends are DuckDB and NetworkX — no external services required.
Pipeline Recovery: The Fix That Was Long Overdue
Before 1.6.0, a crashed cognify run meant restarting from scratch. OOM error, dropped connection, interrupted process — all progress lost. That is a real problem when ingesting a large codebase or document corpus.
Now Cognee preserves completed documents across crashes and resumes where it left off. Equally important: datasets now track which embedding model they used. Previously, switching embedding models on an existing dataset would silently mix vectors from different model spaces and degrade search results. Now Cognee warns you and blocks the mismatch. This is the kind of silent failure that takes hours to diagnose.
MCP Server: Wire Cognee Into Claude Code or Cursor Once
The cognee-mcp package exposes three tools via Model Context Protocol: remember, recall, and forget. Add it to your Claude Code or Cursor config and every agent session in that project shares the same persistent knowledge graph — context that survives between sessions, not just within the conversation window.
Add this to your .claude/settings.json or Cursor’s MCP settings:
{
"mcpServers": {
"cognee": {
"command": "python",
"args": ["-m", "cognee_mcp.server"],
"env": {
"COGNEE_API_URL": "http://localhost:8000"
}
}
}
}
The MCP server runs on port 8001, REST API on 8000, UI on 3000. For a fully keyless stack with Ollama handling both embedding and enrichment:
pip install cognee
export LLM_PROVIDER=ollama
export LLM_MODEL=mistral
export EMBEDDING_PROVIDER=ollama
export EMBEDDING_MODEL=nomic-embed-text
For detailed MCP server setup and available tools, see the cognee-mcp README. The official local setup guide covers the full Ollama configuration including model selection and backend options.
Cognee vs. Mem0 vs. Zep: Pick the Right Tool
Agent memory has three main contenders in 2026. They are not interchangeable.
| Tool | Architecture | Keyless? | Best for |
|---|---|---|---|
| Cognee | Graph + vector hybrid | Yes (v1.6.0+) | Deep retrieval, relationship queries, corrections |
| Mem0 | Vector + key-value | No (cloud-first) | Drop-in simplicity, existing agents |
| Zep | Temporal knowledge graph | Optional | Time-sensitive facts, episodic memory |
The graph-native architecture is Cognee’s real differentiator. When your agent needs to answer “what services depend on this module?” or “which decisions contradict each other?”, flat vector similarity returns N nearest chunks. Cognee traverses graph connections. Those are different answers — and for complex reasoning tasks, the graph answer is more useful.
For a full breakdown across all four major memory tools — including Letta for long-running autonomous agents — this DEV Community comparison is worth reading before you commit to an architecture.
If you are building a new agent stack in 2026, the question is not whether to add a memory layer. It is whether that memory layer requires phoning home to function. With 1.6.0, Cognee no longer does.













