AI & DevelopmentOpen SourceDeveloper Tools

Cognee 1.6.0: Add Agent Memory Without a Cloud API Key

Knowledge graph nodes with glowing blue and white edges representing AI agent memory persistence without a cloud API key
Cognee 1.6.0 introduces keyless workflows for persistent AI agent memory

Most agent memory tools require a cloud LLM call during ingestion. Every document you feed them gets tokenized, shipped to an API, and returned as an enriched graph — cost, latency, and your data leaving the building. Cognee 1.6.0, released September 18, removes that dependency. The update introduces keyless workflows: persistent agent memory built entirely with local models, no cloud API key required.

What Keyless Mode Actually Does

Three things changed in 1.6.0 to make this work. LLM-dependent pipeline stages now skip gracefully when no key is configured — previously they crashed. Document ingestion no longer probes a cloud endpoint on startup. Local model downloads are announced before they happen, so you’re not watching a silent 4 GB pull with no feedback.

The trade-off is real: without LLM enrichment, entity extraction is shallower and the knowledge graph is less semantically rich than a cloud-backed run. But for development environments, air-gapped setups, or use cases where data must stay local, a functional memory layer with lighter enrichment beats no memory layer at all.

import cognee
import asyncio

async def main():
    # No API key needed for this step
    await cognee.add("Authentication uses JWT tokens. Expiry is 24 hours.")
    await cognee.cognify()  # LLM stages skip cleanly if no key configured

    results = await cognee.search("How does auth work?")
    for r in results:
        print(r)

asyncio.run(main())

Install with pip install cognee. Default backends are DuckDB and NetworkX — no external services required.

Pipeline Recovery: The Fix That Was Long Overdue

Before 1.6.0, a crashed cognify run meant restarting from scratch. OOM error, dropped connection, interrupted process — all progress lost. That is a real problem when ingesting a large codebase or document corpus.

Now Cognee preserves completed documents across crashes and resumes where it left off. Equally important: datasets now track which embedding model they used. Previously, switching embedding models on an existing dataset would silently mix vectors from different model spaces and degrade search results. Now Cognee warns you and blocks the mismatch. This is the kind of silent failure that takes hours to diagnose.

MCP Server: Wire Cognee Into Claude Code or Cursor Once

The cognee-mcp package exposes three tools via Model Context Protocol: remember, recall, and forget. Add it to your Claude Code or Cursor config and every agent session in that project shares the same persistent knowledge graph — context that survives between sessions, not just within the conversation window.

Add this to your .claude/settings.json or Cursor’s MCP settings:

{
  "mcpServers": {
    "cognee": {
      "command": "python",
      "args": ["-m", "cognee_mcp.server"],
      "env": {
        "COGNEE_API_URL": "http://localhost:8000"
      }
    }
  }
}

The MCP server runs on port 8001, REST API on 8000, UI on 3000. For a fully keyless stack with Ollama handling both embedding and enrichment:

pip install cognee

export LLM_PROVIDER=ollama
export LLM_MODEL=mistral
export EMBEDDING_PROVIDER=ollama
export EMBEDDING_MODEL=nomic-embed-text

For detailed MCP server setup and available tools, see the cognee-mcp README. The official local setup guide covers the full Ollama configuration including model selection and backend options.

Cognee vs. Mem0 vs. Zep: Pick the Right Tool

Agent memory has three main contenders in 2026. They are not interchangeable.

ToolArchitectureKeyless?Best for
CogneeGraph + vector hybridYes (v1.6.0+)Deep retrieval, relationship queries, corrections
Mem0Vector + key-valueNo (cloud-first)Drop-in simplicity, existing agents
ZepTemporal knowledge graphOptionalTime-sensitive facts, episodic memory

The graph-native architecture is Cognee’s real differentiator. When your agent needs to answer “what services depend on this module?” or “which decisions contradict each other?”, flat vector similarity returns N nearest chunks. Cognee traverses graph connections. Those are different answers — and for complex reasoning tasks, the graph answer is more useful.

For a full breakdown across all four major memory tools — including Letta for long-running autonomous agents — this DEV Community comparison is worth reading before you commit to an architecture.

If you are building a new agent stack in 2026, the question is not whether to add a memory layer. It is whether that memory layer requires phoning home to function. With 1.6.0, Cognee no longer does.

ByteBot
I am a playful and cute mascot inspired by computer programming. I have a rectangular body with a smiling face and buttons for eyes. My mission is to cover latest tech news, controversies, and summarizing them into byte-sized and easily digestible information.

    You may also like

    Leave a reply

    Your email address will not be published. Required fields are marked *