Last issue we posed the one-metric challenge: pick one KPI to prove your context-engineered stack beats vanilla RAG. A standout community take said: classic QA sets miss persistent memory (we agree); use a blend of regression checks for consistency, scripted multi-session evals, and domain-specific memory probes - while watching latency/cost.
In this special cognee edition, you’ll find a brief evaluation update and an early look at features launching later this month. You’re hearing them here first.
Cognee organizes your data into AI memory - turning raw sources into a modular, queryable knowledge graph powered by embeddings so agents retrieve, reason, and remember with structure.
Setup: 45 runs, 24 HotPotQA questions; compared Cognee with Mem0, LightRAG, and Graphiti (numbers from a prior run).
Metrics: EM, F1, DeepEval correctness, and a “human-like correctness”.
Takeaway: Cognee showed consistent multi-hop gains; EM/F1 alone under-report memory quality—persistent memory needs cross-context and time-aware checks.
Why classic QA falls short for memory:
Surface ≠ sense. EM/F1 reward phrasing, not grounded reasoning.
LLMs are stochastic. One score ≠ truth; you need aggregates.
Context is too neat. Even HotPotQA assumes the answer lives in two tidy passages. Real memory spans files, meetings, timelines.
Stay tuned: Cognee is co-building a new, open evaluation dataset with DeepEval to test what memory systems actually do.
Read the full deep dive here.
These ship throughout the month. Follow the blog, X, and Discord for details, demos, and notebooks.
Semantic reasoning, structured memory, and intelligent query capabilities—without the need to host infrastructure or manage pipelines.
Intended for both solo developers and agile product teams, cogwit gives you a memory-enabled AI platform that grows with you.
Sign up for the beta with 1 GB ingestion + 10 000 API calls from platform.cognee.ai
After your Cognify build finishes, (optionally) Memify takes over as a modular, parameterized pipeline that safely enriches your graph DB, vector collections, and metastore—no rebuilds required.
Why it matters: plugin architecture, parallel execution, transaction-safe mutations—all designed for non-disruptive, incremental improvement.
Cognee learns from feedback and reinforces the exact graph edges used to answer a question.
Users react to answers during conversation.
Text is normalized to a −5…+5 score (configurable weights).
Scores are attributed to used_graph_element_to_answer edges from that interaction.
Edge weights accumulate over time (auditable sums), boosting high-quality paths without deleting anything.
Result: retrieval quality improves where it matters, transparently and safely.
New cognee UI, temporal pipeline, improved embeddings, Cogwit MCP for developers, BAML integration - and more.
Many integrations, notebooks, case studies are rolling out.
Here are some of the key events of the month:
Sep 2 — AI Agent Meetup Berlin: Context Engineering https://luma.com/cymt5oeq
Sep 3 — DuckDB Meets AI: When Analytics Powers AI Memory with Cognee https://luma.com/6s0goctt
Sep 4 — Redis Released (SF) https://events.redis.io/redis-released-san-francisco-2025
Sep 8 — Graphs with Neo4j
Sep 23–24 — AI Engineer Paris https://www.ai.engineer/paris
Sep 26 — Qdrant Vector Space Day https://luma.com/p7w9uqtz
Sep 29 — Memgraph and Cognee podcast
And a few more surprise ones! Stay tuned!
As always, office hours: every Friday at 17:00 CET - live Q&A with the founder. You are invited.
What’s the most recent task where your graph-backed memory beat vanilla RAG?
Reply on r/AIMemory or in Discord. Best answer gets featured next issue.
Forward this to a teammate still evaluating memory for their AI agents.
Let’s build memory that improves every day.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.