Get started with Aura
Aura gives your AI agent persistent memory — no required embedding or LLM API for memory operations. The Rust core runs locally; new-query recall measured 11.61 ms median on 5,000 records in one local test.
Available since v1.58.0
Immutable evidence lineage, deterministic context capsules, and observable recall outcomes make agent context verifiable, bounded, and easier to operate in production.
Quickstart
Install
pip install aura-memoryPrebuilt wheels: Linux and macOS for Python 3.9–3.13, Windows for 3.10–3.13. No system dependencies.
Your first memory in 30 seconds
from aura import Aura
brain = Aura("./my_brain")
# Store knowledge
brain.store("Deploy to staging first. Never push straight to prod.")
brain.store("Our database is PostgreSQL. Auth uses JWT, 15-minute tokens.")
brain.store("User prefers concise answers without bullet points.")
# Recall relevant context locally; latency depends on the stored corpus
results = brain.recall_structured("deployment rules", top_k=3)
for r in results:
print(r["content"])
# Output:
# Deploy to staging first. Never push straight to prod.The brain persists to disk. Next time you run the script — it remembers everything.
Add memory to any LLM
from aura import Aura
import anthropic
brain = Aura("./brain")
client = anthropic.Anthropic() # ANTHROPIC_API_KEY
def chat(message: str) -> str:
# Pull relevant memories before each message
context = brain.recall(message, token_budget=1500)
response = client.messages.create(
model="claude-haiku-4-5-20251001",
max_tokens=1024,
system=f"You are a helpful assistant.\n\nWhat you remember:\n{context}",
messages=[{"role": "user", "content": message}]
)
brain.store(f"User asked: {message}")
return response.content[0].text
# Teach the agent
brain.store("The user is building a FastAPI service with PostgreSQL.")
brain.store("The user prefers async endpoints.")
print(chat("What database should I use?"))
# → Answers with PostgreSQL from memoryCore Concepts
What is Aura?
Aura is a cognitive memory layer that runs alongside your AI agent. It stores what the agent learns, organizes it automatically, and surfaces the most relevant context when the agent needs to answer a question.
Unlike a vector database, Aura doesn't just store and retrieve. It forms beliefs from accumulated evidence, discovers causal patterns over time, and generates advisory hints — all without any LLM calls.
The brain
Everything starts with Aura("./path"). This creates or opens a brain at the given directory. The brain stores everything to disk — it persists between script runs, server restarts, and deployments.
from aura import Aura
brain = Aura("./my_agent_brain")
# Creates: my_agent_brain/brain.cog, beliefs.cog, etc.Store and Recall
Two operations you'll use constantly:
# Store — write something to memory
brain.store("content", tags=["optional", "tags"])
# Recall as ready-to-use context for a prompt (a string)
context = brain.recall("your query", token_budget=1500)
# Or as ranked records you can inspect
results = brain.recall_structured("your query", top_k=5)
for r in results:
print(r["content"]) # the stored text
print(r["score"]) # relevance score
print(r["source_type"]) # recorded / retrieved / inferred / generated
print(r["tags"]) # associated tagsHow memory becomes knowledge
Aura automatically organizes stored records into higher layers during maintenance cycles:
Run brain.run_maintenance() to trigger a cycle. In production, use the background daemon.
Integrations
Aura works with any LLM framework. The pattern is always the same: recall before the LLM call, store after.
Claude (Anthropic)
from aura import Aura
import anthropic
brain = Aura("./brain")
client = anthropic.Anthropic()
context = brain.recall(user_message, token_budget=1500)
response = client.messages.create(
model="claude-haiku-4-5-20251001",
max_tokens=1024,
system=f"Assistant with memory:\n{context}",
messages=[{"role": "user", "content": user_message}]
)Gemini (Google)
from aura import Aura
import google.generativeai as genai
brain = Aura("./brain")
genai.configure(api_key="YOUR_KEY")
model = genai.GenerativeModel("gemini-2.5-flash-lite")
context = brain.recall(user_message, token_budget=1500)
prompt = f"Memory:\n{context}\n\nUser: {user_message}"
response = model.generate_content(prompt)
brain.store(f"User asked: {user_message}")Ollama (local, no API key)
from aura import Aura
import requests
brain = Aura("./brain")
context = brain.recall(user_message, token_budget=1500)
r = requests.post("http://localhost:11434/api/generate", json={
"model": "gemma3n:e4b",
"prompt": f"Memory:\n{context}\n\nUser: {user_message}",
"stream": False
})
brain.store(f"User asked: {user_message}")CrewAI
from aura import Aura
from crewai.tools import tool
brain = Aura("./brain")
@tool("remember")
def remember(content: str) -> str:
"""Store important information in long-term memory."""
brain.store(content)
return "Stored."
@tool("recall")
def recall_memory(query: str) -> str:
"""Search long-term memory for relevant information."""
return brain.recall(query, token_budget=1000) or "Nothing found."Memory Layers
When you store something, you can specify how long it should persist. Aura uses four levels — memories decay naturally and promote based on how often they're accessed.
| Level | Lifespan | Use for |
|---|---|---|
| Level.Working | Hours | Current session context, temporary notes |
| Level.Decisions | Days | Decisions made, tasks in progress |
| Level.Domain | Weeks | Project knowledge, team preferences |
| Level.Identity | Months+ | Core user traits, permanent rules |
from aura import Aura, Level
brain = Aura("./brain")
# Persists for months — core identity
brain.store("User is a senior backend engineer", level=Level.Identity)
# Persists for weeks — project knowledge
brain.store("We use PostgreSQL, not MySQL", level=Level.Domain)
# Persists for days — active work
brain.store("Working on auth module this week", level=Level.Decisions)
# Default — current session
brain.store("User asked about deployment just now")API Reference
Core methods
brain.store(content, level?, tags?, source_type?, namespace?)Store a memory. Returns the record ID.
brain.recall(query, token_budget?, min_strength?, namespace?, format?)Recall relevant memories as one context string, ready to put into a prompt. Since 1.61 it keeps first-hand memory apart from fenced, quoted outside text; format="levels" returns the older level-grouped text.
brain.recall_structured(query, top_k?, min_strength?, namespace?)Recall relevant memories as a ranked list of dicts (id, content, score, level, tags, source_type, created_at).
brain.explain_recall(query)Why each memory was or was not surfaced for a query.
brain.run_maintenance()Run a maintenance cycle. Forms beliefs, concepts, causal patterns, policy hints from stored records.
brain.get_surfaced_policy_hints(limit?)Get advisory hints derived from accumulated memory patterns.
Record fields
r = brain.recall_structured("something", top_k=1)[0]
r["id"] # unique record ID
r["content"] # stored text
r["level"] # Working / Decisions / Domain / Identity
r["tags"] # list of tags
r["source_type"] # recorded / retrieved / inferred / generated
r["strength"] # current strength (decays, refreshed by use)
r["created_at"] # unix timestampMCP server
Aura includes an MCP server (stdio) for Claude Desktop, Cursor, VS Code and any MCP client, and an HTTP server for Make.com and n8n. The HTTP server listens on 127.0.0.1 by default and refuses other addresses without an API key; install it with pip install "aura-memory[http]".
# MCP server over stdio (point your MCP client at this command)
python -m aura mcp ./my_brain
# HTTP server for Make.com, n8n and remote calls
pip install "aura-memory[http]"
python -m aura serve ./my_brain --api-key YOUR_KEYExamples
Ready-to-run scripts for every major use case. Clone the repo and run directly.
basic_usage.pyStore, recall, levels, tags — full basics walkthrough
claude_sdk_agent.pySystem prompt injection + tool use with Claude
gemini_aura_demo.pyCheap model + memory vs expensive model alone
ollama_agent.pyFully local — no API key required
crewai_agent.pyMulti-agent crew with persistent memory tools
langchain_agent.pyDrop-in memory via prompt injection
fastapi_middleware.pyPer-user memory isolation in a web API
research_bot.pyResearch agent that accumulates findings
maintenance_daemon.pyBackground brain maintenance in production
encryption.pyChaCha20 encrypted brain at rest
Try in browser — no install needed
Open the Colab notebook and run Aura in your browser in under 2 minutes.
Open in Colab