Agent memory,
at enterprise scale.
Agents spend most of their tokens re-reading what they already know. Hexian remembers it — store once, recall in milliseconds, send 95% fewer tokens — with every fact traced to its source.
One memory layer between your agents
and everything they touch.
Agents gather context, call tools, answer, and forget — then pay to do it all again. Hexian turns that repeated work into durable, governed memory that the whole fleet shares.
Durable agent memory
What an agent learns in one session is available in the next — across tools, users, and time.
Source-linked recall
Every memory keeps the document, ticket, message, or byte span behind it. Answers cite their basis.
Freshness-aware tool gating
Repeated read-only calls are answered from memory while fresh. Stale data and writes pass through.
Metered savings
Served tokens, saved tokens, avoided calls — reported per session, auditable by finance.
Hexian is the product.
Prism DB is the moat.
Memory that answers in milliseconds — with time travel, graph context, and byte-level provenance — is not possible on a stitched-together stack, or by wrapping someone else's database. So we built Prism DB from scratch: one canonical store doing the jobs of five systems at once, behind a single query. That engine is the reason Hexian works — and the reason it's state of the art.
semantic similarity over embeddings
BM25 + phrase search
multi-hop relationship traversal
keys, IDs, structured fields
facts as of any point in time
byte-span source provenance
numeric and date windows
nested-object identity
geo lookups
usage and savings metering
Measured, not promised.
Performance is tracked continuously on real product workloads — recall, ingest, vector, graph, and full agent sessions. Single node, commodity hardware.
Memory that understands when things change.
When new information contradicts the graph, Hexian supersedes the old fact — and keeps it as history. Ask what's true now, or what was true on any past date. Every fact traces back to the exact source that produced it.
Accurate. Efficient. Auditable.
On BEAM 100K — a public long-memory benchmark over ~100,000-token conversations — a reader model answered from Prism DB's evidence packs instead of full replay. Scored by BEAM's official evaluator, unmodified.
| Score | 71.7 · BEAM-canonical macro-average |
| Context tokens | −95.68% vs full-conversation replay |
| Evidence | 100% source-linked, SHA-256 run manifest |
BEAM 100K TIER · THIS HARNESS CONFIGURATION · NO CLAIMS ABOUT OTHER TIERS OR BENCHMARKS
Recall first. Run tools only when needed.
Agent asks
The agent needs context for a customer, ticket, codebase, or workflow — the same context it needed yesterday.
Hexian recalls
Exact, text, vector, graph, and temporal memory retrieve the relevant facts — each one carrying the source it came from.
Freshness gates the tools
Read-only calls with fresh answers are served from memory and counted. Stale fields and writes pass through to your systems.
Model gets evidence
The model sees compact, sourced context under a hard token budget — and the session gets a receipt for what was saved.
Five minutes to first recall.
One MCP config block connects Claude Code, Cursor, or your own platform. No SDK rewrite, no schema design, no pipeline to babysit — and savings are metered from the first session.
{
"mcpServers": {
"hexian": {
"command": "python3",
"args": ["-m", "prism_client.mcp_server"],
"env": { "PRISM_URL": "http://127.0.0.1:8080" }
}
}
}Pick an agent. Watch the bill drop.
The same task, run with and without memory. Hexian pays back fastest where agents repeatedly touch long-lived context: customers, tickets, codebases, and company knowledge.
Resolve ticket #4821 — recurring SLA escalation
on this task
More accurate, not just cheaper — Answers cite the case history and the prior fix — every fact source-linked.
This one measured session saved 11,268 tokens (−89.6%). Across 1,000 agents running 40 sessions a day, that rate is ≈ 450M tokens saved daily — about $494,000 a year at $3.00/1M input tokens.
Estimate the cost of forgetting.
Input-token spend scales with sessions × repeated context. Hexian reduces the repeated-context side and reports the savings — put your fleet's numbers in.
ASSUMES $3.00/1M INPUT TOKENS · 85% REDUCTION — MIDPOINT OF THE MEASURED 80–90% RANGE
Governed at the substrate.
Authorization, retention, audit, and provenance live in the engine — not bolted on around it. The enterprise question isn't whether agents can remember; it's whether you can inspect, isolate, and delete what they remember.
Isolation in the planner
Tenant, namespace, and scope prune before any index is touched. One customer's memory is structurally invisible to another's — not filtered at the edge.
Audit and provenance
Every fact traces to the source that produced it. When an agent acts, the basis for the action is enumerable.
Retention and deletion
TTLs, tombstones, quarantine-before-delete, and rebuildable indexes. Right-to-forget is a workflow, not an incident.
Your VPC, your keys
A single binary with zero runtime dependencies. Compute, data, and keys stay inside your boundary — no hosted memory silo.
Hexian ships as one binary with zero runtime dependencies. Your VPC holds compute, data, and keys — the trust boundary never leaves your account.
Stop paying your agents to re-learn what they already know.
95%+ less context. Millisecond recall. Every fact traceable to its source — and the savings metered on every session.

