From one laptop to petabytes

Agent memory on PostgreSQL. Models where your GPUs are.

Keep fastmemory's topology in PostgreSQL + Apache AGE, its vectors in pgvector and your models on dedicated GPU servers. Each agent hydrates a small, relevant working set. The design aims to keep decisions over thousands to millions of intents, entities, tools or policies fast, with every answer citing its source (architecture; latency at that scale not yet measured). No prompt has to hold them all.

Tested stack

StorePostgreSQL 17.7 · Apache AGE 1.7.0 · pgvector 0.8.0 · pg_trgm
ModelsLaya decision model + MiniLM embedder, via NVIDIA Triton configs checked against the ONNX files
OrchestrationDocker Compose (validated) · Kubernetes (schema-validated)
EngineMahaBodi Rust core; Python, Node, Java, C# and Go bindings
LicenseMIT (MahaBodi) · Apache-2.0 (Laya, MiniLM) · MIT (fastmemory)

Four planes, scaled independently

Ingestion, storage, models and agents are separate, so you scale each on its own and each fails on its own.

01 · Ingest

Stateless workers

Chunk documents into ATFs, extract graph edges and embed passages. Writes go to one namespace at a time.

02 · Store

PostgreSQL shards

Relational passages with full-text search, a trigram vocabulary, pgvector HNSW, and one AGE graph per namespace.

03 · Models

GPU model servers

Laya and MiniLM on Triton. Tokenisation stays in the client, so the server only runs the fused graph.

04 · Serve

Agents

Hybrid SQL + Cypher search, then hydrate a working set into in-process MahaBodi and decide.

Documents Ingest workerschunk · ATF · embed PostgreSQL shard passages + FTS + pg_trgmpgvector (HNSW)Apache AGE graph / namespace AgentMahaBodi in process Model serverLaya · MiniLM hybrid searchhydrate

Five deployment patterns

Start embedded, then move to a shared store and sharding when your data grows. The dashed arrow in the diagram marks the remote-model client, which is not built yet (see Status).

P1

Embedded

One process with memory in RAM and JSON snapshots. This is what exists today.

Fits in one machine's RAM
P2

Store node + model server

One PostgreSQL + AGE + pgvector node and a Triton server. Docker Compose provided.

GBs to low TBs
P3

Sharded by namespace

Application-level sharding: a namespace lives on one shard, and a router maps namespaces to shards. Kubernetes manifests provided.

TBs to PBs
P4

Tiered working set

Search the shared store, pull the Cypher neighbourhood, restore it into in-process MahaBodi, then decide. Tested end to end.

Recommended for agents
P5

Model hosting

Run models in process, as a sidecar on the node, or from a central GPU pool with autoscaling. Configs match the real models.

Scale models independently

What exists, what's tested, what's next

An enterprise buyer should know exactly where the edge is.

PieceStatusEvidence
MahaBodi engine and six language bindingsexistsfull test suite across all languages passes on macOS x86_64 and Ubuntu
PostgreSQL + AGE + pgvector schema, sync, hybrid search, hydrationtestedend-to-end test: source found in the top 5 for 20/20 probes; hydrated engine answered 20/20
Triton model configschecked12/12 inputs and outputs match the ONNX files; Triton not run
Docker Compose / Kubernetesvalidateddocker compose config; kubeconform strict: 7/7 valid; not deployed
Postgres storage backend inside the enginenot builttoday, sync goes through snapshot/hydrate
Remote model client (Triton gRPC)not builtmodels run in process today
Measured scale test at 10⁶+ recordsnot runsizing is extrapolated, and labelled as such

Lessons from testing at scale

We found and fixed these before calling anything production-ready.

AGE needs explicit indexesWithout GIN indexes on properties, property matches are sequential scans, and a 2,000-document load had not finished after 8 minutes. The schema now creates the indexes per namespace.
Hub concepts explode graph queriesCommon terms link thousands of records, and one 2-hop query ran for more than 20 minutes. Vertices now carry their degree, and traversal skips hubs.
Search quality falls as memory growsClean-query recall drops from beating BM25 at 300 paragraphs to a tie at 2,000. The design therefore keeps every query scoped to a namespace.
Grounding cuts confident errors, not all errorsOn BoolQ, answering from memory lifts accuracy from 0.42 to 0.78 and cuts confident errors from 45 % to 17 %. It still falls short of the oracle passage.