Keep fastmemory's topology in PostgreSQL + Apache AGE, its vectors in pgvector and your models on dedicated GPU servers. Each agent hydrates a small, relevant working set. The design aims to keep decisions over thousands to millions of intents, entities, tools or policies fast, with every answer citing its source (architecture; latency at that scale not yet measured). No prompt has to hold them all.
| Store | PostgreSQL 17.7 · Apache AGE 1.7.0 · pgvector 0.8.0 · pg_trgm |
| Models | Laya decision model + MiniLM embedder, via NVIDIA Triton configs checked against the ONNX files |
| Orchestration | Docker Compose (validated) · Kubernetes (schema-validated) |
| Engine | MahaBodi Rust core; Python, Node, Java, C# and Go bindings |
| License | MIT (MahaBodi) · Apache-2.0 (Laya, MiniLM) · MIT (fastmemory) |
Ingestion, storage, models and agents are separate, so you scale each on its own and each fails on its own.
Chunk documents into ATFs, extract graph edges and embed passages. Writes go to one namespace at a time.
Relational passages with full-text search, a trigram vocabulary, pgvector HNSW, and one AGE graph per namespace.
Laya and MiniLM on Triton. Tokenisation stays in the client, so the server only runs the fused graph.
Hybrid SQL + Cypher search, then hydrate a working set into in-process MahaBodi and decide.
Start embedded, then move to a shared store and sharding when your data grows. The dashed arrow in the diagram marks the remote-model client, which is not built yet (see Status).
One process with memory in RAM and JSON snapshots. This is what exists today.
One PostgreSQL + AGE + pgvector node and a Triton server. Docker Compose provided.
Application-level sharding: a namespace lives on one shard, and a router maps namespaces to shards. Kubernetes manifests provided.
Search the shared store, pull the Cypher neighbourhood, restore it into in-process MahaBodi, then decide. Tested end to end.
Run models in process, as a sidecar on the node, or from a central GPU pool with autoscaling. Configs match the real models.
An enterprise buyer should know exactly where the edge is.
| Piece | Status | Evidence |
|---|---|---|
| MahaBodi engine and six language bindings | exists | full test suite across all languages passes on macOS x86_64 and Ubuntu |
| PostgreSQL + AGE + pgvector schema, sync, hybrid search, hydration | tested | end-to-end test: source found in the top 5 for 20/20 probes; hydrated engine answered 20/20 |
| Triton model configs | checked | 12/12 inputs and outputs match the ONNX files; Triton not run |
| Docker Compose / Kubernetes | validated | docker compose config; kubeconform strict: 7/7 valid; not deployed |
| Postgres storage backend inside the engine | not built | today, sync goes through snapshot/hydrate |
| Remote model client (Triton gRPC) | not built | models run in process today |
| Measured scale test at 10⁶+ records | not run | sizing is extrapolated, and labelled as such |
We found and fixed these before calling anything production-ready.