Build agents with a System 1
Use it where an agent picks among many options (intents, tools, known issues, policy answers) and prompting an LLM with every option stops scaling. MahaBodi is one Rust engine with one JSON API, identical across Rust, Python, Node.js, Java, C#/.NET and Go. It stores memory as a fastmemory topology, finds things again with a hybrid search cascade, and decides with Laya's typed decision models. Every answer says whether it matched, how, and whether to hand off.
Build
Pre-release: build from source. The Laya model is exported once to ONNX (about 1.7 GB), and the export is verified against Laya's PyTorch output.
Rust 1.88+ · libonnxruntime 1.23+ · Python 3.11python research/export_onnx.py --out models/laya-v2./scripts/test_all.shQuickstart
from mahabodi import Bodi
b = Bodi()
b.ingest(open("kb.md").read(), source="kb") # memory
r = b.query("refund policy", k=5) # r["matched"], r["handoff"], r["stage"], r["hits"]
b.load_laya("models/laya-v2") # decisions
a = b.decide({"message": "I was charged twice"},
{"refund": {"type": "noul", "instructions": "Does `message` ask for a refund?"}})
g = b.decide_with_memory({"question": "Can I get my money back after 30 days?"},
{"answer": {"type": "noul", "instructions": "Based on `passage`, is the answer to `question` yes?"}},
style="passages") # grounded in memory, cites its hits
const { Bodi } = require('mahabodi');
const b = new Bodi();
b.ingest(fs.readFileSync('kb.md', 'utf8'), { source: 'kb' });
const r = b.query('refund policy', 5);
b.loadLaya('models/laya-v2');
const a = await b.decide({ message: 'I was charged twice' },
{ refund: { type: 'noul', instructions: 'Does `message` ask for a refund?' } }); // off the event loop
try (Bodi b = new Bodi()) {
b.ingest(Files.readString(Path.of("kb.md")), "kb");
String r = b.query("refund policy", 5); // JSON in, JSON out
b.loadLaya("models/laya-v2");
String a = b.decide("{\"message\":\"I was charged twice\"}",
"{\"refund\":{\"type\":\"noul\",\"instructions\":\"Does `message` ask for a refund?\"}}");
}
using MahaBodi;
using var b = new Bodi();
b.Ingest(File.ReadAllText("kb.md"), source: "kb");
var r = b.Query("refund policy", 5);
b.LoadLaya("models/laya-v2");
var a = await b.DecideAsync(JsonNode.Parse("{\"message\":\"I was charged twice\"}")!,
new JsonObject { ["refund"] = new JsonObject { ["type"] = "noul", ["instructions"] = "Does `message` ask for a refund?" } });
e, _ := mahabodi.New(nil)
defer e.Close()
e.Ingest(kb, "kb")
r, _ := e.Query("refund policy", 5)
e.LoadLaya("models/laya-v2")
a, _ := e.Decide(`{"message":"I was charged twice"}`,
`{"refund":{"type":"noul","instructions":"Does `+"`message`"+` ask for a refund?"}}`)
use mahabodi_core::Bodi;
let b = Bodi::from_json("")?;
b.call("ingest", &json!({"text": kb, "source": "kb"}))?;
let r = b.call("query", &json!({"q": "refund policy", "k": 5}))?;
b.call("load_laya", &json!({"dir": "models/laya-v2"}))?;
let a = b.call("decide", &json!({"state": {"message": "I was charged twice"},
"questions": {"refund": {"type": "noul", "instructions": "Does `message` ask for a refund?"}}}))?;
The snippets are illustrative: exact signatures are in each binding's source and tests (bindings/*).
API methods
Every binding sends these through the same JSON dispatcher, call(method, args).
| Method | What it does |
|---|---|
ingest, ingest_batch | Add documents (ATF markdown, fastmemory tags or prose). A batch rebuilds once. |
query, context, traverse | Hybrid search cascade, grounding text, and graph neighbourhood. |
density, ensure_density | Measure and repair the concept graph so every record stays reachable. |
load_laya, predict, decide, decide_batch | Typed decisions. predict is Laya-exact; decide adds tournament, gate and cache. |
decide_with_memory | Retrieves context, attaches it, decides, and reports which memory was used. |
learn, forget, load_embedder, embed_text | Experience memory from labelled cases, and the dense embedder. |
snapshot, restore, clear, stats | Persistence (JSON snapshot) and introspection. |
Result fields
| Field | Meaning |
|---|---|
matched | Something in memory actually matched the query. A hub fallback or empty memory gives false. |
handoff | System 1 shouldn't act on this: escalate to System 2 or ask. Treat it as "no answer". |
stage | How it was found: exact, substring, stem, fuzzy, dense or hub. |
confidence, term_coverage | Stage-based confidence, and the share of query terms that matched. |
answers[q].bodi | For decisions: strategy (direct or tournament), whether memory was used, and the handoff flag. |
Decision options
| Option | Default | Notes |
|---|---|---|
max_options_per_pass | 16 | Above this, tournament shortlisting runs (Banking77: 0.660 vs 0.492). |
experience_below_margin | 0.5 | Memory changes an answer only when Laya's top-2 margin is below this, so it never loses to Laya on the 6 suites tested. |
experience_override_agree, experience_override_min_trust | 6, 0.6 | Lets memory override a confident answer when at least 6 of 8 nearest stored cases agree and the task memory has proven reliable. Fresh test: no loss vs Laya; +8.2 points on Banking77 over the gate alone. |
experience_memory_first_margin | 0.2 | Only for tasks learned with learn(..., calibrate=N) (opt-in; N extra decisions). Memory answers first where its leave-one-out accuracy beats Laya's by this margin. Fresh test: never below Laya or kNN on 5 suites; ties kNN on Banking77. |
script_gate | true | Hands off non-Latin text sent to the English checkpoint. This is a subset of Laya's Router. |
cache | true | Answer cache; benchmarks run with it off. |
Reproduce the benchmarks
Every number in BENCHMARKS.md comes from research/results/*.json, which holds per-item predictions. It is regenerated by research/report.py and tested with exact McNemar.
.venv/bin/python research/bench.py --n 500 # zero-shot, 6 suites + Laya's mitigations
.venv/bin/python research/bench_massive.py --per-lang 100 # 51 languages
.venv/bin/python research/bench_grounding.py --style-file research/results/tune_grounding.json
.venv/bin/python research/report.py # regenerates BENCHMARKS.md
Known limits
oos_min_similarity (it needs load_embedder): in the CLINC150 benchmark it caught about 72 % of out-of-scope requests on fresh items and cost some in-scope accuracy (the built-in option reproduces the benchmark exactly: all 1,600 items on Ubuntu with the final build). Tune its threshold on data with your real out-of-scope rate.
Memory is in process, with a full rebuild per ingest batch.
It has been built and tested on macOS x86_64 and Ubuntu (all 8 test suites pass on both). Version 0.1.2 is published on crates.io, PyPI, npm and NuGet, and as Go module v0.1.2; prebuilt binaries cover Linux x86_64 (glibc 2.28 or newer) and macOS x86_64 only. Java is built from source; it is not on Maven Central yet.
BENCHMARKS.md lists every tie and loss.