RetailboxMessage us
#15AI / Developer Tools202630/30 adversarial writes blocked in red-team testing

FleetMemory: Shared Memory for a Fleet of AI Agents

Retailbox R&D: CockroachDB x AWS Hackathon ("Build with Agentic Memory")

Retailbox R&D: CockroachDB x AWS Hackathon entry, submitted 2026

Our entry for the CockroachDB x AWS "Build with Agentic Memory" hackathon: two AI agents sharing one memory store, governed by a write-gate and an adversarial verifier that rejects facts nobody can source. Public repo, live demo, no login required.

LangGraphAWS Bedrock (Claude Haiku 4.5)AWS AgentCore MemoryCockroachDB (C-SPANN vector index)AWS Lambda
Demo video: FleetMemory: Shared Memory for a Fleet of AI Agents · 1:15 · Watch on YouTube

What was the problem?

AI agent fleets that share memory usually share garbage, too. Point two agents at the same store with no gate, and one bad write becomes "known" to every agent that reads it next: a hallucinated number, an unsourced claim, whatever slipped through. For this hackathon we set out to prove a fleet of agents could share memory and still refuse to believe something nobody actually said.

How did Retailbox solve it?

We built FleetMemory: two LangGraph agents, Alex (front-line) and Maria (account manager), reading and writing one shared memory store on CockroachDB. Every candidate fact passes a write-gate first (junk, duplicate, contradiction, and confidence checks) before it reaches the table. Anything that fails quarantines and goes to an adversarial LLM verifier (AWS Bedrock, Claude Haiku 4.5) whose default is reject; it weighs provenance, so a customer's own words can supersede a prior fact, but an unsourced guess can't. Corrections never delete: facts are bi-temporal, so a point-in-time read can reconstruct what the fleet believed last week even after it's been overturned. Recall runs two ways, structured fact lookup and semantic search over a CockroachDB C-SPANN vector index, both scoped so one customer's data never leaks into another's session. AWS AgentCore Memory handles short-term session state; Lambda hosts the public demo.

What does it look like?

FleetMemory demo screen: agent Alex's chat on the left, agent Maria's on the right, and the shared-memory panel in the middle, where a $1,800 budget card stamped Held for review sits above the $18,000 budget that passed the gate. (opens the full-size image in a new tab)
A rogue agent asserts a $1,800 budget. The gate holds it for review, directly above the $18,000 fact Alex learned from the customer; Maria still answers from $18,000.
The $1,800 budget card struck through and stamped Refused, with the verifier's reason: the value lacks a direct customer utterance and is unsourced agent-generated memory, while the existing fact comes from a direct customer statement. (opens the full-size image in a new tab)
The adversarial verifier refuses the claim and writes its one-sentence reason on the card: no customer statement behind $1,800, a direct one behind $18,000.
psql output against CockroachDB Cloud: seven active facts for the demo customer with confidence scores, then the gate_decisions journal with six clean accepts and two quarantined writes that the verifier rejected. (opens the full-size image in a new tab)
The same customer read straight from CockroachDB: seven active facts, and the decision journal with six clean accepts and two contradicting writes quarantined, then rejected by the verifier.
FleetMemory architecture: agents Alex and Maria (LangGraph) send every write through the write-gate; quarantined writes go to a verifier subagent that defaults to reject; facts, the C-SPANN vector index and the gate_decisions journal live in CockroachDB; AgentCore Memory, Bedrock and the Lambda demo run on AWS. (opens the full-size image in a new tab)
Architecture: every agent write goes through the gate, quarantined claims go to the verifier, and facts plus the decision journal live in CockroachDB. AWS runs session state (AgentCore Memory), the models (Bedrock) and the public demo (Lambda).

What were the results?

  • ✓Submitted to the CockroachDB x AWS "Build with Agentic Memory" hackathon on 2026-08-14, four days ahead of the deadline.
  • ✓Open-sourced under MIT with the write-decision journal, red-team script, and eval harness in the repo.
  • ✓Red-teamed 30 adversarial writes across 5 attack classes: the first run found a real gap, we closed it, and the re-run blocked 30/30 while still passing 9/10 legitimate customer-confirmed updates (p50 gate latency 293ms).
  • ✓10/10 on a self-run mini-LongMemEval covering extraction, multi-session recall, temporal correctness, and abstention.
  • ✓Live AWS demo, no login required: talk to agent Alex, come back as agent Maria, and watch a hallucinated fact get quarantined and rejected on the record.
  • ✓Found two CockroachDB incompatibilities in LangGraph's Postgres checkpointer and filed them upstream with a standalone repro and a fix direction (langchain-ai/langgraph#8620). An independent contributor reproduced both on LangGraph main and reports a fix passing the checkpoint-postgres test suites on PostgreSQL 17.5 and CockroachDB v26.2.4. The issue is open, pending maintainer review.

Have a similar problem?

30-minute call. Free. We'll tell you honestly if we can help.