The SamurAI · for a Middle East cybersecurity client · 2026
Market-intelligence system
A multi-agent system that reads the market in English and Arabic and proposes leads to a human. One graph, one door, one human.
- My role
- Architecture, governance design and implementation
- Period
- 2026, pilot in progress
Stack
- Python
- Pydantic
- SurrealDB
- LLM gateway
- OpenAI
- Claude
- MCP
- Docker
- GitHub Actions
- Evals
- NIST AI RMF
- ISO 42001
On the resume
- Designing and building a multi-agent market-intelligence system for a Middle East cybersecurity client: bilingual (English/Arabic) data ingestion, classification and schema-enforced LLM extraction into a SurrealDB knowledge graph, with human review before any CRM write.
- Designing its AI governance: an AI gateway for every model and tool call (credential, budget, trace), controls mapped to NIST AI RMF and ISO/IEC 42001, and evaluation metrics on a labelled set run on every change.
The problem
The client sells cybersecurity and AI assurance services in the Gulf. Every day regulators act, companies get breached, tenders open, laws get deadlines. Nobody can read all of it, and the question is narrow: which companies in our territory just did something that means they need us, and why?
A scraper fails because it produces pages, not decisions. A chatbot fails because it can invent, cannot be audited, and would push unverified claims into the CRM. Three failure modes shaped the design: a model that invents, one company that exists twice in the data, and controls that live only in a prompt.
What I built
The first agent is a fixed pipeline, not a reasoning loop. Allow-listed sources are fetched and normalised in English and Arabic. One LLM call per article returns either "not relevant" or a typed card: company as named, event type and pillar from closed lists, event date, a verbatim quote, the reason to call, a confidence. Code then verifies the quote exists in the article; if it does not, the card is quarantined whatever the model said.
Verified cards are resolved against the account graph and land in a review queue where one named reviewer approves, edits or rejects with a reason. Only an approval triggers the CRM write, through a small tool that holds the one integration credential. The agent never holds it.
Status in October 2026: ingestion and normalisation run on real sources; extraction and verification are built and under review; company resolution, the review queue and the CRM write are the next steps; a labelled evaluation set is being assembled from reviewer decisions.
How it works
Plays on its own. Click a step to pause.
Decisions and trade-offs
- One store for documents, graph and vectors
- Instead of a vector database plus a graph database plus a document store: one schema, one client, one backup, and a loop that needs graph hops a vector store cannot express. The cost is a young engine, so load-testing comes before calling it production.
- Every relationship is an edge, never a field
- A fact stored twice eventually disagrees. Match metadata (exact, alias, similarity, human) lives on the edge itself.
- An LLM classifies and extracts for the pilot
- Training a small classifier first would have needed labelled data we did not have. The pilot produces the first real card weeks earlier, and reviewer decisions become the labelled set. The cost: higher per item, and LLM confidence is not calibrated.
- One scheduled command instead of a workflow engine
- About twenty sources, a straight line, every step a testable function. The human wait sits outside the pipeline. An engine comes back if branches or several human waits appear.
- Code asks for a model nickname, never a provider
- The gateway maps the nickname to a model with a fallback route. The model is chosen by how it scores on labelled cards, not by brand.
- Exactly one thin UI
- The review queue is the only screen built for the pilot. Everything else reuses existing tools.
Governance
- One door: the AI gateway
- The application holds one gateway key; provider keys live only in the gateway. It maps nicknames to models with fallback, retries once, times out, counts tokens against a per-run budget, caps answer length and records the real model behind every call. A boundary test enforces that only one file in the codebase may call a model.
- NIST AI RMF
- Govern: an owner per agent and one decision record per choice, with controls enforced at the gateway rather than in prompts. Map: fetched text treated as untrusted data, closed lists for event types and pillars, a risk register. Measure: the evaluation set and its metrics. Manage: human gates on input and output, quarantine, and stop conditions on steps and tokens.
- ISO/IEC 42001
- Policy as data (the source table, the pillar table). An audit trail in the graph from card to item to source, with prompt version, model version, source-list fingerprint, reviewer and reason. Version fields on every node.
- Regional data rules
- Regional AI and data-protection guidance mapped to human oversight, an in-country data zone and gated enrichment of personal data.
- Evaluation on every change
- A versioned set of labelled cards is re-run whenever the prompt, the model or the source list changes. Metrics: precision, match accuracy, cost per run, with recall and faithfulness (is the quote really in the text?) proposed as additions.
What I learned
- AI sits in two boxes, a human sits in one, and ordinary code and a database do everything else. Most of the reliability came from the ordinary code.
- Schema enforcement has to live at three levels: database asserts, typed models that mirror them with a drift test, and extraction that returns the schema or nothing.
- Real sources misbehave: bot protection, JavaScript shells with no links, PDFs, unreliable page dates. Quarantine with a reason beats silently skipping.
- A database client that does not raise on a failed statement inside a transaction will teach you to read every statement's status. Once.
Built inside a private company repository. Names, credentials and internal identifiers are left out on purpose.
Next case study
Dawn