Skip to content
Sarah Henia.
All work
The SamurAI

The SamurAI · in production since March 2026

Dawn

An autonomous project-management agent where nothing is written until a human approves it.

My role
Designed and built the whole system
Period
2026, in production since March

Stack

  • Python
  • SurrealDB
  • Redis
  • OpenRouter
  • Llama
  • Microsoft Graph
  • AWS
  • Linux
  • GitHub Actions
  • Logfire
  • Langfuse

On the resume

  • Designed and built Dawn, an autonomous AI agent for project management, live in production from March 2026: turns recorded meetings into tasks with owner, priority and due date, with human-in-the-loop approval before any write to Microsoft Planner.
  • Engineered its graph memory: a SurrealDB knowledge graph with hybrid RAG (vector search plus recency boost) and a Redis context cache (CAG), integrated with the Microsoft Graph API.
  • Designed its governance and security layer: identity gate, per-agent policies, HMAC-verified inbound messages, audit log, confidence gating, and sensitive data routed to a local Llama model.
  • Built multi-model routing through OpenRouter with automatic fallback; deployed on hardened AWS EC2 with GitHub Actions CI/CD and observability in Logfire and Langfuse.

The problem

A small consulting team records a lot of meetings. The commitments made in them were not reliably turning into tracked work: someone had to read the transcript, decide who owns what, and type it into Microsoft Planner. Follow-up depended on memory.

The brief was not "automate project management". It was narrower and harder: get the right tasks into Planner, with the right owner and date, without ever letting an AI write something nobody checked.

What I built

Dawn reads meeting transcripts, pulls in the team context it needs, extracts proposed tasks with an owner, a priority, a due date and a confidence score, and sends them to two oversight leads by email. They reply APPROVE or REJECT. Only then does Dawn create the tasks in Planner, store what happened in its graph memory, and start nudging assignees when work stalls.

Microsoft blocked app-only direct messages in Teams, so email became the approval channel. It shipped weeks earlier than waiting for admin access would have, and the leads liked it.

How it works

Plays on its own. Click a step to pause.

Decisions and trade-offs

Human approval before any write, instead of auto-create above a confidence level
Trust had to come first, and the blast radius of a bad task had to be zero. The early design auto-created high-confidence tasks; the shipped rule is that nothing reaches Planner without a reply.
CAG and RAG, not one or the other
Small stable facts (who is on the team, who is overloaded) are cached in Redis. Large history (what we decided in past meetings) is retrieved. They solve different problems.
Hybrid re-ranking instead of naive top-k retrieval
Similarity alone let old meetings outrank recent ones. Re-ranking on similarity and recency, with a hard threshold, keeps noise out of the prompt.
SurrealDB instead of a separate graph database plus a vector store
Documents, graph edges and vectors in one engine: one schema, one client, one backup. Edges are first-class, so "task extracted from meeting" is a relationship, not a field.
Direct API calls instead of an agent framework
Explicit control flow, grep-able AI calls, and state that is visible in Redis and SurrealDB rather than hidden in a chain.
Models as aliases behind a router
Swapping a provider is a one-line change. Sensitive content is forced to a local model. No model name is hardcoded in business logic.
Archive at 90 days instead of delete
History has pattern value and storage is cheap. Archived meetings leave the active context but stay queryable.

Governance

Identity gate
Only senders from the company domain can trigger actions. External senders get a redirect and no data. Sender trust levels decide what the agent will do.
Signed inbound messages
Every inbound Teams webhook request is verified with HMAC-SHA256. Invalid signatures are rejected before parsing.
Sensitive-data routing
Keywords tied to regulated or client-sensitive content force a local model and flag the item for review.
Audit trail
Every proposal keeps its reasoning, its confidence, who resolved it and when. Assignees never see confidence scores, model names or record identifiers.
Hardened host
Non-root service user, secrets outside the repository, encrypted volume, least-privilege IAM, SSH key only, automatic patching, no exposed bot port. Deployed by GitHub Actions; traced in Logfire and Langfuse.

What I learned

  • Every production bug was a distributed-systems classic: duplicates, a self-loop, lost updates. The fix was idempotency: stable identifiers, status flags and exactly-once writes, not a smarter prompt.
  • Human-in-the-loop is a state machine (pending, approved, rejected, duplicate, stalled), and it deserves the same design care as the model call.
  • A webhook that must answer in five seconds and a job that takes longer are two different programs. Separating them removed a whole class of silent crashes.
  • The best interface for approval was the one the leads already lived in. Email beat a custom UI.

Built inside a private company repository. Names, credentials and internal identifiers are left out on purpose.

Next case study

Collaboris