1 · Clean purchase — autonomous approve
Clearbit Sample · lead-enrichment · €300
- Call protected vendor-risk endpoint
- Evaluate deterministic policy
- Jev System One fraud gate
- Issue X402 payment
- Command Code finance memo
- Persist to Supabase ledger
AgentPayOps
Live Architecture Demo
Hits the same production APIs the dashboard uses: X402 402 challenge, deterministic policy, Jev System One gates, Command Code memos, and the Supabase ledger. Nothing is mocked on this panel.
Clearbit Sample · lead-enrichment · €300
Metro Cloud Brokers · cloud-credits · €12,600
Veritas Risk Graph · vendor-risk-data · €180 + BEC context
Clearbit Sample · lead-enrichment · €300 (again)
Invoice text: new IBAN, 'don't call to verify', urgent discount
Tip: run it twice — flow 4 (and then flow 1) trips the duplicate guard because the first purchase is now in the ledger. That is the memory story. Reset anytime with npm run demo:reset.
Every button on the demo above hits these seven layers for real.
Layer 0
PDF, image, or text invoices are parsed: vendor, amount, category, due date, findings. Extraction uses Command Code vision with a deterministic fallback, so the pipeline never dies on a bad upload.
Rule: Garbage in must never become a payment.
On this page: Flow 5 on this page sends a fraud-flavored invoice through intake.
Layer 1
Plain code checks what code can see: vendor allow/block lists, category, amount ceilings, approval thresholds, and duplicate purchases against the persisted ledger. Finance edits these live in Payment Controls — no redeploy.
Rule: If a rule can decide it, a rule decides it. Free, instant, testable (accuracy 1.000, false-approve 0.000).
On this page: Flows 1, 2 and 4 are decided here alone.
Layer 2
typesafe/jev takes state + typed questions and answers with raw probabilities in ~1s for fractions of a cent. It reads the payment packet and documents for patterns rules cannot encode: lookalike domains, changed wire instructions, manufactured urgency, 'don't call to verify'.
Rule: Jev runs only AFTER rules approve, and can only TIGHTEN (approve → escalate). It never releases funds. Blocked paths skip it — zero cost on dead ends.
On this page: Flows 3 and 5 exist to show this gate firing.
Layer 3
stealth/space-bunny-alpha (fallbacks: xiaomi/mimo-v2.6-pro, meta/muse-spark-1.3-contributor) turns the decision into a CFO-ready memo: headline, risk level, evidence, next action. Generation is the only job left for the expensive model.
Rule: The LLM explains decisions; it does not make them.
On this page: Every flow ends with a real memo, source + latency shown.
Layer 4
The vendor-risk endpoint is a real HTTP 402 Payment Required challenge: price, network, accept header. Demo mode settles instantly; one env switch moves to Base mainnet USDC settlement.
Rule: Money moves in code, only on an approved decision.
On this page: Flow 1 receives the 402, pays, and attaches the report.
Layer 5
Everything escalated — by rules or by Jev — waits for a finance controller to release or cancel, with an optional note. A model's 98% confidence is not permission; a human signature is.
Rule: Humans own irreversible money moves.
On this page: Flows 2 and 3 land in the queue on the dashboard.
Layer 6
Transactions, audit events, agent runs, and editable policies persist. This is what makes duplicate-blocking real across sessions, what the approval queue reads, and what compliance exports as CSV/JSON.
Rule: If it isn't in the ledger, it didn't happen.
On this page: Flow 4 re-runs flow 1 — and gets blocked by memory.
The most expensive mistake in agent engineering is paying a frontier LLM to answer yes/no. Here the LLM writes, Jev decides, code executes. Each layer is independently testable and independently priced.
Jev cannot approve anything rules rejected, cannot release funds, and its escalations always route to a human — never to an auto-cancel. A wrong ML call costs a reviewer's minute; a wrong auto-approval costs real money.
Policy: free. Jev: ~$0.00002 and ~1s. Memo: seconds of cheap inference. The ledger exports with actor types, so finance can audit who decided what — model, agent, or human.
No key, timeout, or model outage returns null and the system runs rules-only. The demo never breaks mid-pitch — the AI layers are amplifiers, not dependencies.
Agents request money the way APIs do (X402). Deterministic policy filters everything code can see. Jev reads context and can only tighten the noose. The LLM writes the memo. A human is the only thing that releases funds past an escalation, and Supabase proves the whole story afterwards.