Arbiter - adjudication infrastructure, the Decision OS for regulated work.
Arbiter is a platform, not a point tool: one governed core that runs a fleet of AI digital-worker agents, each rendering regulated decisions that are cited and warranted.
One platform, a fleet of governed agents
Many governed agents on one shared spine. A new vertical is a new agent, not a new product.
Ingests documents, extracts every field with provenance, and runs fraud / tamper checks - the evidence layer every other agent reads from.
Calculates and verifies qualifying income, then clears or holds income, asset and credit conditions.
Spreads financials, computes DSCR / LTV / debt-yield, extracts and tests covenants, drafts the credit memo.
Triages submissions, checks appetite, and scores risk 0-5 with a cited rationale.
Handles FNOL intake, verifies coverage, recommends reserves, and flags leakage / fraud.
Samples auto-rendered decisions for QC, watches for drift and fairness, and feeds the evaluation harness.
Every agent inherits the same governed core - PolicyForge, Checkpoint, Ledger, Proofset and the Institutionalizer learning loop. Adding a vertical means adding an agent, not rebuilding the platform.
Every agent, fully instrumented
Pick an agent and watch it run - full observability, audit, explainability and integration.
- 1INTAKE
- 2EXTRACT
- 3DERIVE
- 4TEST
- 5CHECKPOINT
- 6LEDGER
Classify the file & documents
AutoIdentified a conforming purchase with self-employed income; loaded 18 documents.
The shared core every agent runs on
The same ten components per agent. Tap any for its job, parameters and an example.
Turn any inbound artifact (PDF, image, fax, feed, email) into structured, provenance-tagged evidence.
- extractor_routeOCR vs layout-LLM vs structured-parser, chosen per document class
- confidence_floorper-field extraction confidence required to auto-accept (e.g. 0.94)
- provenance_requiredevery field carries {source_doc, page, bbox, extractor, timestamp}
- freshness_ttlmax age a document is valid for (e.g. paystub ≤30 days)
Autonomy is earned, one rung at a time
Nothing auto-renders until it's proven. Confidence and risk set the tier.
Worker decides; a human approves every case.
Example: Worker clears the income condition; underwriter clicks approve.
A worker earns its way up the ladder - it only auto-renders once the evaluation harness proves it on a golden set. The same file can sit at different tiers for different decisions.
Standing behind the decision is the business model
We sell de-risked decisions, not automation. Lower defects compound into margin - drag the sliders.
Drag defect rate down: the warranty gets cheaper and margin climbs. That is the flywheel - fewer defects fund a cheaper warranty, which wins more volume, which yields more outcome data, which lowers defects.
Honest caveat: the defect rate here is an unmeasured assumption, not a track record. In a real deployment it would be measured in shadow mode against senior-underwriter ground truth - and no warranty would be priced or written before that baseline exists.
A solution factory, not a services shop
Three layers, so each new customer and vertical ships faster than the last.
Built once, shared by everyone.
- WorkerRuntime
- PolicyForge
- Checkpoint
- Ledger
- Proofset
- Institutionalizer
- connector framework
The vertical IP - reused per industry.
- CanonGraph ontology
- policy packs
- golden sets
- worker state machines
- reason-code maps
- hero metrics
Tenant-specific overlay - set per lender.
- lender overlays
- RBAC + approval matrix
- connector credentials
- autonomy tiers
- branding
The moat stack
Ranked by depth. The deep ones compound per decision.
Judgment network effect
Every override and outcome flows into versioned policy + memory + recalibrated thresholds. Incumbents capture data - not structured judgment lineage with a closed loop to convert it.
Compounds: More decisions → better golden sets → higher safe-autonomy → more decisions auto-rendered → more outcome labels.
Governance as switching cost
Their guidelines, overlays, RBAC, approval matrices and audit history live in our ledger. Leaving means rebuilding their control framework and losing the immutable trail exams depend on.
Compounds: Every passed audit and accepted policy version deepens the dependency.
Regulatory trust / liability transfer
We stand behind decisions with an insurer-backed warranty - selling de-risked outcomes, not automation. That requires a clean eval + outcome history rivals don't have.
Compounds: Lower realized defect rate → cheaper warranty → more volume → more outcome data → lower defect rate.
Accelerator-factory speed
Productized blueprints + a connector framework collapse time-to-value to weeks. Services-led rivals' delivery cost stays linear; ours falls with every customer.
Compounds: Each deployment hardens the domain packages used by the next.
Distribution via neutrality
The only serious player with no system-of-record and no balance sheet to protect - so we run on Encompass and its rivals, and automate past the ceiling incumbents won't cross.
Compounds: More connectors → more deployable accounts → more decisions → feeds moats 1-3.
The PRD contract
Fifteen sections - minimalist, but enough to build the worker from.
The thin-slice workflow, why it's the highest-leverage entry, what breaks today.
Who operates / approves / owns, and their jobs.
Inputs, the decisions it renders, the decisions it must NOT render, its autonomy tier.
States, transitions, guards, terminal decisions, recovery points.
Ontology entities, document classes, required fields, provenance + freshness.
Typed tool registry, model routing, step/token budgets, determinism-vs-LLM split.
Policy pack (rules + citations), RBAC, approval matrix, control packs.
Autonomy tiers, confidence thresholds, risk gates, escalation paths + SLAs.
Reason-code set (mapped to regulation), citation format, immutable ledger schema.
Golden dataset, regression gates, fairness tests, drift monitors, shadow-mode plan.
Target systems, connector, sync mode, field mappings, write-back permissions.
The 3-4 hero outcome metrics + operational SLOs (latency, auto-rate, QC rate).
Override-to-policy threshold, memory config, bandit reward, outcome-label lag.
Shadow → L1 → L2 → L3 gates, success criteria per phase, next slice.
Explicitly out-of-scope decisions, systems and edge cases (caps scope + liability).