Disclaimer: Independent product concept by Kaushal Khodifad. Not a live commercial product.return to portfolio
Disclaimer: Independent product concept by Kaushal Khodifad.
Product · the build

Arbiter - adjudication infrastructure, the Decision OS for regulated work.

Arbiter is a platform, not a point tool: one governed core that runs a fleet of AI digital-worker agents, each rendering regulated decisions that are cited and warranted.

Platform · a fleet of agentsLive demo ↓Overlay, not rip-and-replace
Forward-deployed services
Arbiter · warranted agents
Months of forward-deployed services
Weeks, as a productized overlay
You own the liability
Insurer-backed decision warranty
Rip-and-replace the system
Overlay - writes back to your tools
An illustrative category contrast - services-led delivery versus a productized overlay. Not a benchmark against any named vendor.
The product is the platform

One platform, a fleet of governed agents

Many governed agents on one shared spine. A new vertical is a new agent, not a new product.

Arbiter VerifierCore
Shared · all verticals

Ingests documents, extracts every field with provenance, and runs fraud / tamper checks - the evidence layer every other agent reads from.

reads: any document, image, feedL2
Arbiter UnderwriterBeachhead
Residential mortgage

Calculates and verifies qualifying income, then clears or holds income, asset and credit conditions.

reads: URLA, W-2, paystubs, tax returns, AUSL2
Arbiter AnalystExpansion
Commercial lending

Spreads financials, computes DSCR / LTV / debt-yield, extracts and tests covenants, drafts the credit memo.

reads: tax returns, financial statements, rent rolls, leasesL1
Arbiter RiskExpansion
Insurance underwriting

Triages submissions, checks appetite, and scores risk 0-5 with a cited rationale.

reads: ACORD, loss runs, SOV, MVRL1
Arbiter AdjusterExpansion
Insurance claims

Handles FNOL intake, verifies coverage, recommends reserves, and flags leakage / fraud.

reads: FNOL, declarations, photos, estimatesL1
Arbiter ReviewerCore
Shared · all verticals

Samples auto-rendered decisions for QC, watches for drift and fairness, and feeds the evaluation harness.

reads: decisions + outcomesL3
SPINE

Every agent inherits the same governed core - PolicyForge, Checkpoint, Ledger, Proofset and the Institutionalizer learning loop. Adding a vertical means adding an agent, not rebuilding the platform.

See it work · live console

Every agent, fully instrumented

Pick an agent and watch it run - full observability, audit, explainability and integration.

Arbiter UnderwriterL2Residential mortgageSynthetic
Illustrative · synthetic - not measured results
58%
Auto-cleared
of clean conforming files
1.8s
Draft latency
p50 / 6.2s p95
100%
Grounded
cite-or-abstain
0.18%
Income defect
vs ~1.1% baseline
Conforming purchase · self-employed borrower · A. Rivera - sole proprietor (Schedule C)
  1. 1INTAKE
  2. 2EXTRACT
  3. 3DERIVE
  4. 4TEST
  5. 5CHECKPOINT
  6. 6LEDGER

Classify the file & documents

Auto

Identified a conforming purchase with self-employed income; loaded 18 documents.

tool callclassify_file()
in: 18 docs (URLA, 1040s, P&L, AUS)
out: {loanType: conforming-purchase, employmentType: self-employed}
PII redaction - SSN + account numbers masked
Confidence signal99%
INTAKE-OK
Real vs simulated in this demo
Real Synthetic
REALDecision math & routing - deterministic TypeScript, server-side each run (engine.ts)
REALModel calls + run ledger - OpenRouter when keyed; SHA-256 hash-chain you can re-verify
SYNCase files & metrics - synthetic files; headline metrics labeled “Illustrative”
SYNExternal systems - AUS / LOS / bureau calls are scripted, never live
Under the hood

The shared core every agent runs on

The same ten components per agent. Tap any for its job, parameters and an example.

Ingest- Turn messy paperwork into clean, traceable facts.
Reason- Decide the case the way an expert would.
Govern- Keep every action inside the rules, with sign-off.
Learn- Get better from every correction and outcome.
Integrate- Sit on top of the systems they already use.
IngestKitIngest
Ingestion & provenance

Turn any inbound artifact (PDF, image, fax, feed, email) into structured, provenance-tagged evidence.

Key parameters
  • extractor_routeOCR vs layout-LLM vs structured-parser, chosen per document class
  • confidence_floorper-field extraction confidence required to auto-accept (e.g. 0.94)
  • provenance_requiredevery field carries {source_doc, page, bbox, extractor, timestamp}
  • freshness_ttlmax age a document is valid for (e.g. paystub ≤30 days)
Example
A W-2 image → IncomeDocument{employer, wages, ytd, page:2, bbox, confidence:0.94}.
JD →Integration patterns · golden datasets · starter kits
Trust dial

Autonomy is earned, one rung at a time

Nothing auto-renders until it's proven. Confidence and risk set the tier.

L1 · ApproveHuman approves each decision.

Worker decides; a human approves every case.

Example: Worker clears the income condition; underwriter clicks approve.

A worker earns its way up the ladder - it only auto-renders once the evaluation harness proves it on a golden set. The same file can sit at different tiers for different decisions.

Why a warranty wins

Standing behind the decision is the business model

We sell de-risked decisions, not automation. Lower defects compound into margin - drag the sliders.

Assumptions
120,000
55%
$18
0.4%
$9.0K
Per-year economics
Illustrative
Auto-rendered decisions
66,000
Revenue
$1.2M
Warranty payouts
$2.4M
Defects covered
264
Gross profit after warranty
$-1.2M
Margin
-100%

Drag defect rate down: the warranty gets cheaper and margin climbs. That is the flywheel - fewer defects fund a cheaper warranty, which wins more volume, which yields more outcome data, which lowers defects.

Honest caveat: the defect rate here is an unmeasured assumption, not a track record. In a real deployment it would be measured in shadow mode against senior-underwriter ground truth - and no warranty would be priced or written before that baseline exists.

How it scales

A solution factory, not a services shop

Three layers, so each new customer and vertical ships faster than the last.

1Core platform

Built once, shared by everyone.

  • WorkerRuntime
  • PolicyForge
  • Checkpoint
  • Ledger
  • Proofset
  • Institutionalizer
  • connector framework
2Domain package

The vertical IP - reused per industry.

  • CanonGraph ontology
  • policy packs
  • golden sets
  • worker state machines
  • reason-code maps
  • hero metrics
3Customer config

Tenant-specific overlay - set per lender.

  • lender overlays
  • RBAC + approval matrix
  • connector credentials
  • autonomy tiers
  • branding
Defensibility

The moat stack

Ranked by depth. The deep ones compound per decision.

1

Judgment network effect

Every override and outcome flows into versioned policy + memory + recalibrated thresholds. Incumbents capture data - not structured judgment lineage with a closed loop to convert it.

Compounds: More decisions → better golden sets → higher safe-autonomy → more decisions auto-rendered → more outcome labels.

2

Governance as switching cost

Their guidelines, overlays, RBAC, approval matrices and audit history live in our ledger. Leaving means rebuilding their control framework and losing the immutable trail exams depend on.

Compounds: Every passed audit and accepted policy version deepens the dependency.

3

Regulatory trust / liability transfer

We stand behind decisions with an insurer-backed warranty - selling de-risked outcomes, not automation. That requires a clean eval + outcome history rivals don't have.

Compounds: Lower realized defect rate → cheaper warranty → more volume → more outcome data → lower defect rate.

4

Accelerator-factory speed

Productized blueprints + a connector framework collapse time-to-value to weeks. Services-led rivals' delivery cost stays linear; ours falls with every customer.

Compounds: Each deployment hardens the domain packages used by the next.

5

Distribution via neutrality

The only serious player with no system-of-record and no balance sheet to protect - so we run on Encompass and its rivals, and automate past the ceiling incumbents won't cross.

Compounds: More connectors → more deployable accounts → more decisions → feeds moats 1-3.

Out-flanking ICE Mortgage Technology
Their moat: Owns the system of record (Encompass) + the data rails; ~40%+ of US loan volume.
Our flank: Don't fight for the system of record - make it irrelevant to where the decision is rendered. ICE monetizes the seat and under-automates to keep humans in Encompass; we monetize the decision and automate past that ceiling, running on Encompass and its rivals. Lenders multi-vendoring against ICE lock-in actively want us.
Out-flanking Rocket
Their moat: A captive, decades-deep proprietary loan-outcome dataset at scale.
Our flank: Rocket's flywheel can't be sold to competitors without arming them. We federate anonymized cross-lender judgment into shared market intelligence broader than any single captive dataset - while each lender keeps its private edge. Many lenders' collective judgment beats one giant's private one.
Build-ready

The PRD contract

Fifteen sections - minimalist, but enough to build the worker from.

01Problem & wedge

The thin-slice workflow, why it's the highest-leverage entry, what breaks today.

02Users & personas

Who operates / approves / owns, and their jobs.

03Digital-worker job spec

Inputs, the decisions it renders, the decisions it must NOT render, its autonomy tier.

04Workflow / state machine

States, transitions, guards, terminal decisions, recovery points.

05Data & document model

Ontology entities, document classes, required fields, provenance + freshness.

06Agent & tool design

Typed tool registry, model routing, step/token budgets, determinism-vs-LLM split.

07Policy & guardrails

Policy pack (rules + citations), RBAC, approval matrix, control packs.

08HITL & escalation

Autonomy tiers, confidence thresholds, risk gates, escalation paths + SLAs.

09Explainability & audit

Reason-code set (mapped to regulation), citation format, immutable ledger schema.

10Eval harness & golden set

Golden dataset, regression gates, fairness tests, drift monitors, shadow-mode plan.

11Integrations

Target systems, connector, sync mode, field mappings, write-back permissions.

12Metrics & SLOs

The 3-4 hero outcome metrics + operational SLOs (latency, auto-rate, QC rate).

13Learning loop

Override-to-policy threshold, memory config, bandit reward, outcome-label lag.

14Rollout / thin-slice plan

Shadow → L1 → L2 → L3 gates, success criteria per phase, next slice.

15Non-goals

Explicitly out-of-scope decisions, systems and edge cases (caps scope + liability).

Concept · Illustrative · Independent studyArbiter is an independent product concept by Kaushal Khodifad - a coined demo brand, not a real company. The math, ledger and model calls in the live console are real; case files and metrics are synthetic and labeled.

↑ Independent product concept by Kaushal Khodifad. Arbiter is an independent product concept by Kaushal Khodifad; it is not a real company or a commercial product. It explores the adjudication infrastructure for regulated lending & insurance space. Not a live commercial product. Data is illustrative.

Return to portfolioOpen the live demo