Sentinel hands you a black-box safety score. Spotter hands you the decision - and the receipts.
Spotter AI’s Sentinel pulls MVR, PSP, FMCSA and CDLIS on a CDL upload in ~30 seconds and returns an AI safety score. This is a teardown of that pattern: Paste a synthetic CDL and watch a real, auditable engine assemble a driver risk file - every point of risk traced to a specific FMCSA violation and the CFR behind it, with an auto-approve / manual-review / deny verdict you could defend to an auditor.
Underwriting trust, not data access, is the moat.
Glass-box FMCSA-grounded scoring - replicate the real SMS construction so every point is explainable line-by-line, not a black box.
Decision, not just data - an explicit auto-approve / manual-review / deny verdict with a reason ledger tied to specific violations and the governing CFR.
Counterfactual transparency - toggle any signal and the file is re-underwritten from scratch; every delta is verified against an independent spec recompute, and a built-in fuzz harness shows the checks failing against deliberately broken engine variants.
Grounded incident triage - an LLM that retrieves the driver record + carrier policy and produces a cited RCA that abstains on thin evidence.
A black-box score vs. a reconstructable reason ledger
SambaSafety and Spotter AI’s Sentinel hand a carrier a number; the same synthetic driver, run through this engine, returns the number and the line-by-line derivation - each point traced to a violation and its governing CFR.
Illustrative comparison on synthetic driver Driver B-2207. SambaSafety and Spotter AI are real companies; the score, line items, and verdict here are produced by this prototype’s engine on synthetic data, not by those vendors’ systems.
Four signal pulls (CDLIS · MVR · PSP · Clearinghouse) resolve into one risk file.
Driver Intake →An auditable engine reproduces the real FMCSA SMS - every point a line item.
Workbench →Line items roll up into six FMCSA categories - cap-30 worst-first, time-weight decay.
BASIC Breakdown →A printable record: verdict, index, reason ledger, CFRs, DQF checklist, FCRA note.
Decision Record →A grounded LLM retrieves record + policy, returns a cited RCA, abstains on thin evidence.
Incident Triage →Who pays, and what a defensible verdict is worth
A worked example, not a business case: every row is badged for what it is, and every derived number is plain arithmetic over the badged inputs. Nothing below is a measured outcome of this prototype.
Pricing hypothesis (directional): per driver-file underwritten, anchored to the screening spend Sentinel claims to cut ~75% (vendor claim) and to the ~$202/file headroom one avoided crash creates. Not a quote; no unit economics are measured by this prototype. Sources and audit status for the verified figures are on the Methodology page.
- Deterministic engine + ledger arithmetic
- Falsifiable check suite + fuzz harness
- OpenRouter LLM triage - labeled live vs. fallback per run
- SMS mechanics: time weights, +2 OOS, cap-30
- Dec-2025 severity overhaul
- 49 CFR 382/391 hard gates
- 11 driver personas + violation histories
- Severity table values (representative)
- Exposure divisor + 0-100 index map + peer bands