Ask the office of the CFO anything.
Grounded answers over Northwind Cloud (fictional Series-C SaaS · ~600 employees) - every number computed by a deterministic engine, every claim cited to a source row, and an honest abstain when the data can't support it.
Same question, verifier OFF vs ON
We take a raw model answer that includes a fabricated figure not in the grounded facts, then run the production numeric-grounding gate over it. The right column is what the model wanted to say; the left is what actually reaches you. The diff below is the gate's own output - not a script.
Pick a question to run the gate live. (Uses the deterministic engine - no API key needed.)
Try to sneak a number past the gate
Each row injects a fabricated sentence into the production verify() gate - the same module the API route imports - against live engine facts, client-side at render. Rows marked PASSED v1 are the exact scaled-form and free-pass holes the previous gate accepted; gate v2 is unit-aware and rounding-exact, so they must all show green. If any row is red, the gate is broken and this panel will say so.
| Attack class | Injected sentence | Expected | Gate v2 |
|---|---|---|---|
Pure fabrication | A further $874,221,000 was lost to fraud this week. | must strip | stripped |
×100 scale collision old gate admitted this | Total cash is $2,240,000,000. | must strip | stripped |
÷1e3 rescale (K for M) old gate admitted this | We hold $22.4K in cash. | must strip | stripped |
Small-integer smuggling old gate admitted this | The board approved 9 new credit lines. | must strip | stripped |
Year smuggling old gate admitted this | A covenant breach is projected for 2043. | must strip | stripped |
±1% tolerance nudge old gate admitted this | Runway is 9.9 months. | must strip | stripped |
Unit swap ($ ↔ months) | Runway is $9.8M. | must strip | stripped |
Invented inline citation | Fraud was confirmed in txn_99999_zz per the ledger. | must strip | stripped |
CONTROL · real narration | Runway is 9.8 months - $22.40M of cash across Mercury and Brex ÷ $2.29M/mo trailing net … | must pass | untouched |
CONTROL · verbal rounding | We have about $22.4M of cash and 9.8 months of runway. | must pass | untouched |
CONTROL · counts + windows | The rule engine flagged 6 anomalous outflows over the trailing 60 days - including a dup… | must pass | untouched |
Write a sentence with any figure and run it against the runway facts ($22.40M cash · $2.29M/mo burn · 9.8 months). Only correctly-rounded, correctly-united grounded figures survive.
Honest limit: this gate proves every figure is grounded - it cannot judge a numberless causal claim. Sentences with no figures and no invented row ids pass; the methodology page says so explicitly.
What I deliberately can't answer
A finance buyer trusts a tool more when it refuses cleanly than when it guesses. These are real CFO-grade questions the copilot abstains on by design - each reason is computed by the live router, from the absence of the required row type. Click one to fire it through the copilot and watch the computed abstain.
Refusing well is a feature. The reason text is generated by the production router, so it can never drift from what the copilot actually does.