Observe how operators resolve unclear records, duplicates and money movement. Establish the time and error baseline.
Product lab 01 / Financial workflows
When should AI be allowed to post an expense?
A product decision is hidden inside a simple automation promise: which suggestions can move straight into the ledger, and which need a person? Change the policy below and see the trade-off.
Working policy prototype · 18 fictional records · one recorded model run and one stress test · no live model or customer data
The product problem
Finance teams spend time checking routine records. Automating every entry reduces review work but can put the wrong account, a duplicate, or a transfer into the books. The product goal is to remove avoidable review without silently creating costly exceptions.
Why I built this
I owned Xero data-export integrations at Spenmo and later led a regulated financial product at Foundation. This is a new fictional scenario, not a feature or result from either company. It shows how I would frame the decision, define a guardrail, and evaluate a release.
Try the policy
The review gate
Higher thresholds send more records to a person. Exception rules force review of known risky cases even when the candidate output is confident.
*Illustrative assumption: each manual review takes two minutes, catches suggestion errors, and auto-posting takes no reviewer time. The model run uses synthetic records, one prompt and self-reported confidence. This is not a measured customer saving or production benchmark.
Inspect the cases
What changes route?
The expected account is held separately from the candidate suggestion. A confident, wrong suggestion is still wrong.
| Record / amount | Context | Suggestion | Expected | Confidence | Route | Reason / outcome |
|---|
The recorded prompt and outputs are public so the exercise can be inspected. Expected accounts are fictional labels, not financial advice.
Product decision
I would not auto-post on this evidence.
In the recorded run, a duplicate gets a 90% confident “Needs investigation” suggestion. A threshold-only policy would still auto-post it. The failure is in the product decision, not simply the model label. I would first ship suggestions with human confirmation, log corrections, and shadow finance operators. Only then would I test automation on a narrow, low-risk segment.
Label a representative holdout set. Slice errors by merchant, account, amount and exception type; inspect confidence calibration.
Start with suggestions. Gate any auto-posting on critical-error targets, an audit trail, reversal, and measured review-time reduction.
What this demonstrates
Product framing, AI evaluation design, data-informed trade-offs, human review, and a working prototype. The records and labels are fictional. One prediction set comes from a recorded model run; the other is an illustrative stress test. No claim here is based on real users or production performance.