~/eugene

Product lab 01 / Financial workflows

When should AI be allowed to post an expense?

A product decision is hidden inside a simple automation promise: which suggestions can move straight into the ledger, and which need a person? Change the policy below and see the trade-off.

Working policy prototype · 18 fictional records · one recorded model run and one stress test · no live model or customer data

The product problem

Finance teams spend time checking routine records. Automating every entry reduces review work but can put the wrong account, a duplicate, or a transfer into the books. The product goal is to remove avoidable review without silently creating costly exceptions.

Why I built this

I owned Xero data-export integrations at Spenmo and later led a regulated financial product at Foundation. This is a new fictional scenario, not a feature or result from either company. It shows how I would frame the decision, define a guardrail, and evaluate a release.

Try the policy

The review gate

Higher thresholds send more records to a person. Exception rules force review of known risky cases even when the candidate output is confident.


50% · more automation99% · more review
Auto-posted—
Sent to review—
Unsafe auto-posts—
Review time avoided*—
Auto-post Human review

*Illustrative assumption: each manual review takes two minutes, catches suggestion errors, and auto-posting takes no reviewer time. The model run uses synthetic records, one prompt and self-reported confidence. This is not a measured customer saving or production benchmark.

Inspect the cases

What changes route?

The expected account is held separately from the candidate suggestion. A confident, wrong suggestion is still wrong.

Record / amountContextSuggestionExpectedConfidenceRouteReason / outcome

The recorded prompt and outputs are public so the exercise can be inspected. Expected accounts are fictional labels, not financial advice.

Product decision

I would not auto-post on this evidence.

In the recorded run, a duplicate gets a 90% confident “Needs investigation” suggestion. A threshold-only policy would still auto-post it. The failure is in the product decision, not simply the model label. I would first ship suggestions with human confirmation, log corrections, and shadow finance operators. Only then would I test automation on a narrow, low-risk segment.

01 / Discover

Observe how operators resolve unclear records, duplicates and money movement. Establish the time and error baseline.

02 / Evaluate

Label a representative holdout set. Slice errors by merchant, account, amount and exception type; inspect confidence calibration.

03 / Release

Start with suggestions. Gate any auto-posting on critical-error targets, an audit trail, reversal, and measured review-time reduction.

What this demonstrates

Product framing, AI evaluation design, data-informed trade-offs, human review, and a working prototype. The records and labels are fictional. One prediction set comes from a recorded model run; the other is an illustrative stress test. No claim here is based on real users or production performance.