EXECUTIVE BRIEFING
PSG EXECUTIVE BRIEFING · NO. 02
Where ML genuinely beats the spreadsheet — driver-based forecasts that hold up under backtesting, and learned baselines that catch the invoice, payroll, and revenue anomalies static thresholds miss. What it takes, what it costs, and how to start.
PLATINUM STRATEGY GROUP · 2026
WWW.PSG-INC.COM
Spreadsheet forecasts extrapolate: last year plus a growth assumption. ML forecasts learn the relationship between outcomes and their drivers — pipeline, bookings, headcount, seasonality, pricing actions, even weather for some businesses — and produce a range, not a single number. The practical gains: seasonality handled automatically, driver changes flow through instantly, and every forecast carries a confidence band that tells you how much to trust it. What it needs is modest: two to three years of monthly (or weekly) history, driver data in usable form, and a definition of “actual” that doesn’t move.
THE METHOD LADDER — CLIMB ONLY AS HIGH AS THE DATA SUPPORTS
Statistical baseline
Seasonal-trend models on history alone. Days to build.
The honesty benchmark — anything fancier must beat this in backtests or it isn’t earning its complexity.
Driver-based ML
Gradient-boosted models on history plus drivers. Weeks to build.
The workhorse for revenue, demand, and cash. Interpretable driver weights — the CFO can see why the number moved.
Judgment overlay
Humans adjust the model’s output — and every adjustment is logged.
Over time the log shows whose judgment adds accuracy and whose subtracts it. Uncomfortable, and worth it.
The discipline that makes it real: backtesting. Before any model touches a board pack, replay it against the past — “standing in January, what would it have said about March?” — across at least a year of holdout periods, and compare its error to the incumbent process. If it doesn’t beat the spreadsheet on the spreadsheet’s own history, it doesn’t ship.
Static rules — “flag invoices over $10,000” — catch only the failure modes someone imagined in advance, and fraud adapts to known thresholds quickly. An anomaly model instead learns what normal looks like for each vendor, employee, account, and season, and scores every new transaction against that baseline. The duplicate invoice 4% below the approval limit, the vendor whose billing cadence quietly doubled, the payroll entry that’s normal in size but abnormal in timing — these surface because they deviate from their own history, not because they crossed a line someone drew.
WHERE IT PAYS FIRST
Accounts payable
Duplicates, near-threshold amounts, new-bank-account changes, unusual vendor cadence. Typically the fastest payback — the losses are direct dollars.
Payroll & expenses
Ghost employees, off-cycle payments, expense patterns that drift from role norms. Catches at entry what audits catch a year later.
Revenue & billing
Unbilled work, pricing leakage, credit-memo abuse, usage that stopped matching the contract. Recovers revenue rather than preventing loss.
Operations
Inventory shrink patterns, quality drift, site-level cost outliers across multi-site businesses — the variances buried in monthly averages.
The design rule: tune for trust, not volume. An anomaly system lives or dies on its exception queue. Every flag goes to one queue, with the evidence attached, owned by the process owner — and the model is tuned so that at least one flag in three is worth the look. A system that cries wolf daily gets ignored by week three; better ten alerts a week that get worked than a hundred that don’t.
Pick one target and assemble the history. One forecast (revenue, demand, cash) or one anomaly domain (AP is the usual first choice). Pull 2–3 years of history plus drivers, and resolve definition conflicts now — data assembly is half the project, every time.
Build the baseline and the model. Statistical baseline first, ML challenger second. For anomalies: train the normal-behavior model on history and score the last six months — the flags it raises on known ground truth are your first quality read.
Backtest and shadow-run. The model runs alongside the incumbent process — forecast versus forecast, flags versus month-end findings — with the process owner grading output weekly. Nothing changes in production yet; credibility is being built.
Go live with ownership. The forecast feeds the planning cycle with its confidence band shown; the exception queue goes to its owner with a retraining cadence scheduled. Models are operated, not installed — drift monitoring is part of go-live.
HOW TO KNOW IT’S WORKING
Error vs. incumbent
Forecast error measured against the process it replaced, on identical periods. The only comparison that counts.
Hit rate
Share of flags worth investigating, and dollars recovered or prevented per quarter, from the exception queue’s own log.
Cycle time
Days from period close to trusted forecast. If ML didn’t shorten it, the model is decoration.
Want to know what your data could predict?
A 30-minute diagnostic consultation — your forecast process, your transaction flows, and a scoped pilot plan within 48 hours.