By R Douglas Orsagh August 6, 2026
What Your Data Needs to Be Ready for Agentic Finance (or How the Misapplied Laws of Thermodynamics Can Explain Agentic AI Failure)

Executive summary
For chief financial officers (CFOs), the readiness question is not whether the AI works. The question is instead whether Finance can trace every number, apply consistent business logic across systems, and defend decision thresholds under audit. Finance must validate lineage, hierarchy integrity, and governance on one process first. Otherwise, automation may accelerate existing data issues rather than improve efficiency and control.
When adopting artificial intelligence (AI) in accounting processes, organizations must overcome various pitfalls when working with agentic Finance. Here, the laws of thermodynamics offer some insight. The laws apply when experiments occur within a closed system. In expanded or more open systems (i.e., the entire galaxy), some observations aren’t consistent with what happens in a closed system. Why? Because different, unimagined variables emerge. Entropy thus rears its inevitable head as the system expands and results move toward an inevitable disorder.
So what do the laws of thermodynamics have to do with accounting? Many AI projects are done on a small piece of a process. After a positive result, the small (piece of a) process is rolled out to the entire org. Often, disorder follows. The underlying issue in such projects is that AI was imagined to fix just one process. Instead, successful enterprise-wide agentic AI adoption comes from imagining and examining AI usage at the policy level.
Here’s a hypothetical to show how the accounting entropy reveals itself.
It’s Day 3 of a typical 10-day close. The accounting team leaders are on a Slack call discussing the day’s progress (or, as it usually seems at this point, the lack thereof). On the call, the EMEA team is raising a concern: “There’s a 12M euro variance that seems really ‘slippery.’”
While not a GAAP reference, “slippery” is a term accounting teams know all too well. Slight panic also usually ensues. Why? Because time and energy are spent tracking down this variance. And it’s time the team doesn’t have. In a pre-post-mortem, the controller asks, “How could this have been prevented?”
Why should the team worry about the variance? After all, variances happen every close. They’re almost always explainable and fixed. Here, the actual problem is that the correct policy wasn’t in place. In the subsequent moving of data from multiple systems, no one caught the manual journal entry from an analyst who left the company several months ago.
That's not a model problem. No algorithm fixes a broken trail. Instead, the problem is a data-authority one. Why? Because policy wasn’t designed to properly capture who owns that number at each step. In this case, the variance was such an esoteric issue that policy didn’t dictate that a human being should track the variance. Nor did current accounting systems know to flag it.
Advisory AI vs. agentic AI, without the handwaving
AI has rapidly evolved into an invaluable tool to help accounting and Finance teams. However, not all flavors of AI are the same. Different types of AI tools deliver different results. To illustrate, let’s return to the hypothetical introduced above.
The advisory version: A model flags the 12M euro swing the morning it appears. In response, AI enables a (more efficient) game of “telephone.” A financial planning and analysis (FP&A) analyst then takes the following steps:
- Pulls the variance report
- Calls the regional controller
- Spends 4 hours tracing the variance to an invoice booked to the wrong entity from a stale customer mapping
The variance gets journaled, the close proceeds, and the mapping fix lands in the backlog — real but never quite urgent enough to fix.
The agentic version: The system traces the swing to the invoice, checks the customer mapping against the entity hierarchy, and finds the mismatch. Then, the system drafts the correcting entry with the evidence attached. Under a pre-defined threshold, that draft is routed for same-day human posting. If above the threshold, the variance goes to escalation instead. The 4 hours spent tracing the variance with advisory AI tools becomes a quick review against the evidence with agentic AI. Then it’s just a click to post.
The agentic version is clearly preferable. However, that version only works if three things are true before the variance appears:
- Invoice traces through every system touched, not just the one the invoice landed in
- Customer and entity hierarchies agree on one answer for “correct entity,” not two
- Someone decided in advance what's safe for same-day posting vs. what escalates
Yet at the process level, most Finance organizations have none of these things defined — an issue that could have been addressed by policy. The model finds the variance just fine. However, the model then stalls exactly where a human would have picked up the phone (next to spreadsheets, the most popular process band-aid). The problem? No one was told to expect the call.
Why agentic AI can break in practice
1. The number that can't be traced during close
This breaking point is the EMEA scenario and the most common failure pattern in practice. Here, the reconciliation does its job by flagging the variance in the first place. What breaks is what comes next. While the system finds the same variance fast, the trace then stops at a hop nobody instrumented. That hop wasn’t explicitly laid out in the processes. Nor was it based on policy.
What does that look like? It’s usually a manual adjustment or a legacy interface that never got wired in because no one expected a machine would need to follow the policy.
While the diagnosis is correct, the evidence chain needed for a controller to sign off is missing. Policy ultimately never imagined that was necessary. As a result, the fix happens manually, and the pilot gets written up as “not ready yet.”
2. Hierarchies that conflict across business units
A manufacturing client, disguised here but real in substance, had two distinct enterprise resource planning (ERP) systems after an acquisition. General ledger (GL) code 4010 meant product revenue in one system and services revenue in the other. When drafting entries for a controller's daily batch sign-off, a reclassification agent moved $340K of services revenue into a product line. Both systems read that code the same way. The controller signed off on the batch.
The problem? The review checked dollar thresholds, not hierarchy logic. As a result, the misclassification went straight through. The chart of accounts was never really one chart, just two that happened to share numbering. During post-merger integration, this issue shows up often when “one version of the truth” gets asserted in a deck before the data supports that truth.
3. The decision nobody can explain at audit
An automated write-off engine drafted a $48K bad-debt write-off under a threshold set at $50K. Since the write-off fell within policy, the controller signed off without a second look. Six months pass, and then an auditor asks why $50K was the number and whether it still holds. The write-off was real policy, approved back when the average write-off was a fraction of today's size, and nobody had revisited it since.
In other words, the write-off itself was correct, but the underlying threshold was stale. That's what surfaces in audit, and what turns an internal review into an external finding. While that small of a write-off may not matter as an isolated event, the question to ask is how many more are out there.
The half-adoption trap
Currently, adopting agentic AI cleanly is difficult. It gets piloted in a controlled, closed system. That typically happens with one process and on a subset of something like intercompany reconciliation or variance investigations while everything else stays manual. The expectation? That the entropy experienced in this closed system pilot will transfer directly to everything on the balance sheet.
While the process is addressed, the overriding policy issues are not. A one-off experiment also only proves that the one-off use case has succeeded. Instead, it should be transferred into an overriding policy with agentic AI architected to support POLICY prior to process.
Now, let’s examine the pattern and its pros and cons. The pilot runs clean for 6 – 8 weeks. Then it makes one visible mistake, something like a correcting entry posted to the wrong period because a fiscal calendar override wasn't synced. The controller adds a manual review step back in, “just until we're sure.” The controller has created a de facto policy that won’t be undone in the pilot and now migrates into overall policy outside of the closed system once the pilot is done. That step never comes out.
Six months later, the system is technically “live,” but a human checks every action before it posts. The system is now back to running like an advisory tool with an agentic label. Here, chief financial officers (CFOs) lose confidence in the number. The model didn't fail. Nobody scoped the calendar override into governance on day one.
Here’s the other half-adoption failure: running two systems of record through one agentic layer without reconciling them first — the two-ERP problem above at scale. If there wasn’t a clear process in place to reconcile this prior to running the agentic layer, then there is no way it will work.
A broader design policy to avoid the issue would dictate having a process in place to ensure the problem didn’t occur. Instead, automation just runs the inconsistency at machine speed instead of catching the issue. Policy thus allows for a proactive process that resolves the issue or fences it off before an agent touches it. Ultimately, the math doesn't care which system posted first.
Policy, then process, then people
Most of the failures above trace back to one thing: the order in which decisions get made. There's a real hierarchy here. First, a CFO and their team set policy. Then a controller or chief accounting officer (CAO) designs the process that carries out that policy. For the day-to-day, people get hired to run the process. Most agentic AI projects start in the middle, instead of at the top, by inserting an agent into a process that already exists. In practice, that means deciding as policy what role AI plays before the process gets built around it.
Account reconciliations exhibit this point well. In a typical month, something like half of all recs net to a zero balance with no real movement. Policy can safely auto-clear those without a human or an agent looking twice. Another large group involves real change that needs a quick check, an invoice or two to verify, a flux to explain. And that's exactly where an agent earns its keep, cutting a 10-minute review down to a minute or two.
Meanwhile, the remainder, usually intercompany and cross-entity cases, requires human decision-making no matter how good the tooling gets. Which tier gets which treatment is a policy decision, made before anyone builds the process around it, not a technology choice.
If the order is backward, agentic AI creates more work, not less. It’s like reverse entropy. That’s great for the time travel that might be required if the agentic AI implementation is misapplied — a so-called cosmic mulligan, which does not exist outside Marvel movies. Teams that bolt an agent onto an existing process end up reviewing everything the agent flags, on top of everything the team was already doing. Why? Because nobody decided in the policy step who owns checking the agent's output and when.
The fix is deciding, as policy, what the agent is trusted to do before it's turned loose on the process. That does not involve assembling a bigger team to keep up with what the agent flags.
This situation is also why AI policy can't sit in IT alone. Instead, policy must be a joint call between the accounting office and IT. Why? Because the accounting office is who must defend the number afterward. A policy IT writes without accounting at the table tends to optimize for system access and security, which misses the parts that actually get tested in an audit.
None of this works with a general assistant bolted onto the process, either. The tools built into this policy need real financial context, not just language fluency. Thus, the underlying data model matters just as much as the AI model.
The three requirements, made concrete
1. Lineage
“Document lineage” only means something once it's this specific. Here's an actual lineage map for one metric (EMEA product revenue) before any agent acts on it:
| Hop | System | What happens here | Who can explain it |
| 1 | Salesforce (EMEA) | Opportunity closes, booking created | EMEA sales operations manager |
| 2 | Billing system | Booking converts to invoice, currency locked | Billing manager |
| 3 | ERP GL | Invoice posts to revenue account, entity-mapped | Regional controller |
| 4 | Consolidation layer | Entity rolls up, FX translation applied | Consolidation manager |
If you can't produce a table like this within a day (not perfectly, just directionally), the metric isn't ready for agentic execution.
2. Hierarchy integrity
The fastest way to find a hierarchy conflict isn't a data quality audit. Instead, five accounts that exist in more than one system should be selected. Then one question should be asked: Does this code mean the same thing, at the same granularity, in both places? That question would have surfaced the GL 4010 conflict mentioned above in an afternoon, before automation touched the conflict.
If the team gets it wrong, the same transaction resolves two different ways depending on which entity touches it first. Most teams skip the question because it feels too basic to be the blocker, even though it usually is.
3. Governance boundaries
A well-run Finance org already has this policy, and it’s written for a human to execute. Agentic AI simply operates inside the policy, specific enough that anyone unfamiliar with the process, human or machine, can apply the policy correctly. Here’s a starting version for three common Finance decisions:
| Decision | Same-day posting if… | Escalate to controller if… | Owner of policy |
| Intercompany reconciliation write-off | Under $25K, matched to a known intercompany mapping | Over $25K or no mapping found within three attempts | Assistant controller |
| Revenue variance investigation | Under $50K, traceable to a single source document | Over $50K or traces to more than one source system | FP&A director |
| Accrual reversal | Exact match to the original accrual schedule | Any deviation from the original schedule | Controller |
Notice what's missing: Any mention of AI or automation platforms. This policy is the same one that already governed a human doing the work. Skip the policy, and a CFO can't defend the number under scrutiny. The model isn't the exposure here. Instead, the missing threshold brings exposure.
A 90-day playbook, with owners and artifacts
Weeks 1 – 2: Lineage for one metric
- Owner: FP&A manager for the metric, plus the systems owner at each hop.
- Artifact: A lineage table like the one above. Trace to source in under 5 minutes, unassisted.
Weeks 3 – 4: Resolve one hierarchy conflict
- Owner: Financial systems manager, with sign-off from both controllers if the conflict spans business units.
- Artifact: A mapping table resolving where the same code means different things. One answer, regardless of which system the answer is queried from.
Weeks 5 – 6: Threshold policy for one recurring decision
- Owner: Whoever currently approves this decision, typically a controller or assistant Controller.
- Artifact: A threshold table like the one above, usable by someone outside the process without help.
Weeks 7 – 12: Shadow mode, then live
- Owner: Whoever owns the process end to end, usually the controller but sometimes FP&A, depending on the decision.
- Artifact: A decision log comparing system output to actual outcome. A 95% match before the draft moves from full review to standard sign-off, with a reason logged for every mismatch.
No platform decision, no steering committee. Just a controller, an FP&A manager, and a financial systems manager, each focused on one metric, one hierarchy, and one decision. Can't name those three owners? That's the blocker.
The objections, and the logic behind the truth
“This adds risk.” Only if governance arrives after automation. Sequenced first, a written policy holds up under audit in a way personal judgment doesn't, once the person who made the original call has moved on.
“Our data isn't ready.” This objection may be correct, but that's a finding rather than an excuse to wait. By design, Finance data is transactional. Significant transformation is needed for this level of real-time decisioning to work enterprise-wide. Instead of trying to fix everything at once, the place to start is with the playbook above on one metric.
“We tried this, and it didn't work.” This objection is the closed system example we cited earlier with entropy escaping once rolled out writ large. One visible mistake? A review step bolted back on, never removed. Here, the lesson isn’t that the technology wasn’t ready. The lesson is instead that the threshold that would have made the mistake impossible was never defined.
The diagnostic
Pick one decision your team makes weekly. A reconciliation write-off. A variance investigation. An accrual reversal. Then answer three questions, in writing:
- Can you trace the number behind this decision to its source, through every system it touches, in under 5 minutes, with a named owner at each hop?
- If this decision were made in two different business units, would the result be the same, using the same hierarchy logic, both times?
- Is there a written threshold that says when this decision is safe to fast-track and when it escalates, with a named owner for that policy?
Did you answer no to any of those questions? Then don't let a system post that decision without someone in the loop, no matter how good the model or how the vendor's demo looked.
Want to learn more about the future of Finance? Read the Forward Finance Playbook.
R Douglas Orsagh, CFA, CPA, is the Vice President of Business Transformation and Growth at OneStream. Doug was brought into OneStream to create and lead a team of former CFOs and industry experts to help enterprises understand how OneStream’s unique capabilities can improve business outcomes. He works directly with the C-suite at OneStream customers, and he and his team have helped more than 500 prospects decide to become OneStream customers.
Prior to OneStream, Doug created and led a similar team at Oracle. He is a recovering CFO (twice), with tenures at both a global telecommunications company and a real estate company. Doug is a PwC alum and a former investment banker. He holds an MBA in Finance from Georgia State University, as well as dual degrees in Economics and Communications from the University of Notre Dame.




