Blog · AI and self-service analytics
Which parts of financial analysis AI can be trusted with, which it cannot, and how to tell the difference. A budget variance question is worked through on six ledger accounts, with the checks that make the answer trustworthy, where AI helps most in FP&A, and the data and governance questions to settle first.
AI for financial analysis can be trusted to read a question, choose the right analysis and draft the commentary. It should not be trusted to produce the figures itself. An answer is worth using when every number came from a tool that computed it from your ledger, the rows behind it are cited, and the totals tie to the trial balance.
Financial analysis has two parts that look like one: deciding what to calculate, and calculating it. A language model is good at the first and unreliable at the second, because it predicts text rather than adding columns. The working rule, explained in full in why the AI must never do the arithmetic:
The external AI model explains. Tools compute.
The external AI model reads the question, picks the calculation and writes the commentary. A deterministic tool computes every figure from the ledger rows, and the answer cites those rows.
Check the definitions before the figures. A variance against the wrong budget version is computed perfectly and still wrong.
| Task | Trust level | Why |
|---|---|---|
| Classifying accounts into a reporting hierarchy you have already defined | Trust, then sample | A pattern-matching job with a known answer set |
| Drafting month-end commentary from computed variances | Trust | The figures are fixed; only the words are generated |
| Explaining a bridge a tool has computed | Trust | The explanation can only rearrange figures already on the page |
| Mapping new GL accounts | Check | A guess at a new account can move a line between categories |
| Choosing the comparison period or budget version | Check | The answer is only as right as the plan it compares against |
| Totals, ratios and variances typed by the AI itself | Never unaided | Generated arithmetic is plausible, not reliable |
| Forecast figures presented as numbers | Never unaided | They belong to statistical models fitted on your data or your driver model |
The clearest public test is FinanceBench, a benchmark of questions about public companies' SEC filings. In the FinanceBench paper (November 2023), the authors report that GPT-4-Turbo used with a retrieval system "incorrectly answered or refused to answer 81% of questions" in their sample. That figure describes the systems tested at that date, not every model available today. The lesson that has not changed is the failure mode: a wrong answer arrives in the same confident sentence as a right one, so the reader cannot tell them apart without recomputing.
That is why the useful question is not "how accurate is the AI?" but "where did this number come from?". If the answer is "a tool computed it from these rows," you can check it. If the answer is "the assistant wrote it," you cannot.
The question typed: "Why did operating expenses exceed budget in Q2?" The data: the GL export for Q2 2026 and the approved budget, version B2. The external AI model picks the variance tool; the tool computes every column below; the external AI model writes the answer from the tool's output.
Variance = Actual − Budget (positive = overspend, for costs)
Variance % = Actual / Budget − 1
Share of net variance = Line variance / Total variance
| Account | Budget (USD) | Actual (USD) | Variance (USD) | Variance % | Share of net variance |
|---|---|---|---|---|---|
| Salaries | 1,240,000 | 1,262,000 | +22,000 | +1.8% | 25.3% |
| Contractors | 180,000 | 246,000 | +66,000 | +36.7% | 75.9% |
| Software subscriptions | 95,000 | 118,500 | +23,500 | +24.7% | 27.0% |
| Travel | 60,000 | 41,200 | −18,800 | −31.3% | −21.6% |
| Marketing programs | 210,000 | 204,300 | −5,700 | −2.7% | −6.6% |
| Rent and facilities | 150,000 | 150,000 | 0 | 0.0% | 0.0% |
| Total | 1,935,000 | 2,022,000 | +87,000 | +4.5% | 100.0% |
In Excel, with budget in B, actual in C, the variance in D and the total variance in D8:
D2: =C2-B2
E2: =C2/B2-1
F2: =D2/$D$8
The answer returned: operating expenses were $87,000 (4.5%) over budget B2 in Q2. Contractors explain $66,000, 76% of the net overspend; software subscriptions add $23,500 and salaries $22,000; travel ($18,800 under) and marketing ($5,700 under) offset part of it.
What the answer does not say matters as much. It does not say why contractor spend rose, because nothing in the ledger records why. It says so and lists the contractor invoices for review, the behavior described in when the assistant says "I cannot compute that". For the full method behind the table, see variance analysis.
Three checks, each one a line you can recompute:
In FP&A the time goes on the words around the numbers and on answering the same questions for many budget holders. That is where AI earns its place:
Forecasts are different. Forecast figures should come from statistical models fitted on your data or from the driver model the team owns, with the external AI model describing the result, never generating the numbers.
Settle four questions before any ledger leaves the building: what data is sent, where it is processed, how long it is kept, and who can see the answers. Send account-level totals when they answer the question, rather than payroll-level detail.
The NIST AI Risk Management Framework, released January 26, 2023 and voluntary, gives a common structure for this: map the use, measure the risks, manage them, and govern who is accountable. Banks have a longer-standing reference in the Federal Reserve's SR 11-7 guidance on model risk management (April 4, 2011), which expects validation, documentation and independent review of the models that inform financial decisions. An AI tool whose figures come from tested, deterministic calculations is far easier to bring inside either framework than one whose figures are generated text.
In Covirage the external AI model reads the question and picks the variance tool; the tool computes every line from your ledger and budget file, checks the lines sum to the total, and the answer cites the rows and the budget version. The external AI model explains the result and never does the arithmetic. See AI analytics for FP&A to ask budget and variance questions of your own ledger, the wider AI analytics overview, and AI data analyst for the same approach on sales data. For the report the variance questions come from, see budget vs actual; for the wider list of uses, artificial intelligence in business; for the team that does this work, what is FP&A.
It can read them, extract figures and describe trends, but its own arithmetic is unreliable. For analysis you act on, have tools compute ratios and variances from the source data and let the external AI model explain the computed results, with the rows cited.
Mainly to draft variance commentary, answer ad hoc questions from budget holders, classify transactions and summarize results. Forecast figures should come from statistical models fitted on your data or from your driver model, not from text generation.
Only if you know where the data goes, how long it is kept, whether it is used for training and who can see it. Remove what the question does not need, such as payroll detail, and prefer tools that process files without retaining them.
It removes routine work such as first drafts and data gathering. Choosing the right comparison, judging whether a variance matters and explaining it to the business still need an analyst, and someone has to answer for the numbers.