Blog · AI and self-service analytics
What AI data analysis is, split into the part the AI model does well (reading the question, choosing the method, explaining the result) and the part it must leave to deterministic tools (the arithmetic). A worked gross-margin fall split into mix and rate shows the split in practice, with the checks that make an answer trustworthy and a five-question test for any tool.
AI data analysis means asking questions of your data in plain language and getting back computed answers with an explanation. Done well, it is a division of labor: the AI model reads the question, picks the calculation and explains the result, while deterministic tools (SQL, spreadsheet functions, statistics code) do the arithmetic on every row. Language models predict text, not sums, so the figures should never come from the text they write.
Two kinds of software do two different jobs.
The AI model: reads the question, maps it to columns, chooses the method, explains the result.
Deterministic tools: compute every figure from every row, the same way every time.
A deterministic tool returns the same answer from the same input, which is what makes a figure checkable. A language model generates the most likely next words, which is excellent for reading and writing and unreliable for adding up 4,000 invoice lines. Why AI gets numbers wrong explains the mechanism.
This is also the standard frameworks ask for. NIST's AI Risk Management Framework, released January 26, 2023, is built around incorporating "trustworthiness considerations into the design, development, use, and evaluation of AI products." For numbers, trustworthy means reproducible from the rows.
Inv_Amt_Net is net revenue and Cust_No is the customer key, and say so for you to confirm.Spreadsheet makers ship this pattern too: Microsoft describes Analyze Data in Excel as a way to "ask questions about your data without having to write complicated formulas," with Excel computing the result.
Why the AI model must never do the arithmetic sets out the principle in full.
A manager asks: "Revenue is up 10%, so why is gross margin down?" The file has revenue and gross profit by product line for two quarters.
| Product line | Revenue last Q (USD) | Margin last Q | Gross profit last Q (USD) | Revenue this Q (USD) | Margin this Q | Gross profit this Q (USD) |
|---|---|---|---|---|---|---|
| Core | 600,000 | 30% | 180,000 | 580,000 | 30% | 174,000 |
| Premium | 250,000 | 45% | 112,500 | 230,000 | 44% | 101,200 |
| Value | 150,000 | 18% | 27,000 | 290,000 | 18% | 52,200 |
| Total | 1,000,000 | 31.95% | 319,500 | 1,100,000 | 29.76% | 327,400 |
Revenue rose 10%, gross profit rose 2.5%, and gross margin fell 2.19 points. The AI model reads the question and chooses a mix and rate split. The tools compute it.
Gross margin % = total gross profit / total revenue
Mix effect = sum over lines of (share this period − share last period) × margin last period
Rate effect = sum over lines of share this period × (margin this period − margin last period)
| Product line | Share last Q | Share this Q | Mix term (points) | Rate term (points) |
|---|---|---|---|---|
| Core | 60.0% | 52.7% | −2.182 | 0.000 |
| Premium | 25.0% | 20.9% | −1.841 | −0.209 |
| Value | 15.0% | 26.4% | +2.045 | 0.000 |
| Total | 100.0% | 100.0% | −1.98 | −0.21 |
The line terms of the mix formula do not read well on their own, because each is measured against zero. Measured against the 31.95% average margin instead, which leaves the total unchanged, Value accounts for −1.59 points of the mix effect, Premium −0.53 and Core +0.14.
The explanation the AI model writes from those computed figures: margin fell mainly because the low-margin Value range grew fastest, from 15.0% to 26.4% of revenue, not because prices were cut. Only Premium lost rate, one point, worth 0.21 points of the total.
Totals that reconcile. The lines sum to 1,100,000 revenue and 327,400 gross profit, the same figures as the ledger.
An identity that holds. Mix plus rate must equal the change: −1.98 + (−0.21) = −2.19, and 29.76% − 31.95% = −2.19 points. If the two parts do not sum to the whole, the split is wrong.
A citation to the rows. Each figure links to the rows behind it, so a controller can open the Value lines and see 290,000 before the slide goes out.
An answer with all three can be checked in a minute. An answer with none of them has to be rebuilt by hand before anyone can rely on it.
Run this on any tool before trusting it. AI analytics tools compared lists the tools; this test is how to judge them on your data.
Can ChatGPT analyze sales data? runs a test of this kind on a general chat assistant.
Upload what the question needs. A margin question needs product line, revenue and cost; it does not need customer names, emails or phone numbers. Where a customer must be identified, a pseudonymized key such as an account number does the job and keeps personal data out of the tool. Check how the vendor stores, retains and uses the file before the first upload; is it safe to upload customer data to an AI analytics tool lists the questions to ask.
In Covirage, the AI model reads the question and picks the tool. Deterministic tools compute the figures on every row, check identities such as mix plus rate equals the change, and cite the rows behind each figure; the external AI model then explains and never does the arithmetic. Ask your own files a question: the tools compute it, the AI model explains it. For the same approach applied to a dashboard, see AI dashboard. For the role it plays on a team, see AI data analyst, and for finance questions, see AI for financial analysis.
Chat assistants can write and run code on an uploaded file in some versions, which is much safer than arithmetic in text. The reliability depends on whether the figures come from executed code or generated text, and on checking totals yourself. Features vary by version and plan.
It changes the work more than it removes it. AI speeds up writing queries, drafting explanations and answering routine questions. Defining measures, checking data, and judging what a result means for a decision still need people who know the business.
It depends on the tool's terms: whether data is used for training, where it is stored, who can access it and how long it is kept. Remove or pseudonymize personal data where you can, and check the vendor's security documentation first.
The one that gives correct, checkable answers on your own data. Test candidates with questions whose answers you already know, including one they cannot answer from the file; the right response to that one is to say so.