Predictive analytics uses statistical models fitted on historical data to estimate a future or unknown value. This page sets it beside descriptive, diagnostic and prescriptive analytics, gives examples from finance and sales, describes the main methods in plain words, and tests a trend forecast on quarters it has not seen, against a naive baseline.
Predictive analytics uses statistical models fitted on historical data to estimate a value you do not know yet: next quarter's revenue, which customers will stop buying, which invoices will be paid late. A prediction is only worth the error it shows on data it was not fitted on. In the worked example below, a straight trend line misses two unseen quarters by 1.26% on average, against 6.75% for simply repeating the last quarter.
Analytics is usually split into four kinds, by the question each answers.
| Kind | Question | Finance and sales example |
|---|---|---|
| Descriptive | What happened? | Revenue by region last quarter |
| Diagnostic | Why did it happen? | Which customers and products drove the fall |
| Predictive | What is likely to happen? | Next quarter's revenue; which accounts may churn |
| Prescriptive | What should we do about it? | Which overdue accounts to call first |
Predictive analytics is the third row. It takes the patterns in past data, fits a method to them, and applies that method to cases or periods where the answer is not known yet.
Regression. Fits a line, or a surface, through past data so one value can be estimated from others. A straight trend over time is the simplest case: revenue = intercept + slope × quarter number. NIST's handbook describes the idea as splitting the variation in one quantity into a deterministic function of other quantities plus a random part.
Time-series methods. Use the series' own past. Exponential smoothing weights recent periods more than old ones and can carry trend and seasonality; ARIMA describes how each period depends on earlier periods and earlier errors. Both are covered in Hyndman and Athanasopoulos, Forecasting: Principles and Practice.
Classification. Estimates the chance of a yes-or-no outcome: churns or not, pays late or not, wins or not. Logistic regression gives a probability from a weighted sum of inputs; decision trees split the rows by simple rules; gradient-boosted trees combine many small trees, each correcting the last.
Eight quarters of revenue, in millions of dollars. Fit the trend on Q1-Q6 only and hold Q7 and Q8 back as the test.
| Quarter | Revenue (USD m) | Trend fit on Q1-Q6 (USD m) | Trend error | Naive: repeat Q6 (USD m) | Naive error |
|---|---|---|---|---|---|
| Q1 | 4.10 | ||||
| Q2 | 4.32 | ||||
| Q3 | 4.45 | ||||
| Q4 | 4.71 | ||||
| Q5 | 4.86 | ||||
| Q6 | 5.02 | ||||
| Q7 | 5.30 | 5.225 | 1.42% | 5.02 | 5.28% |
| Q8 | 5.47 | 5.410 | 1.10% | 5.02 | 8.23% |
| MAPE | 1.26% | 6.75% |
On Q1-Q6, the fitted trend has a slope of 0.1851 a quarter and an intercept of 3.9287:
Trend = a + b × t, with b = SLOPE(y, t) and a = INTERCEPT(y, t)
So Q7 is 3.9287 + 0.1851 × 7 = 5.225 and Q8 is 5.410. With quarter numbers 1-8 in A2:A9 and revenue in B2:B9, the forecast for Q7 in C8 is:
=FORECAST.LINEAR(7,B2:B7,A2:A7)
which returns 5.225. Microsoft describes FORECAST.LINEAR as predicting a future value "by using linear regression". The error per quarter in D8 is:
=ABS(B8-C8)/B8
and MAPE is the average of those errors:
MAPE = average of |actual − forecast| / actual
The trend misses Q7 by 1.42% and Q8 by 1.10%, a MAPE of 1.26%. The naive baseline, which repeats Q6's 5.02, misses by 5.28% and 8.23%, a MAPE of 6.75%. The trend beats the baseline on quarters it never saw, so it has earned its place. Only then refit on all eight quarters: slope 0.1946, intercept 3.9029, giving Q9 = 5.655 and Q10 = 5.849.
Hold-out test. Fit on the earlier data, test on the later data, and report the error on the test only. Hyndman and Athanasopoulos put it plainly: "A model which fits the training data well will not necessarily forecast well."
A naive baseline. Compare against "same as last period". The same book treats simple methods as benchmarks: if a new method does not beat them, "the new method is not worth considering."
An error measure people can read. MAPE is a percentage, so it compares across series of different size. How to measure forecast accuracy and bias in Excel adds bias, the tendency to miss in one direction.
A range, not a point. Report a prediction with the range it is likely to fall in. In Excel, FORECAST.ETS.CONFINT returns the half-width of a confidence interval for a FORECAST.ETS forecast.
Neutral, and in rough order of effort:
=FORECAST.ETS(target_date,values,timeline)), and LINEST for regression with several inputs.The right choice depends on how much data there is, how many series, and who will keep the work running after the first version.
Statistical models fitted on your data make the prediction, and they can be tested for error on periods they have not seen. A language model can explain the result in plain words: which quarters drove the trend, why the range is wide, what would change the estimate. It should not produce the numbers, because its output cannot be tested the same way and does not show the rows behind it. Why the language model must never do the arithmetic sets out the reasoning.
Covirage fits statistical models on your own history with its tools, tests them against a naive baseline on periods they have not seen, and reports the error beside every prediction; the external AI model explains the result and never produces the numbers. See AI analytics for how the two parts divide the work, and what is a good forecast accuracy for what error to accept. For the Excel functions in detail, see the FORECAST function in Excel.
Estimating which customers are likely to stop buying next quarter from their order frequency, recency and support history; forecasting next month's revenue from past months; or scoring open invoices by the chance they will be paid late.
Predictive analytics estimates what is likely to happen. Prescriptive analytics recommends what to do about it, often by optimizing a decision such as price, inventory level or which accounts to call first.
Not exactly. Most predictive analytics uses statistical and machine-learning methods fitted on historical data. Language models can explain predictions in plain words, but the prediction itself should come from methods that can be tested for error.
Spreadsheets with forecasting functions for simple series; Python and R for full modeling; BI and data platforms with built-in forecasting; and specialist tools for demand and churn. The right choice depends on data size and who maintains the work.