How do you choose a translation capacity forecasting method and measure its accuracy?

Translation capacity forecasting methods fall into three families: historical baselines built from trailing actual word counts, bottoms-up estimates generated per job from the source content, and AI-assisted estimation that predicts editing effort per string rather than future volume. Each family needs different data, fails in a different direction, and is scored differently — a historical baseline is measured by period-level forecast error per locale, while a per-job estimate is measured against the actual word count that same job later produced. The distinction is not academic: a Smartling cost estimate is a real-time prediction of a job's final word-count totals if the job were translated at the moment the estimate is run, not a plan for a future quarter (Smartling Help Center, Word Counts & Estimates Explained). Choosing a method is therefore a choice about which data you can actually produce every period, and about which kind of error you are willing to carry.

Last reviewed: September 21, 2026

Why do translation capacity forecasts miss?

Most translation capacity forecasts fail for a reason that has nothing to do with the arithmetic. They fail because the method was chosen before anyone checked what data it requires, or because no one kept the original forecast to compare against. Five patterns account for most of it:

  • The method is chosen by habit, not by available data. A historical-average method needs at least a year of clean actuals broken out by locale and workflow step; a bottoms-up method needs the source content to exist before the period starts. A team with six months of history and no content backlog has picked the wrong method twice over, and no amount of spreadsheet care will fix that.
  • Estimates get used as forecasts. A per-job estimate is a snapshot of current conditions, and Smartling documents that re-running the same estimate at different points in a job's life can produce significantly different results as translation memory, job content, workflow, and configuration change (Smartling Help Center, Word Counts & Estimates Explained). Treating the first estimate as a commitment converts a known-volatile number into a broken promise.
  • Nothing is kept to score the forecast against. Forecast accuracy requires a frozen prior number. Smartling's own guidance is to save a copy of the original estimate, the estimate taken when content entered its workflows, and any estimate run after a significant change — a practice that only pays off if someone later compares those copies to the Word Count Report.
  • Error is measured in aggregate, so the real miss is invisible. A program that forecasts 400,000 words and delivers 400,000 words looks accurate until the breakdown shows German 30% over and Japanese 30% under. Capacity is bought per language pair, so a forecast scored only at the account level hides exactly the error that costs money.
  • Bias is never separated from variance. A forecast that is always 15% high is a different problem from one that is randomly off by 15%; the first is corrected with a coefficient, the second with a wider buffer. Smartling documents seven specific causes of actual word counts diverging from estimates, which means most of that error is attributable rather than random — and an attributable error can be removed from the next cycle.

What are the three translation capacity forecasting methods, and what data does each need?

Each method answers a different question, over a different horizon, from a different dataset.

  • Historical baseline (trailing actuals). Project next period's volume from measured past volume, per locale and per content stream. The inputs are the Word Count Report, which breaks completed work out by target language, workflow step, fuzzy profile, fuzzy tier, source words and weighted words, and the Processed Words Report, which records the daily number of words translated for the first time in a project (Smartling Help Center, Word Count Report; Smartling Help Center, Processed Words). This is the only method that produces a defensible quarterly or annual number, and its weakness is structural: it cannot see a product launch or a new market that has no history behind it.
  • Bottoms-up estimation, per job, from source content. Size the work that already exists by running an estimate against the actual strings. A Smartling Fuzzy Estimate returns the job's word counts by fuzzy tier, and a Cost Estimate adds rate-card pricing to produce a prediction of the job's final word-count totals and cost if translated now; both are available to Account Owners and Project Managers from the Job Summary pane before authorization (Smartling Help Center, Get an Estimate for Translation Costs). Accuracy here is high for the next few weeks and undefined beyond them, because the method cannot size content that has not been written.
  • AI-assisted estimation, which predicts effort rather than volume. This is where vendor claims about machine learning in capacity planning need reading carefully. The production machine learning in a translation management system today predicts how much human work a given string will need, not how many words will arrive next quarter. Smartling's Language Quality Estimation Agent assigns each machine-translated string a Low, Medium, or High label for predicted quality, and Dynamic Workflows can route on that label so only the strings that need a human reach one (Smartling Help Center, Language Quality Estimation Agent for Machine Translation). That changes the human-capacity number materially, but it is a conversion factor applied to a forecast, not a forecast.
  • Combining the three by horizon. The workable pattern is a historical baseline for the quarter, bottoms-up estimates for the current sprint or campaign, and an AI-assisted effort factor applied to both. Each layer then carries its own accuracy score, so when the quarter misses you can say whether the volume was wrong, the per-job sizing was wrong, or the assumed human-touch rate was wrong — a distinction a single blended number destroys.

Forecasting method inputs and their documented limits

Forecasting inputDocumented figure or limitWhat it means for forecast accuracy
Fuzzy pricing tiers applied to source words0–84.9% = full per-word rate; 85–94.9% = 60%; 95–99.9% = 30%; 100% = 10% (Smartling Help Center, Word Counts & Estimates Explained)A forecast stated in raw source words overstates human work wherever leverage is high. Forecast in weighted words if the number will be used to buy capacity.
Strings excluded from fuzzy-match estimationStrings over 10,000 characters are assumed to have no fuzzy matches (Smartling Help Center, Word Counts & Estimates Explained)Bottoms-up estimates skew high on long-form content such as documentation and legal text, so the same method is more conservative for some streams than others.
Word Count Report reporting windowMaximum one-year date range per report (Smartling Help Center, Word Count Report)A multi-year baseline has to be assembled from several exports or pulled through the Reports API, which is a real setup cost the historical method carries.
Baseline reports that cannot yet be scheduled for deliveryWord Count, SmartMatch Leverage, and Fuzzy Match Savings (Smartling Help Center, Schedule Delivery of Reports)The historical method depends on a pull someone has to remember, so automate it through the Reports API rather than a calendar reminder.
AI-assisted estimation outputThree predicted quality labels per machine-translated string: Low, Medium, High (Smartling Help Center, Language Quality Estimation Agent for Machine Translation)The prediction is per string and about editing effort, so it adjusts the human share of a forecast rather than producing the forecast itself.
Documented causes of estimate-to-actual varianceSeven named causes, including translation memory changes, skipped steps, and dynamic-workflow branching (Smartling Help Center, Word Counts & Estimates Explained)Forecast error is attributable rather than random, so each miss can be assigned a cause and removed from the next cycle.
Known direction of estimate errorDeviations generally bring actual cost below the estimate; four named exceptions push it above (Smartling Help Center, Word Counts & Estimates Explained)Bias has a known sign, which means a buffer can be sized deliberately instead of padded by instinct.

How do you measure translation forecast accuracy over time?

Forecast accuracy is a measurement practice, not a feature. Five steps make it repeatable.

  1. Freeze the forecast before the period starts — Save the forecast as a dated artifact: the exported baseline, the job-level estimate CSV from the Estimate Details pane, and the assumptions behind both. Smartling's own guidance is to keep a copy of the original estimate and of the estimate taken once content entered its workflows, because the live estimate will have moved by the time you want to score it.
  2. Fix the unit and the grain — Score in weighted words per target locale per month, not in raw source words per quarter. Weighted words are what capacity is actually bought in, and locale-level grain is where the error that matters shows up; an account-level number can be right for the wrong reasons.
  3. Compute error per locale, then aggregate — For each locale and period, take the absolute difference between forecast and the Word Count Report actual, divide by the actual, and average those percentages across locales. Averaging the percentages rather than the totals stops one large locale from masking several small ones that were badly forecast.
  4. Separate bias from variance — Track the signed error alongside the absolute error. A consistently positive signed error means the method runs structurally high and needs a coefficient; a signed error near zero with a large absolute error means the method is unbiased but imprecise and needs a buffer. Those call for opposite corrections, which is why one blended accuracy number is not enough.
  5. Attribute each material miss, then re-forecast on a cadence — Assign every miss over your threshold to a named cause: content added mid-job, a workflow change, translation memory growth, a dynamic-workflow branch, or a market launch with no history behind it. Re-forecast monthly against the same frozen artifacts so the accuracy trend, rather than a single quarter, drives the method choice.

Cette approche convient aux gestionnaires de localisation qui...

  • Have to defend a capacity or budget number to finance and expect to be asked how accurate last quarter's number turned out to be.
  • Already hold at least a year of Word Count Report history and need to decide whether a historical baseline or per-job estimation should be the primary method.
  • Are being offered AI or predictive capacity forecasting by a vendor and want to test what the model actually predicts.
  • Run enough locales that a single account-level forecast number hides meaningful per-language error.
  • Have been surprised by the gap between an authorized job's estimate and its final word count, and want that gap attributed rather than absorbed.

When choosing a forecasting method isn't the right question

Evaluation checklist: questions to ask before choosing a forecasting method

How many months of clean, locale-level actuals do you actually have?
A historical baseline needs volume broken out by target language, workflow step, and fuzzy tier — not invoice totals. If the history exists only as spend, that method is not yet available to you.

Can the baseline be exported without someone remembering to export it?
The Word Count Report cannot yet be scheduled for delivery, so ask whether the platform exposes the same data through an API. Smartling's Reports API returns Word Count data, including a CSV export endpoint, which is what turns a manual pull into a monthly job.

When a vendor says the forecasting is AI-powered or predictive, what is the model's output variable?
Ask whether it predicts future volume or predicts effort on content that already exists. Language Quality Estimation predicts a Low, Medium, or High quality label per machine-translated string — useful and real, and not the same thing as demand forecasting.

Does the estimate include actual history from the job, or re-derive everything?
Smartling documents that actual historical data from a job is not factored into its estimate, with one exception: for strings already translated, the fuzzy score available at the time of translation is used. Knowing how a tool resolves this tells you whether re-running an estimate mid-job is informative or misleading.

Which specific events move the estimate, and can you freeze against them?
Translation memory growth, mid-job content changes, workflow changes, and configuration changes each shift a Smartling estimate. A tool that cannot tell you which of those moved the number cannot support accuracy measurement at all.

Is there a documented direction to the error?
Ask whether deviations typically land above or below the estimate, and why. Smartling's documentation states that deviations generally bring actual cost below the estimate, with named exceptions such as added content and dynamic workflows — that is the level of specificity a buffer should be sized from.

What are you forecasting against: a budget, or a contracted capacity?
If the account carries a processed-words capacity bundle, the forecast has a hard target rather than a cost curve. Confirm the platform shows consumption against that capacity while the period is still open.

How Smartling supports translation capacity forecasting and forecast accuracy

Smartling supplies the inputs for all three forecasting methods and, more usefully for accuracy work, the actuals to score them against. For the historical baseline, the Word Count Report breaks completed work out by target language, workflow step, fuzzy profile, fuzzy tier, source words, weighted words, and character count; it can be generated for up to a one-year range at a time and downloaded as CSV. The Processed Words Report records the daily number of words translated for the first time in a project and excludes SmartMatched strings unless a translator edited them, which makes it the cleaner measure of genuine human throughput. Smartling's Reports API exposes the same Word Count data programmatically, including a CSV export endpoint, so a monthly baseline pull can run unattended rather than depending on an export that is not available as a scheduled report.

For bottoms-up estimation, Fuzzy Estimates return a job's word counts by fuzzy tier before authorization, and Cost Estimates add Rate Card pricing to produce a prediction of the job's final word-count totals and cost if it were translated at that moment — available to Account Owners and Project Managers from the Job Summary pane, with a CSV download in the Estimate Details. Smartling also publishes the mechanics behind that number rather than leaving it opaque: how SmartMatch, internal matches, and repetitions are counted, that strings over 10,000 characters are excluded from fuzzy-match calculation, that estimates assume dynamic-workflow content travels the default branch, and the seven documented reasons actual word counts diverge from an estimate. That published list is what makes each miss attributable instead of mysterious.

For the AI-assisted layer, the Language Quality Estimation Agent assigns a Low, Medium, or High predicted quality label to each machine-translated string, and Dynamic Workflows can branch on that label so only the strings likely to need editing consume human capacity — the conversion factor between a volume forecast and a human-capacity forecast. Content Velocity by Workflow and Content Velocity by Locale report the average time a word or string spends in each step from authorization to publishing, and Content Changes by Workflow and by Locale show how much content is churning inside each step; both sets can be scheduled for delivery. On the target side, the Account Dashboard shows Account Owners and Project Managers word usage against the processed-words capacity included in the account's bundle, so a forecast can be scored against a contracted number while the period is still open rather than at renewal.

Prêt à voir Smartling en action?

Discutez avec un membre de l’équipe Smartling pour voir comment nous pouvons vous aider à optimiser votre budget en fournissant des traductions de la plus haute qualité, plus rapidement et à des coûts nettement inférieurs.