Amazon Web Services · Amazon EC2 Auto Scaling

Predictive scaling — ML forecast-directed capacity scale-out

Amazon EC2 Auto Scaling predictive scaling uses machine-learning-based forecasting of historical workload patterns to predict future demand and derive hourly capacity requirements. In ForecastAndScale mode, Amazon EC2 Auto Scaling can automatically scale out an Auto Scaling group when forecast capacity exceeds current capacity, causing EC2 instances to be launched ahead of predicted demand. Predictive scaling itself does not perform forecast-driven scale-in when predicted demand falls; AWS documents dynamic scaling as the mechanism for removing excess capacity.

Recorded characteristics

Function
Predictive scaling analyses up to 14 days of historical load data for an Auto Scaling group, detects daily and weekly patterns, and produces hourly load forecasts for the next 48 hours, refreshed every 6 hours. AWS then calculates future capacity needs from the load forecast and the customer-configured target utilisation; the GetPredictiveScalingForecast API returns both a load forecast and a capacity forecast. Core action statement: Historical workload observations → ML-derived future-load forecast → AWS calculates forecast capacity using the forecast and customer-configured target utilisation → ForecastAndScale automatically applies forecast-directed scale-out → Amazon EC2 Auto Scaling launches EC2 instances to satisfy the increased capacity requirement. Public evidence supports an intermediate capacity calculation; it does not state that an ML model directly outputs the final EC2 instance count. ForecastOnly vs ForecastAndScale (negative control): in ForecastOnly mode AWS generates capacity forecasts but does not scale the Auto Scaling group based on those forecasts. In ForecastAndScale mode the forecast capacity is applied automatically at the start of each hour (or earlier when a scheduling buffer is configured). Forecasting ≠ infrastructure execution. Scale-out boundary: predictive scaling can increase capacity in anticipation of forecast demand. A forecast decrease does not cause predictive scaling itself to scale in; AWS states customers must use dynamic scaling policies to remove excess capacity. Predictive scaling does not decide when to terminate EC2 instances. Dynamic scaling interaction: dynamic scaling operates independently from predictive scaling. Predictive scaling can provide anticipated/baseline capacity while dynamic scaling responds to current demand and can perform scale-in. Predictive scaling does not have exclusive authority over the group's final capacity. Outside action: Amazon EC2 Auto Scaling acts beyond its immediate service boundary by using its service-linked role to launch EC2 instances in the customer's AWS account. This follows the Registry's immediate product/service boundary, not ownership of the target environment; no outside organisation or third-party infrastructure is involved. Human confirmation: an administrator configures predictive scaling and enables ForecastAndScale, but AWS does not document individual human approval before each forecast-directed scale-out action. Configuration authorisation is not runtime confirmation. Cost: forecast-directed capacity increases can alter AWS resource consumption and therefore may affect EC2 costs. This is not financial authority. Explicit exclusions: ForecastOnly; manual scaling; administrator-authored scheduled scaling; target tracking without predictive forecasting; step scaling; simple scaling; ordinary CloudWatch alarm → deterministic scaling; dynamic scale-in; forecast-driven instance termination; health-check replacement; AWS Compute Optimizer recommendations; human-reviewed capacity recommendations; generic EC2 APIs; generic agent/tool calls invoking EC2; Spot-capacity mechanisms; ordinary desired-capacity enforcement unrelated to predictive forecasting.
Data access
Uses the Auto Scaling group's historical CloudWatch load metric data (up to 14 days; AWS may use temporary forecasts from aggregate data when history is short). Execution identity: the service-linked role AWSServiceRoleForAutoScaling (or a custom-suffix variant), which trusts the service principal autoscaling.amazonaws.com and carries the AutoScalingServiceRolePolicy permissions to call other AWS services (including launching and terminating EC2 instances) on the customer's behalf. Auto Scaling receives service-specific delegated permissions rather than simply inheriting the configuring user's interactive permissions. The role belongs to the Auto Scaling service, not a distinct predictive-scaling agent. Auditability: customers can view load forecast, capacity forecast, actual capacity and launched-instance count graphs, target utilisation and policy configuration, CloudWatch predictive-scaling metrics, the GetPredictiveScalingForecast API, Auto Scaling scaling activity, and service-role actions for auditing. Operational tracing is available, but complete ML decision provenance is not publicly established; administrators cannot reconstruct the model's complete reasoning for an individual forecast.
Actions
Can take actions
External actions
Yes
Human confirmation
Not required
Permission basis
Separate permissions
Administrative control
Customer establishes the operational envelope; ML-derived forecast demand materially determines forecast capacity within that envelope. Customer-configured controls: Auto Scaling group; scaling/load metric; target utilisation; minimum capacity; maximum capacity; mode (ForecastOnly / ForecastAndScale); scheduling buffer; maximum-capacity behaviour (HonorMaxCapacity, the default, or IncreaseMaxCapacity); maximum-capacity buffer. IncreaseMaxCapacity: where enabled, predictive scaling can increase the group's maximum capacity according to forecast capacity and the configured buffer. AWS warns that a raised maximum becomes the new normal maximum for the group until it is manually updated; it does not automatically decrease. Reversibility: administrators can return the policy to ForecastOnly, modify the policy, delete the policy, change minimum/maximum capacity, and reset an automatically raised maximum. These controls stop or alter future forecast-directed execution. Predictive scaling itself does not undo previous forecast-driven launches through predictive scale-in. Blast radius: direct controlled object is the capacity requirement for an Auto Scaling group. A forecast-derived scale-out can result in multiple EC2 instances being launched. Capacity is normally constrained by the configured group maximum; where IncreaseMaxCapacity is enabled, forecast capacity plus the configured buffer can alter that boundary. Universal maximum automated capacity increase: Not Publicly Established.
Default state
Disabled
Availability
Current Amazon EC2 Auto Scaling feature with current user-guide and API documentation including supported-Regions information; no preview or deprecation marking. Default: "When you first enable predictive scaling, it runs in forecast only mode." The PredictiveScalingConfiguration API "Mode" setting defaults to ForecastOnly when not specified. The disabled classification rests on this explicitly documented non-executing default, not merely on configuration being required.
Licensing
Predictive scaling is a feature of Amazon EC2 Auto Scaling; forecast-directed capacity increases may affect EC2 costs. Specific pricing for the feature is not asserted here.
External model or provider
AWS first-party forecasting. The explicit machine-learning characterisation comes from the AWS Compute Blog introduction of native predictive scaling ("Predictive scaling uses machine learning to predict capacity requirements based on historical usage and continuously learns...") and the current Amazon EC2 Auto Scaling product page ("Use machine learning to predict and schedule the right number of EC2 instances"). Current AWS technical documentation describes historical-data analysis, pattern detection, forecasting and automatic scaling but does not itself label the mechanism as machine learning. The explicit ML characterization is supplied by separate first-party AWS product and technical material.
Limitations and uncertainty
Evidence qualification: current technical documentation does not itself use the term "machine learning"; explicit ML characterization comes from other first-party AWS sources (2021 AWS Compute Blog; current product page). This is not an implication that ML is absent. Not Publicly Established: exact production ML architecture; constituent production models; model/version responsible for an individual forecast; complete training dataset; complete feature set/weighting; ensemble/model-selection logic; confidence intervals; calibration methodology; full internal forecast-to-capacity implementation beyond documented relationships; forecasting-service outage behaviour; stale-forecast behaviour; retry/backoff semantics; treatment of every missing/late metric-data condition; complete precedence across every possible scaling-policy combination; universal maximum automated instance increase; complete ML-to-individual-instance causal provenance; current cross-customer learning/data boundaries; model-update cadence; independent customer reproducibility of AWS forecasts; internal model-quality thresholds; how temporary/backfill forecasts from aggregate data are produced. Execution identity is established (service-linked role AWSServiceRoleForAutoScaling, principal autoscaling.amazonaws.com) and is not marked Not Publicly Established.

Evidence