A typical structure for this case is a five-part plan: define the decision problem, model censored demand, validate forecasts chronologically, convert forecasts into inventory decisions with asymmetric costs, and roll out safely while tracking waste and availability.
1. Clarify the business problem
Start by pinning down the operational question.
- Decision unit: SKU-store-day, SKU-region-week, or product-category-store?
- Decision cadence: daily ordering, weekly replenishment, or both?
- Objective: minimize total expected cost, where cost includes waste, lost sales, holding cost, and possibly markdowns.
- Constraints: shelf life, delivery lead time, minimum order quantities, supplier lead times, storage capacity.
- Data available: sales, inventory, orders, stockout flags, waste/spoilage records, promotions, prices, holidays, weather, store attributes, product hierarchy.
The core challenge is that observed sales are not true demand.
2. Handle censored demand
For each item-store-time unit, define:
Dit=true latent demand
Sit=available inventory
Observed sales are:
Yit=min(Dit,Sit)
When the item sells out, sales are right-censored: we know true demand was at least the observed sales, but not how much higher.
Useful modeling choices:
- Use a censored likelihood with Poisson, negative binomial, or Tweedie distributions.
- Use a Tobit-style regression or survival analysis if demand is continuous.
- Use gradient boosting or quantile regression with a custom loss that treats censored observations as lower bounds.
- If stockout indicators are available, treat censored observations explicitly. If not, infer likely stockouts from zero ending inventory or unusually low sales before replenishment.
Features typically include seasonality, trend, price, promotion, holiday, weather, day of week, store cluster, product category, and lagged demand.
3. Time-based validation
Do not use random cross-validation for this problem. Demand is temporal and autocorrelated.
Use walk-forward validation:
- Train on data up to time t.
- Forecast t+1,…,t+h.
- Evaluate against actuals.
- Advance the training window and repeat.
This can be an expanding window or a sliding window. For probabilistic forecasts, evaluate:
- Pinball loss at each quantile.
- CRPS for the full predictive distribution.
- Quantile coverage, e.g., the 90% interval should contain the actual about 90% of the time.
- Sharpness, meaning intervals should not be unnecessarily wide.
Business-level backtesting should also estimate waste, availability, and profit under the recommended order quantities.
4. Probabilistic forecasts
The model should output a predictive distribution, not just a point forecast.
Candidate methods:
- Quantile regression with LightGBM/XGBoost.
- Distributional models: Poisson, negative binomial, Tweedie.
- Hierarchical models that share information across stores and products.
- Deep learning models such as DeepAR or Temporal Fusion Transformer if data volume is large enough.
Typical output quantiles:
5%,25%,50%,75%,95%
These quantiles feed directly into the inventory decision.
5. Asymmetric stocking costs and optimal order quantity
Overstocking and understocking have different costs.
Let:
co=cost per unit overstocked
cu=cost per unit understocked
For food:
- co includes waste, disposal, markdowns, and holding cost.
- cu includes lost margin, lost customer loyalty, and possible substitution effects.
The newsvendor critical ratio gives the optimal service level:
Q∗=F−1(cu+cocu)
where F is the predictive distribution of demand.
If waste cost is high, the optimal quantile is lower. If lost sales and availability are more expensive, the optimal quantile is higher.
6. Waste and availability metrics
Track these explicitly after choosing an order quantity Q:
Expected waste:
E[(Q−D)+]
Expected lost sales:
E[(D−Q)+]
Total expected cost:
coE[(Q−D)+]+cuE[(D−Q)+]
Availability or service level:
P(D≤Q)
These metrics should be reported by product category, store, region, day of week, and season.
7. Rollout plan
Do not deploy everywhere at once.
Recommended rollout:
- Pilot in a small number of stores or regions.
- Run the model in shadow mode first: generate recommendations but do not execute them.
- Compare model recommendations with current ordering behavior.
- Use a staggered rollout or A/B test across comparable stores.
- Monitor forecast drift, stockout flags, data quality, and retraining frequency.
- Build fallback rules for missing data or model failure.
Success should be measured by:
- Reduction in waste.
- Maintenance or improvement of availability.
- Increase in profit or reduction in total cost.
- Model calibration and coverage over time.
The final system should output a recommended order quantity per SKU-store-day by minimizing expected asymmetric cost using a probabilistic forecast built from censored demand data and validated through time-based backtesting.