Decision-Focused Learning for Operating Room Scheduling Under Uncertain Surgery Durations with Two-Stage Stochastic Optimization

Abstract

Operating rooms (ORs) account for roughly 40% of hospital expenses and generate 60–70% of hospital revenue, making scheduling quality a primary lever for both cost control and timely patient access. However, high variability in surgery durations—driven by patient heterogeneity, procedure complexity, and unexpected intraoperative events—frequently causes cascading delays, staff overtime, last-minute cancellations, or unrecovered idle time. Standard scheduling approaches either rely on fixed duration estimates or decouple machine learning prediction from mathematical optimization under the Predict-then-Optimize (PtO) paradigm, training models to minimize prediction error (MSE) rather than the downstream quality of the resulting surgical schedule.

In this work, we present a framework combining Decision-Focused Learning (DFL) with two-stage stochastic programming (2SP) to build OR schedules that are both cost-aware at the prediction stage and adaptive at execution time. We formulate a modified scheduling objective using a predictive cost vector that balances accumulated patient waiting times against predicted surgery durations, and we benchmark two DFL architectures: an iterative Regret-Aware Gradient Boosting approximation and an end-to-end Smart Predict-Then-Optimize (SPO+) model trained via a convex surrogate loss that propagates optimization subgradients directly back to the duration predictor. Because even a decision-focused schedule remains vulnerable to intraoperative disruptions once execution begins, our second-stage Sample Average Approximation (SAA) recourse mechanism dynamically adapts schedules as true surgery durations are revealed in real time through three clinically grounded actions: opportunistic case advancement within a feasible preparation window, inter-OR swaps where surgeons follow their assigned patients, and bounded overtime cancellations.

We validate the framework on a real clinical dataset of 5,663 historical neurosurgical records across 114 surgeons from the Instituto de Neurocirugía Dr. Raúl Asenjo in Santiago, Chile. Benchmarking four prediction pipelines (a deterministic mean-duration baseline, Vanilla PtO, Regret-Aware Gradient Boosting, and SPO+) with and without 2SP recourse across 1,000,000 simulated OR days (100 patient queues × 10,000 Monte Carlo realizations), we show that DFL and stochastic recourse are strongly complementary. While SPO+ alone maximizes utilization by generating aggressive schedules, augmenting SPO+ with 2SP recourse absorbs the resulting tail risk—achieving 94.4% OR utilization, increasing daily throughput to 18.7 patients (up from 16.8 in the baseline), and reducing 95th-percentile overtime from 2.21 hours to 1.41 hours.

Finally, we analyze the operational and computational trade-offs required for clinical deployment. While SPO+ coupled with 2SP is ideally suited for overnight batch schedule generation (12.80 seconds per instance), Regret-Aware Gradient Boosting provides a lightweight alternative (1.36 seconds per instance) suitable for near real-time intraoperative rescheduling. Moreover, we demonstrate that the duration penalty parameter λ in the SPO+ objective serves not merely as a technical hyperparameter, but as an interpretable, governance-controlled policy knob that enables hospital administrators to explicitly map institutional overtime budgets and staffing constraints onto the utilization–overtime frontier.

Publication
European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD 2026)
Ali Elaswad
Ali Elaswad
The American University in Cairo
Rodrigo A. Carrasco
Rodrigo A. Carrasco
Associate Professor & Director of Data and Computing
Nouri Sakr
Nouri Sakr
The American University in Cairo

Related