December 7, 2025
CarbonQueue: Carbon-Aware Batch Job Scheduling System
A queue that schedules deferrable jobs in the cleanest feasible window before their deadline.
Computing demand is placing increasing pressure on grid capacity, but the carbon intensity of electricity also varies substantially by time of day as the generation mix changes. CarbonQueue addresses this temporal dimension by shifting flexible batch workloads toward lower-carbon periods without changing the underlying computation. The approach applies carbon-aware scheduling at the scale of a single machine or small team, where workloads can be deferred within explicit deadlines.
A job consists of a shell command, an estimated duration, a device power estimate, and a deadline. CarbonQueue forecasts hourly carbon intensity for the local grid zone, evaluates every feasible execution window before the deadline, and selects the one with the lowest mean forecast intensity over the job duration. After execution, it estimates CO2e emissions from the realized execution window and compares them with the corresponding estimate for immediate execution. The system also records scheduling decisions and job state changes in a durable event log for reproducibility and auditability.
System Design
Jobs are stored in a durable SQLite database, with each state change appended to an event log. This preserves job state across restarts and provides an auditable record of scheduling decisions. Jobs that fail three times transition to a dead-letter state for operator review rather than being retried indefinitely. The implementation applies standard task-queue and dead-letter patterns at single-machine scale, where persistence is particularly important for workloads that may remain queued for hours.
Carbon-intensity data is provided through two modes. Live mode uses a lightweight client to query the Electricity Maps API for the configured grid zone, while a cron-driven tick command performs each scheduler pass by fitting the forecaster, generating a forecast, placing queued jobs, and executing jobs that are due. Simulation mode instead replays an hourly trace from disk without network access. All results presented in this document use the simulation mode.
The forecaster uses ridge regression with lag values at 1, 2, 3, 24, 48, and 168 hours together with calendar features. It is evaluated through walk-forward validation against a seasonal-naive baseline that repeats the value from 24 hours earlier. The ridge model is selected only when it improves day-ahead MAE over the baseline. Otherwise, the baseline remains available as a fallback. Because the scheduler uses the forecast to rank candidate windows rather than predict an exact future value, the approach does not require a highly complex forecasting model.
Results
The reported run replays a 28-day hourly trace with 12 jobs drawn from six workload profiles, including model retraining, database backups, CI suites, and media transcoding. Jobs arrive at random times and have 24-hour deadlines. The synthetic trace includes an overnight plateau, midday solar dip, evening ramp, weekly cycle, and autocorrelated noise. It serves as a substitute for a cached historical export, and the simulation harness also accepts a historical trace as a CSV input. At each scheduling decision, the forecaster uses only observations available up to that point. Final emissions are calculated from the realized trace rather than the forecast.
Under these conditions, CarbonQueue avoided 42.5% of modeled emissions across the batch, with 930 g CO2e from scheduled execution compared with 1,617 g CO2e for immediate execution. The median execution delay was 9.5 hours. The day-ahead ridge forecast achieved an MAE of 26.7 gCO2e/kWh, compared with 29.3 gCO2e/kWh for the seasonal-naive baseline across six evaluation windows. The trace figure shows the resulting shift from higher-intensity evening periods toward lower-intensity overnight and midday periods.


Comparison with a perfect-foresight oracle provides a measure of forecast effectiveness. Selecting each job’s window from the realized trace would have avoided 44.4% of emissions, meaning the forecast-based scheduler captured 96% of the oracle’s achievable saving. This result shows that useful scheduling performance does not require highly precise intensity forecasts when the objective is to rank candidate windows. The deadline-slack sweep further shows the effect of scheduling flexibility. Savings increase from approximately 8% at six hours of slack to 42% at 24 hours, while the gap between forecast-based and oracle performance increases at longer horizons.

Scope and Limitations
Energy consumption is modeled rather than measured. Job energy is estimated from duration and nameplate device power, which does not account for utilization dynamics or idle consumption. The carbon-intensity signal represents grid-average rather than marginal intensity, and the two measures can differ in which hours they identify as lowest-carbon. Marginal intensity data is available through sources with different access requirements. The reported emissions reduction is derived from a synthetic trace and should therefore be interpreted as a demonstration of the scheduling mechanism rather than a measurement for a specific grid. The simulation harness accepts historical exports as a drop-in input without code changes. Finally, electricity prices may peak at different times than carbon intensity. CarbonQueue optimizes for carbon alone, while joint cost and carbon optimization remains future work.
Availability and Reproducibility
There is no hosted demo because the system is designed for multi-day scheduling and live mode requires a long-running process with an API key. The complete evaluation can be reproduced offline. The simulation command regenerates the summary JSON and all figures from scratch in under a minute on a machine with Python, while a smoke test covers the persistent store, scheduler, and dead-letter behavior in seconds. The summary file linked above is the authoritative source for the figures reported on this page.
Measured Power
Modeled power consumption is the least certain input in the current pipeline and can be replaced with measured energy. RAPL counters on Linux provide package-level energy measurements, while Kepler provides per-pod energy metrics for Kubernetes workloads. Measured joules per job would allow avoided emissions to be calculated from observed energy use rather than estimated power. The existing event log could also support calculation of a Software Carbon Intensity score for each workload. A future evaluation using measured power and an extended period of real grid data would provide a stronger validation of the scheduler.
Receipts
- Scope 12 jobs across six workload profiles were simulated over 28 days with 24-hour deadlines.
- Result CarbonQueue avoided 42.5 percent of modeled emissions at a median delay of 9.5 hours and captured 96 percent of oracle savings.
- Method Walk-forward evaluation used only information available at each scheduling decision, with emissions recomputed from the realized trace.
- Reliability SQLite persistence, an append-only event log, and dead-letter handling preserve and audit job state.
- Limits Results use modeled energy consumption, grid-average carbon intensity, and a synthetic trace rather than measured power or grid data.