Files
challenge_ai/PLAN.md
T
mattandClaude Opus 5 54e877db9b Add exercise documents and battery dispatch implementation plan
Commits the three source documents for the battery trading exercise
(task brief, battery specification, and three years of half-hourly
Market 1 and hourly Market 2 prices) together with PLAN.md, the
implementation plan for the dispatch model.

The plan covers a rolling-horizon MILP on the half-hourly grid with a
48-hour lookahead and 24-hour commitment, solved with PuLP/CBC, plus
degradation and cycle-life reporting, an independent schedule
validator, charts and tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-24 12:47:12 +01:00

15 KiB
Raw Blame History

Battery dispatch model — implementation plan

Context

The folder contains a second-round technical exercise (2nd Round Technical Question.pdf) asking for a Python package that models a 2 MW / 4 MWh battery charging and discharging across two wholesale electricity markets to maximise profit, as a price taker. Attachment 1.xlsx gives the battery specification, Attachment 2.xlsx gives three years of prices. Nothing has been built yet — this plan covers the package from scratch. The exercise is explicitly judged on practical design, clear communication of process and outcomes, and a codebase that is easy to review — not on squeezing out the last pound of revenue.

Decisions already taken with the user: model the full three years across both markets, use PuLP with the bundled CBC solver, and include degradation / cycle-life reporting plus charts and a written results file. Investment economics (NPV, IRR, payback) and heuristic benchmark comparisons are deliberately out of scope.

What the data actually looks like

Verified by reading both workbooks directly:

Item Value
Market 1 half-hourly, 52,608 rows, 2018-01-01 00:00 → 2020-12-31 23:30, 465 negative prices, mean £44.45/MWh
Market 2 hourly, 26,304 populated rows (rows 26,306+ of the sheet are blank and must be dropped), mean £47.96/MWh
Battery 2 MW charge, 2 MW discharge, 4 MWh, 5% charge loss, 5% discharge loss, 10 yr / 5,000 cycles, 1e-3 %/cycle fade, £500k capex, £5k/yr opex

Two data quirks the loader must handle explicitly:

  1. Market 2 trailing blanks — the sheet is 52,608 rows long but only the first 26,304 carry prices.
  2. Spring clock-change labelling — Market 1 timestamps duplicate 02:00/02:30 and skip 01:00/01:30 on 2018-03-25, 2019-03-31 and 2020-03-29. Row counts are nevertheless a clean 48/day (1,096 × 48) and Market 2 is a clean 24/day (1,096 × 24). The loader will rebuild a regular half-hourly index from the series start, assert the expected row counts, log the discarded anomalies, and align Market 2 to Market 1 by integer position (2 half-hours per hour) rather than by timestamp join. This is documented as an assumption in the README.

Sizing sanity check: mean daily Market 1 spread is £39/MWh and hourly-averaged Market 1 correlates with Market 2 at 0.983 (mean absolute gap £4.22/MWh), so a two-cycle-per-day strategy lands in the tens of thousands of pounds per year. Cross-market optimisation adds value but intraday arbitrage dominates.

Approach

Discretise everything onto the half-hourly grid — the finest market resolution — and express dispatch as a mixed-integer linear program solved over a rolling horizon: a 48-hour lookahead window of which only the first 24 hours is committed, with state of charge carried forward. This keeps each solve tiny (96 binaries) while avoiding the myopia of independent daily solves, and it makes the three-year run tractable in minutes rather than attempting one 52,608-period monolith.

Market 2's hourly commitment rule is enforced structurally: one decision variable per hour that applies to both of that hour's half-hours, so the constraint is impossible to violate rather than merely checked. The "cannot sell the same unit of energy into both markets" rule falls out for free from a single shared state-of-charge balance.

Because both price series contain negative prices, the LP relaxation can genuinely prefer to charge and discharge simultaneously (get paid at both ends). The solver therefore runs LP-first with an MILP fallback: solve the relaxation, detect any window where charge and discharge are both non-zero in the same half-hour, and re-solve that window with binaries. This is exact, and fast in the common case.

Formulation

Per solve window, index half-hours by t and hours by h, with H(h) the two half-hours in hour h. Δt = 0.5 h, η_c = η_d = 0.95, P_c = P_d = 2 MW, E_max = current degraded capacity.

Variables:

  • c1[t], d1[t] ∈ [0, P] — MW committed to Market 1 for half-hour t
  • c2[h], d2[h] ∈ [0, P] — MW committed to Market 2 for the whole of hour h
  • e[t] ∈ [0, E_max] — energy stored at end of t
  • z[t] ∈ {0,1} — charging indicator (MILP pass only)

Constraints for every half-hour t, with h = hour(t):

c1[t] + c2[h]  <=  P_c * z[t]
d1[t] + d2[h]  <=  P_d * (1 - z[t])
e[t] = e[t-1] + Δt * ( η_c * (c1[t] + c2[h]) - (d1[t] + d2[h]) / η_d )
0 <= e[t] <= E_max

e[-1] is the carry-in state of charge from the previous window. In the LP pass, z[t] is dropped and the two power constraints become c1+c2 <= P_c, d1+d2 <= P_d.

Objective (maximise):

Σ_t  Δt * p1[t] * (d1[t] - c1[t])  +  Σ_h  1.0 * p2[h] * (d2[h] - c2[h])

Discharge accounting matches the spec: to deliver D MWh to the grid the battery draws D / 0.95 MWh from storage; to store S MWh it imports S / 0.95 MWh from the grid — i.e. revenue and cost are settled on the grid-side quantity, which is what the market pays for.

Degradation is applied between windows rather than inside the LP (which would make it non-linear): after each committed day, cumulative equivalent full cycles are updated as cycles += discharged_from_storage_MWh / E_nominal, and usable capacity is set to E_max = E_nominal * (1 - 1e-5 * cycles) for the next window.

Package layout

Built in /Users/matt/aurora_task as a git repository (git init locally; nothing pushed anywhere without asking). The two spreadsheets are copied into data/ so the repo is self-contained and reproducible.

pyproject.toml            # deps: pandas, openpyxl, pulp, matplotlib; dev: pytest
README.md                 # setup, how to run, assumptions, one-paragraph approach summary
RESULTS.md                # headline outcomes, tables, embedded charts
data/                     # copies of Attachment 1.xlsx and Attachment 2.xlsx
outputs/                  # committed reproducible results (csv, json, png)
src/battery_dispatch/
  __init__.py
  config.py               # BatterySpec + RunConfig dataclasses; parse Attachment 1
  data.py                 # price loading, index rebuild, count assertions, market alignment
  markets.py              # Market dataclass (name, periods_per_halfhour, prices)
  optimiser.py            # window problem construction, LP/MILP solve, rolling horizon driver
  battery.py              # SoC state, cycle counting, capacity fade
  validation.py           # independent post-hoc schedule checker
  metrics.py              # annual/monthly KPIs, cycle and capacity reporting
  plots.py                # matplotlib figures
  cli.py                  # `battery-dispatch run ...`
tests/
  test_data.py
  test_optimiser.py
  test_validation.py
  test_battery.py

Implementation steps

Task 1 — Data and configuration

  1. config.py: BatterySpec dataclass (max_charge_mw, max_discharge_mw, capacity_mwh, charge_efficiency, discharge_efficiency, lifetime_years, lifetime_cycles, degradation_pct_per_cycle, capex_gbp, fixed_opex_gbp_per_year) with a from_excel() classmethod that reads data/Attachment 1.xlsx by row label, converting the two loss fractions into efficiency = 1 - loss. Also RunConfig (start, end, window hours, commit hours, solver options, output directory, optional degradation cost).
  2. data.py: load_prices() returning a single half-hourly DataFrame with columns market_1_price, market_2_price, hour_index. Drops blank rows, asserts 1,096×48 and 1,096×24 rows, rebuilds the regular index, logs clock-change anomalies, broadcasts hourly prices onto half-hours, and exposes a slice(start, end) helper that clips to whole days.

Task 2 — Optimiser

  1. optimiser.py: build_window(prices_window, spec, soc_start, capacity, relaxed: bool) returning a PuLP problem plus a handle on the variable dictionaries. One c2/d2 variable per hour, referenced by both of its half-hours.
  2. solve_window(): solve relaxed; if any half-hour has both charge and discharge above 1e-6, rebuild with binaries and re-solve. Raise on non-optimal solver status rather than silently returning a bad schedule. Expose --milp-always / --lp-only overrides.
  3. run_rolling_horizon(): iterate windows of window_hours committing commit_hours, carry SoC, update capacity fade via battery.py after each commit, and concatenate committed half-hours into one result frame (timestamp, p1, p2, c1, d1, c2, d2, soc, capacity, revenue_m1, revenue_m2).

Task 3 — Battery state and degradation

  1. battery.py: BatteryState tracking SoC, cumulative charged/discharged MWh, equivalent full cycles and degraded capacity; apply_cycles() implementing the 1e-3 %/cycle fade; helper for implied end of life, min(10 years, 5000 / cycles_per_year).

Task 4 — Independent validation

  1. validation.py: validate_schedule(result, spec) re-checking, from the output frame alone and without reference to solver internals:
    • 0 <= c1 + c2 <= max_charge, 0 <= d1 + d2 <= max_discharge
    • no half-hour with both charge and discharge non-zero (tolerance 1e-6)
    • Market 2 power identical across both half-hours of every hour
    • SoC recomputed from scratch stays in [0, capacity] and matches the reported SoC
    • revenue recomputed from prices and powers matches the reported revenue Returns a structured report; the CLI fails loudly if any check fails.

Task 5 — Metrics, plots, CLI

  1. metrics.py: totals and per-year / per-month breakdowns — gross trading profit split by market, MWh charged and discharged, equivalent full cycles, capacity remaining, average captured spread £/MWh, share of revenue from each market.
  2. plots.py: (a) one representative week showing both prices, dispatch by market and SoC; (b) monthly revenue stacked by market; (c) cumulative cycles and capacity fade over three years; (d) distribution of captured spread.
  3. cli.py: battery-dispatch run --start 2018-01-01 --end 2020-12-31 --out outputs/ writing dispatch.csv, summary.json, and the PNGs; prints a compact summary table and the validation verdict. Deterministic: no randomness anywhere, fixed solver settings.

Task 6 — Tests

  1. tests/test_optimiser.py — small hand-checkable cases:
    • two half-hours priced £0 then £100, Market 1 only → charges then discharges; revenue equals 2 MW × 0.5 h × 0.95 × 0.95 × £100 = £90.25
    • an hour whose two half-hours have opposite Market 1 signals → confirms Market 2 power stays flat across the hour
    • negative prices in both markets → confirms the MILP fallback triggers and no simultaneous charge/discharge survives
    • a case where the optimum splits power across both markets within one hour, respecting the 2 MW cap
  2. tests/test_data.py — blank-row dropping, row-count assertions, hourly→half-hourly broadcast.
  3. tests/test_validation.py — validator catches each injected violation type.
  4. tests/test_battery.py — cycle counting and capacity fade arithmetic.

Task 7 — Documentation and committed results

  1. Run the full three years, commit outputs/ and write RESULTS.md: headline gross profit, per-year table, revenue split between markets, cycles per year against the 5,000-cycle budget, capacity fade, solver statistics and wall-clock runtime.
  2. README.md: install (uv sync or pip install -e .), run commands, the modelling assumptions (efficiency convention, clock-change handling, rolling-horizon perfect foresight within the window, degradation applied between windows), known simplifications, and the required one-paragraph summary.

Acceptance criteria

  • pip install -e . (or uv sync) succeeds on Python 3.11+ with no non-pip dependencies; CBC ships with PuLP.
  • battery-dispatch run --start 2018-01-01 --end 2020-12-31 completes in under 10 minutes on a laptop and exits 0.
  • validate_schedule reports zero violations on the full three-year run, checked independently of the solver.
  • Every half-hour satisfies c1+c2 <= 2 MW, d1+d2 <= 2 MW, 0 <= SoC <= capacity, and never both charges and discharges (tolerance 1e-6).
  • Market 2 committed power is constant within every hour, verified across all 26,304 hours.
  • Revenue recomputed independently from dispatch.csv matches the optimiser objective to within £0.01.
  • pytest passes, including the £90.25 analytic two-period case to within £0.01.
  • Re-running the CLI twice produces byte-identical summary.json.
  • RESULTS.md reports gross profit per year, revenue split by market, cycles per year, and end-of-life capacity, each traceable to a file in outputs/.

Verification

  1. pytest -q — unit tests, including analytic known-answer cases.
  2. battery-dispatch run --start 2018-01-01 --end 2018-01-07 --out /tmp/smoke — fast smoke run; eyeball the week plot against the price series to confirm it buys the overnight trough and sells the evening peak.
  3. battery-dispatch run --start 2018-01-01 --end 2020-12-31 --out outputs/ — the reproducible headline run.
  4. Confirm the validation report at the end of the full run shows all checks passed.
  5. Spot-check a single day by hand: pull one day from dispatch.csv, recompute revenue and SoC in a scratch script, and confirm agreement.
  6. Edge cases to exercise: a day containing negative prices in both markets, the three spring clock-change days, and the first and last day of the series (window boundary behaviour).

Risks and mitigations

  • CBC runtime on 1,096 windows. Mitigated by the LP-first strategy — binaries only appear in the rare windows where the relaxation cheats. If it is still slow, widen commit_hours or add a --max-seconds-per-window solver cap; both are config, not redesign.
  • Rolling horizon is not globally optimal. A 48-hour window with 24-hour commitment captures essentially all intraday value for a 2-hour-duration battery. This will be stated plainly in the README as a deliberate trade-off rather than left implicit.
  • Efficiency convention could be read the other way. The spec's wording is unambiguous (losses occur before storage on import and before the grid on export), but the convention will be written out in the README and pinned by a unit test so a reviewer can see exactly what was assumed.
  • Clock-change timestamps. Rebuilding a regular index is the pragmatic choice given clean 48/day counts; the discarded anomalies are logged so the decision is visible rather than hidden.
  • Scope creep into investment economics. Explicitly excluded per the user's decision; the package reports cycles and capacity fade but stops short of NPV, keeping the deliverable within the exercise's time box.