Files
challenge_ai/README.md
T
mattandClaude Sonnet 5 c70c9731a6
CI / test (push) Successful in 14s
Run the full three-year dispatch and fix issues it surfaced
Executed battery-dispatch run over the full 2018-2020 Attachment 2
series (225.9s, 1096 rolling-horizon windows) and committed the
results: outputs/ (dispatch.csv, annual/monthly summaries, summary.json,
four figures), README.md (setup, run instructions, assumptions, the
one-paragraph approach summary) and RESULTS.md (headline numbers,
per-year table, charts, commentary).

The full run surfaced three real issues, all fixed rather than papered
over:

- validate_schedule's soc_within_bounds check used a fixed 1e-6 MWh
  tolerance for the "state of charge below zero" case, while its
  sibling soc_matches_power_flows check already scales its tolerance
  with sqrt(n) for the same reason (CBC's own solver precision
  accumulating over a long cumulative sum). Over 52,608 half-hours this
  false-failed on a 5.1e-6 MWh solver-noise dip, not a real violation.
  Scaled it the same way.
- cli.py's summary printed net_of_opex_gbp under the label "Less fixed
  opex", so the terminal output showed the post-opex profit (£193k)
  where a reader would expect the opex figure itself (£15k). Split into
  two correctly-labelled lines.
- plots.py's monthly revenue chart stacked Market 1 (always negative
  here) and Market 2 (always positive) with a single running `bottom`,
  which is only correct for same-signed series: Market 2's bar
  completely overlapped and painted over Market 1's, hiding the
  -£297k Market 1 loss entirely. Now accumulates positive and negative
  contributions on separate baselines and marks the net.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-24 16:29:40 +01:00

132 lines
7.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Battery dispatch model
A Python package that dispatches a 2 MW / 4 MWh battery across two wholesale electricity
markets — Market 1 (half-hourly prices) and Market 2 (hourly prices) — to maximise trading
profit, as a price taker, over the three years of price data in `data/Attachment 2.xlsx`. Battery
specifications come from `data/Attachment 1.xlsx`.
## Approach (one paragraph)
Dispatch is formulated as a linear/mixed-integer programme (PuLP + the bundled CBC solver) on
the half-hourly grid, solved over a rolling 48-hour lookahead window of which only the first 24
hours is committed and state of charge carried forward — this keeps each solve small while
avoiding the myopia of independent daily solves. Market 2's hourly commitment rule is enforced
structurally (one decision variable per hour, referenced by both of its half-hours) rather than as
a constraint that could be mis-specified, and "no selling the same energy twice" falls out for free
from a single shared state-of-charge balance. Because negative prices and cross-market price gaps
let the LP relaxation profitably "cheat" by charging and discharging at once, the solver runs
LP-first with a MILP fallback that adds binary exclusivity only where the relaxation actually
cheats. Every schedule is independently re-validated from its output alone (power limits, no
simultaneous flow, hourly commitment, state-of-charge and revenue recomputed from scratch) so a
formulation bug couldn't mark its own homework, and cycle counting / capacity fade are tracked
between windows per Attachment 1's degradation figures. Investment economics (NPV, IRR, payback)
and heuristic benchmark comparisons were deliberately left out of scope to keep the deliverable
focused on the dispatch decision itself.
## Setup
Requires Python 3.11+. CBC ships with PuLP, so no external solver install is needed.
```bash
uv sync # or: pip install -e ".[dev]"
```
## Running
```bash
# The full three-year run (~4 minutes; this is what produced outputs/)
uv run battery-dispatch run --start 2018-01-01 --end 2020-12-31 --out outputs/
# A fast smoke run over one week
uv run battery-dispatch run --start 2018-01-01 --end 2018-01-07 --out /tmp/smoke
# Tests (38 tests, ~1.5s)
uv run pytest -q
# Lint
uv run ruff check .
```
`battery-dispatch run` writes `dispatch.csv` (every half-hour's prices, power, state of charge and
revenue), `annual_summary.csv`, `monthly_summary.csv`, `summary.json`, four PNG figures, and prints
a headline summary plus the independent validation verdict. It exits non-zero if validation finds
any violation. Useful flags: `--lp-only` / `--milp-always` to force a solver mode,
`--max-seconds-per-window` to cap solve time per window, `--degradation-cost` to price cycle life
into the objective (0 by default — the exercise asks for gross trading profit), `--no-plots` to
skip figure rendering.
## Reproducible results
`outputs/` is committed and holds the full three-year run described above. See
[`RESULTS.md`](RESULTS.md) for headline numbers, charts and commentary. Re-running the CLI with
the same arguments reproduces `summary.json` byte-for-byte (no randomness anywhere, fixed solver
settings).
## Modelling assumptions
- **Efficiency convention.** Attachment 1 quotes charge/discharge *losses* (5% each); the model
treats these as `efficiency = 1 - loss`. Importing `I` MWh from the grid stores `0.95 × I`;
exporting `E` MWh to the grid draws `E / 0.95` from storage. Money is always settled on the
grid-side quantity, since that's what the market meters and pays for. This convention is stated
once in `config.py` and pinned by a unit test (`test_spec_reads_losses_as_efficiencies`, plus the
£90.25 analytic two-period case in `test_optimiser.py`).
- **Clock-change labelling.** Market 1's timestamps duplicate `02:00`/`02:30` and skip
`01:00`/`01:30` on the three spring clock-change days, but every day is still a clean 48 rows.
Rather than guess at intended local-time semantics, the loader rebuilds a regular half-hourly
index from the series start and logs the discarded labels — the series is treated as 52,608
consecutive half-hours, which is what the dispatch model actually needs.
- **Market alignment.** Market 2's sheet is exactly as long as Market 1's but only the first
26,304 rows carry prices; the rest are blank and dropped (with an assertion on the surviving row
count). Market 2 is aligned to Market 1 by integer position (two half-hours per hour) rather
than by timestamp join, since hourly timestamps carry a few milliseconds of floating-point noise
that makes a timestamp join fragile.
- **Rolling-horizon perfect foresight.** Each 48-hour window solves with perfect knowledge of
prices within that window, then commits 24 hours and slides forward. This is not globally
optimal, but a 48-hour window comfortably covers the useful lookahead for a 2-hour-duration
battery — intraday arbitrage dominates the value here (mean daily Market 1 spread is ~£39/MWh),
and cross-window boundary effects are small (see the capacity-seam note in `validation.py`).
- **Degradation applied between windows, not inside the LP.** Making usable capacity a function of
cumulative throughput within the same optimisation would make the programme non-linear. Instead,
equivalent full cycles and the resulting capacity fade (Attachment 1's 0.001%/cycle) are updated
*after* each committed 24-hour block and used as the capacity bound for the next window. Over a
single window this changes capacity by ~1e-5 MWh, far below any decision-relevant threshold.
- **LP-first, MILP fallback.** The LP relaxation can profitably charge and discharge in the same
half-hour whenever prices are negative or the two markets' prices diverge enough to beat the
~10% round-trip efficiency loss. The solver detects this and re-solves only the offending window
with binary exclusivity added. On this dataset the fallback fires in essentially every window
(the persistent Market 1/Market 2 price gap makes it attractive almost daily), but each MILP
solve is still fast in practice — confirmed directly against Attachment 2 at ~0.1s/window,
~4 minutes for the full three years.
## Known simplifications
- No investment economics (NPV, IRR, payback) — the model reports gross trading profit, cycles and
capacity fade, but does not attempt to value the £500k capex against them.
- No heuristic/benchmark comparison strategy — only the optimised dispatch is reported.
- A rolling 48h/24h window is a deliberate trade-off against a single monolithic 52,608-period
solve, stated explicitly rather than left implicit.
- Degradation cost is not priced into the objective by default (`--degradation-cost 0`), so the
model reports the profit-maximising *gross* dispatch; a positive `--degradation-cost` is
available to see how marginal cycling cost would suppress trading.
## Package layout
```
src/battery_dispatch/
config.py BatterySpec + RunConfig; reads Attachment 1
data.py Price loading, index rebuild, market alignment
markets.py Market dataclass (name, commitment block length, prices)
optimiser.py Window construction, LP/MILP solve, rolling-horizon driver
battery.py State of charge, cycle counting, capacity fade
validation.py Independent post-hoc schedule checker
metrics.py Annual/monthly KPIs, headline summary
plots.py Matplotlib figures
cli.py `battery-dispatch run ...`
tests/ 38 tests covering data loading, the optimiser, validation and battery state
```
## CI
`.gitea/workflows/ci.yml` runs lint (`ruff`) and the full test suite on every push, and posts a
pass/fail notification to ntfy.