How merit-order surveillance reconstructs a delivery hour

A day-ahead price is one number standing in for thousands of orders nobody outside the exchange sees. This view rebuilds what the available stack should have produced, compares it against what cleared, and is specific about the hours where the rebuild is the weaker of the two.

What the screen answers

Given a delivery day and one or more bidding zones: which hours cleared at a price the published fundamentals do not explain, and how far past the line each one sits. The board is the whole selection at a glance; opening a zone gives one hour's full derivation.

It answers this for an analyst who already has an alert from somewhere else. It is a second opinion on an hour, not a first detection.

The controls, and what each one changes

Zones
Which bidding zones are on the board. Zones with no ingested generation for the window are marked as such rather than silently dropped.
Delivery day
The day under review. The calibration window is the days BEFORE it, never including it.
Calibration lookback
How many prior days the cost parameters are fitted on. A longer window is a steadier fit and a slower read; it also decides how far back the trailing residual spread is measured.
Screen sensitivity
Moves every threshold at once - standard, or roughly two-thirds of it. The reconstruction does not change: the same hours are rebuilt with the same numbers, and only the line drawn through them moves. Every count on the page is at the setting currently selected, and the setting rides in the URL, so a link shared from a loosened line opens loosened.

How the reconstruction is built

Five steps, in the order the engine runs them. The zone view shows this for the selected hour with the real numbers at each step, and each step ends with the assumption most worth doubting.

1. Inputs
Realised generation per production type, the day-ahead load forecast, installed and available capacity, and the day's outage notices. Where nothing is stored for the window, these are read live at request time.
2. Availability
Installed capacity less what the outage notices remove, per technology. This is the ceiling each fleet could have produced against, and it is the step most often wrong: capacity unavailable for a reason nobody published looks, to the model, like withholding.
3. Cost
Each technology placed by short-run marginal cost, with must-run and zero-marginal-cost output at the base. The cost parameters are the fitted part.
4. Demand and imports
Residual demand for the hour, met from the stack. Net imports enter as a modelled block priced at the median of the interconnected neighbours - modelled, because the physical flow is published but the price it was contracted at is not.
5. Clearing
The block meeting the last megawatt sets the reconstructed price. The gap between that and the settled price is the residual every screen works on.

Why calibration is walk-forward

Cost parameters for a delivery day are fitted only on days that strictly precede it. A fit that has seen the day it scores has already absorbed the very anomaly the screens exist to find, and would report a quiet residual for the hour that most deserves attention.

The residual distribution from that fit is the model's error, and it is out-of-sample by construction. It is reported as MAE and bias on the board, and beside every finding in the zone view. The screens are aware of it: a margin has to beat the reconstruction's own residual spread before it counts, so an imprecise model produces fewer findings, not more.

The screens

Ten screens run on the fundamentals. Each states what was observed with the numbers in it, and none of them claims abuse.

Cleared above fundamentals
The settled price sits materially above the reconstruction, measured in the zone's own residual spread rather than a fixed euro amount. Severity rises where the residual clears roughly 1.7x the threshold.
Cleared below fundamentals
The mirror case. Reported at warn only: a discount has fewer abusive readings than a premium and more benign ones.
In-merit capacity idle
Capacity whose modelled marginal cost sat below the clearing price and which did not generate. Both an absolute floor in MW and a share of demand have to be cleared, so a small zone and a large one are held to comparable evidence.
Demand above modelled supply
The stack could not meet residual demand at all. Always high severity, and usually a coverage problem rather than a market one - it is the first thing to check against the sources table.
Thin reserve margin
Available capacity was close to demand. Graded info or warn, never high: a tight hour explains a high price rather than indicting it.
Steep merit order at the clearing point
The euro move per 100 MW at the margin. A steep clearing point is where a small withheld block moves the price a lot, so it is context that raises the value of the other screens.
Pivotal supplier
The residual supply index below 1: demand could not have been met without that fleet. Deliberately graded low - at technology-fleet granularity this is common in a concentrated zone, and it is context rather than an alert unless deeply pivotal.
Cleared above the whole stack
The settled price exceeded the most expensive modelled block. Either the stack is missing a technology or the hour cleared on something the fundamentals do not carry.
Import-dependent clearing
The net-import block set the price. The import block is modelled, not published, so an hour that clears on it carries more model error than one clearing on a domestic technology.
Coupling divergence
The zone's price diverged from its coupled neighbours by more than the threshold, which is the shape a congested or decoupled hour takes.

A further set runs on the exchange's published auction results rather than on fundamentals: price z-score against the trailing same-hour baseline, price-limit hits, buy/sell one-sidedness, and priced-but-untraded intervals. That lens needs no fundamentals at all, so it reaches market areas the reconstruction cannot, and it sits alongside rather than replacing it.

One delivery period across every auction

A delivery period clears more than once. Day-ahead settles it first, then the intraday auctions IDA1, IDA2 and IDA3 re-clear the same period as the gate moves closer. The zone view puts all of them side by side for the period you have open.

Findings on the intraday auctions are read alongside the day-ahead assessment, never added to it. Each auction is screened against its own trailing band, because each is its own distribution; and an intraday move happens after day-ahead's gate has already closed, so it cannot be evidence about how day-ahead cleared. An hour with a striking intraday finding and an ordinary day-ahead price does not climb the zone's triage queue on that account.

Coverage is uneven and the view says which kind of gap it is looking at. IDA2 and IDA3 in particular do not run for every period, so an auction absent from one hour of a day it otherwise covers is a gap in the period rather than a gap in collection. A zone in none of the clusters that publish aggregated curves has no intraday data at any hour, which is a scope limit rather than a gap at all.

The board, column by column

Zone
The bidding zone. A row appears only where realised generation per production type was ingested for the calibration window; a zone without it is listed as uncovered rather than shown empty.
Risk
The composite triage score, 0-100. It orders a queue. It is not a probability, and two hours with the same score are not equally interesting - read the findings.
Severity
The highest severity among the zone's findings that day: info, warn or high. A zone whose only findings are info has raised context, not an alert.
Hours with alerts
Delivery hours where at least one screen fired at the sensitivity currently selected. Changing the sensitivity changes this count and nothing else about the reconstruction.
Near misses
Hours where no screen fired but one came close. This column exists because a threshold is a line drawn through a continuum, and the hours just short of it are the ones an alerting engine would hide.
Top finding
The highest-ranked single finding in the zone that day, named so the board can be scanned without opening every zone.
Model MAE (day)
The reconstruction's out-of-sample mean absolute error for that zone on that day, in EUR/MWh. A large MAE is a weak basis for calling an hour anomalous, and the number sits beside the findings rather than in a footnote.
Bias
Whether the reconstruction ran systematically high or low that day. A persistent offset is a model-quality problem; an isolated spike against a well-behaved baseline is what the screens are for.

The zone-hour grid and the full ledger

The grid is one cell per zone-hour for the whole day, coloured by whichever metric you choose, whether or not a screen fired. That is the point of it: an alerting engine that publishes only what crossed hides the block that was 180 MW when the screen wanted 200. Every metric the grid can be coloured by is also a column in the ledger below it.

The ledger is the complete reconstruction for the selection, nothing filtered and nothing rounded away, ordered however you choose and exportable as CSV. It carries twenty columns, including both risk score and screen pressure, the closest unfired screen, and how close that screen came as a percentage of its threshold. It is the table to read when you distrust the ranking.

How a near miss is ranked

The under the radar queue holds hours that raised no alert, ordered by how close the nearest screen came to firing. Each row names the binding screen and the distance it still had to travel, so the queue can be read as evidence rather than as a list of near-things.

The composite risk score, where a screen did fire, sums a per-screen weight times how far past its threshold the finding sits, capped at twice the threshold so one extreme reading cannot dominate, then maps the total onto 0-100 through a saturating curve - roughly 55 at one strong finding and 85 by three. It orders a queue and nothing more.

Where this can be wrong

The reconstruction is an estimate and the settled price is a fact. Where the two disagree about how steep an hour was and a published aggregated curve exists, believe the curve: it is the book, and the reconstruction is the thing with a reported error.

Legitimate explanations the public data cannot see produce every shape these screens look for: unit commitment and start-up constraints, heat obligations, network constraints, and hedged positions. The feeds carry no participant identity either, so a finding attributes to a technology fleet in a zone and never to a firm. Verify against outage disclosures and the exchange's own data before drawing any inference.

Read next