Aller au contenu principal
OceanData Consulting
Scoreboard · method

How these figures are produced

What the scoreboard measures, with which data, and what its figures are not yet worth. Without this page, the table cannot be verified.

← Back to the scoreboard

Principle

A post-processing of the physical model, not a replacement

The AI does not produce a forecast from scratch: it corrects that of the official physical model.

Every day at 06:00 UTC, the pipeline retrieves the reference physical forecast for each station, then a post-processing model trained on that same station’s history corrects its output over the next 48 hours.

For significant wave height, this reference is not the same everywhere: five wave models are compared station by station — Météo-France (MFWAM), ECMWF WAM, DWD EWAM and GWAM, NCEP GFS-Wave — and the one that forecasts the site best becomes its reference. The AI is therefore compared with the best of the five, not with a model chosen once and for all; the selected model is named in the table and on each station page. For water level, the reference is a harmonic tidal analysis.

The post-processing model itself is chosen automatically, station by station, at training, and its identity is named in the same place as the physical reference. Three families exist: a gradient boosting (histogram gradient boosting, “hgb”), a regularised linear regression (“ridge”), or a separate gradient boosting per lead-time band — 1 to 12 h, 13 to 24 h, 25 to 48 h (“hgb-per-lead”). A station running ridge is not one where the model failed to do better: it is the project’s comparison floor, linear regression already capturing the gain at that station without a more sophisticated model adding anything measurable. The model name is published only when it can be identified; otherwise none is invented and nothing is displayed in its place.

The model takes as input the physical forecast itself, the recent error of that forecast against observations, and a 10 m wind forcing. For tide gauges, it predicts only the residual — the difference between the observation and the harmonic — and the published level is rebuilt as “harmonic + predicted residual”.

The next day, the forecast issued the day before is compared with the observations that actually arrived, and the day’s mean absolute error is computed for the AI and for the physics alike, on exactly the same points. Nothing is rewritten after the fact: a day without usable observations stays marked as missing and enters no average.

All timestamps of the scoreboard, its API and its charts are in UTC.

Publication

A station that gains nothing publishes no AI figure

A station is published only if the model gains at least 5% of MAE over the physical forecast there. The others stay displayed, without a figure.

Below that threshold, the model finds no usable signal in the current data for that station. It stays listed on the scoreboard, tracked, and without an AI figure — removing it would make the others unverifiable. The scoreboard table names the ones currently in this case, and marks it on each row.

A station can also pass the threshold while carrying an additional caveat: the model does not beat its own baseline there once that baseline is debiased. Its displayed gain is then first of all a constant removed from the physical forecast, not meteo-oceanic knowledge. These stations are flagged as such, in the table and on their detail page.

Stated limits

What these figures are not yet worth

Four caveats that bear on the displayed gain itself. They are measured, not rhetorical.

Part of the gain is only a bias correction

Training evaluation of 3 August 2026, displayed gain and then the gain once only the mean bias is removed from the physical forecast — Les Pierres Noires +25.8% → +25.9%; Belle-Île +14.1% → +11.7%; Cherbourg +19.3% → +3.0%; Anglet +9.8% → +5.4%; Brest +22.9% → −5.4%; Saint-Malo −13.4% → −14.1%. At Les Pierres Noires, the gain owes nothing to bias. At Cherbourg and Brest, it comes almost entirely from it: this is why the first is not published and the second carries a caveat. The figure to quote is always the second.

Tide gauges are still trained on reanalysis wind

A model trained on a wind known after the fact, then served with a forecast wind, is evaluated too favourably: the forecast carries a lead-time error that the reanalysis does not. The retraining of 3 August 2026 lifted this caveat on the wave stations — their training forcing now comes from archives of forecasts actually issued (ARPEGE Europe, ECMWF IFS, ICON-EU), exactly the kind of data received in production. It remains in full on the two tide gauges, still trained on ERA5 reanalysis: their figures are optimistic until they are retrained in turn.

The 25–48 h lead times are scored one day late

A day is first scored only on the lead times whose observation already exists, i.e. the first 24 hours; the rest of the horizon is marked pending and completed in a later run, when the observations arrive. A recent day can therefore be shown over 24 h first, then over 48 h. No value is rewritten: only missing lead times are added.

The reference’s long lead times are approximated at training

The training baseline of the wave stations is rebuilt from Open-Meteo forecast archives, which mostly restore the freshest runs: distant lead times are better represented there than a real 48 h run would have given. The model has therefore learned to correct a reference that is slightly too good at long lead times. The pipeline archives the forecast actually served every day to lift this approximation in turn.

Sources and credits

Where the data comes from

The attributions below are licence obligations attached to the data used, not acknowledgements.

  • Candhis — Cerema

    Wave observations (Hs) from the four buoys: Les Pierres Noires, Belle-Île, Anglet, Cherbourg offshore.

    Données de houle in situ : réseau Candhis, Cerema.

  • REFMAR — SHOM

    Water-level observations from the RONIM tide gauges of Brest and Saint-Malo, in near real time.

    Données marégraphiques : SHOM / REFMAR.

  • Open-Meteo

    Wave forecasts from the five compared physical models — Météo-France (MFWAM), ECMWF WAM, DWD EWAM and GWAM, NCEP GFS-Wave — including the one selected per station as its reference. Also 10 m wind: ARPEGE Europe, ECMWF IFS and ICON-EU forecast archives for training the wave stations, ERA5 reanalysis for the tide gauges, ARPEGE Europe forecast for inference.

    Weather and marine data by Open-Meteo.com — Météo-France MFWAM and ARPEGE, ECMWF WAM and IFS, DWD EWAM/GWAM and ICON, NCEP GFS-Wave, and ERA5 reanalysis (Copernicus Climate Change Service / ECMWF), served under CC BY 4.0.

    Open-Meteo’s free tier is reserved for non-commercial use. This scoreboard is public, with no advertising or paid access: the use is non-commercial and stated as such. Monetising this page would require moving to Open-Meteo’s paid tier or to the primary sources (CDS for ERA5, Météo-France for ARPEGE).

Embed the badge

Display a station’s verdict on your own site

A compact iframe badge with the station’s 7-day MAE and a link to its detail page on the scoreboard. The attribution line under the badge is part of the code to copy: a link placed inside the iframe remains, for a search engine, a link from our site to itself — only the one placed in your page signals the source.

<iframe src="https://oceandataconsulting.fr/scoreboard/widget/brest" width="320" height="180" style="border:0" loading="lazy" title="Metocean Scoreboard — brest"></iframe>
<p><a href="https://oceandataconsulting.fr/en/scoreboard/brest">Metocean AI Scoreboard — OceanData Consulting</a></p>
Who publishes this scoreboard

OceanData Consulting

This scoreboard is produced and published by OceanData Consulting, an independent firm specialising in ocean modelling and marine data analysis. The pipeline that feeds it is open: the code, the station configuration, the publication verdicts and the daily result files can be consulted and replayed.