Skip to content

to decode ensemble members in chunks - #2752

Open
kctezcan wants to merge 1 commit into
ecmwf:develop-ssl-diffusion-v1from
MeteoSwiss:develop-ssl-diffusion-v1-2
Open

to decode ensemble members in chunks#2752
kctezcan wants to merge 1 commit into
ecmwf:develop-ssl-diffusion-v1from
MeteoSwiss:develop-ssl-diffusion-v1-2

Conversation

@kctezcan

@kctezcan kctezcan commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

Description

When running diffusion inference with more than 1 member, we run into cudea OOM issues.

This patch deodes the ensembles in groups/chunks in a tradeoff of time vs memory.

Issue Number

No issue

Is this PR a draft? Mark it as draft.

Checklist before asking for review

  • I have performed a self-review of my code
  • My changes comply with basic sanity checks:
    • I have fixed formatting issues with ./scripts/actions.sh lint
    • I have run unit tests with ./scripts/actions.sh unit-test
    • I have documented my code and I have updated the docstrings.
    • I have added unit tests, if relevant
  • I have tried my changes with data and code:
    • I have run the integration tests with ./scripts/actions.sh integration-test
    • (bigger changes) I have run a full training and I have written in the comment the run_id(s): launch-slurm.py --time 60
    • (bigger changes and experiments) I have shared a hegdedoc in the github issue with all the configurations and runs for this experiments
  • I have informed and aligned with people impacted by my change:
    • for config changes: the MatterMost channels and/or a design doc
    • for changes of dependencies: the MatterMost software development channel

@github-actions github-actions Bot added the model Related to model training or definition (not generic infra) label Aug 11, 2026
@clessig

clessig commented Aug 12, 2026

Copy link
Copy Markdown
Collaborator

@kctezcan : this seems related to this #2750 ?

@kctezcan

Copy link
Copy Markdown
Contributor Author

@kctezcan : this seems related to this #2750 ?

I think Evens PR chunks the writing process, this PR is related to chunking the calculation of predictions from the latent representations. So these should be orthogonal as far as I see.

@clessig

clessig commented Aug 12, 2026

Copy link
Copy Markdown
Collaborator

@kctezcan : this seems related to this #2750 ?

I think Evens PR chunks the writing process, this PR is related to chunking the calculation of predictions from the latent representations. So these should be orthogonal as far as I see.

Ok. We are working on merging the chunking along the forecast step dimension. This would have to come afterwards.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

model Related to model training or definition (not generic infra)

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

2 participants