SPIN Unprocessed
Source Reddit r/MachineLearning reddit.com Forum
September 14, 2026 ai_technology community

How to automatically find the batch size when using Accelerate with FSDP2? [D]

View original on reddit.com

Overview

Hi, For single-GPU training, I’m using Hugging Face SFTTrainer with auto_find_batch_size=True, which automatically reduces the batch size after a CUDA OOM until it finds a batch size that works. I would like to have similar behavior when training on multiple GPUs on a single node using accelerate launch with FSDP2. Is there a supported way to automatically determine or reduce the batch size when using Accelerate + FSDP2? In particular, I’m wondering how this should be handled when one of the dis

SpinGraph analysis pending — check back after processing.

Ask AI about this story

Opens with the SpinGraph .md URL and structured context — one click, prompt included.

More from Reddit r/MachineLearning

View all →

Markdown (.md) · JSON-LD schema (.json) · Machine-readable for AI & GEO