Best Image Size for LoRA Training
The best image size for LoRA training is not one universal width and height. It is the resolution your base model and trainer are designed to process, supported by source images that contain real detail at roughly that scale. For many SD 1.x workflows, 512 pixels is a common starting target. SDXL workflows commonly center on about 1024 pixels. FLUX training tools also often use 1024-class buckets, but the exact supported values belong to the trainer configuration you actually run.
Do not force every source into a square. Modern trainers can group images into aspect-ratio buckets, resize each image toward a target pixel area, and crop only the excess needed for a valid batch. That preserves portraits, landscapes, and full-body frames more honestly than stretching them. This guide focuses on the size decision itself: target resolution, source adequacy, buckets, crop risk, file dimensions, and a quick test before a full run.
Quick Answer
Match training resolution to the model family and the detail present in your sources. Use 512-class training for a typical SD 1.5 setup, 1024-class training for a typical SDXL setup, and the documented target for your selected FLUX trainer. Treat those as workflow starting points, not proof that every image should be exactly 512 by 512 or 1024 by 1024.
Keep original aspect ratios whenever bucketing is available. Exclude sources that are too small, badly compressed, or so tightly cropped that resizing cannot preserve the concept. A larger file is useful only when it contains more genuine information.
Resolution Means Training Canvas, Not Required File Shape
When a trainer asks for resolution 1024, it normally describes a target scale used by its resize and bucket logic. It does not necessarily require a directory of square 1024 by 1024 files.
Consider three valid source shapes:
- A 1024 by 1024 head-and-shoulders image.
- An 832 by 1216 full-body portrait.
- A 1344 by 768 environmental scene.
A bucket-aware loader can assign each to a compatible shape. It may resize and crop modestly so dimensions meet model and batching constraints. The important questions are whether the subject remains readable, whether the crop keeps defining features, and whether the resulting pixels carry enough detail.
Stretching all three files to a square changes proportions. Faces widen, products deform, and circles become ovals. Those distortions are then presented as training truth. Padding every file to a square is not automatically safer because repeated borders or fill colors can become a dataset pattern.
Start With the Base Model Family
A LoRA attaches to a base model architecture, so the base model is the first resolution constraint.
SD 1.x
Stable Diffusion 1.x training commonly uses a 512-class target. Higher-resolution training is possible in some tools, but it costs more memory and does not guarantee better learning. If the source set is mostly modest web images, 512 can be more honest than inflating every file to 768 or 1024.
SDXL
SDXL was developed for higher native image sizes and commonly uses 1024-class training with multiple aspect ratios. A 1024 target can retain small costume, material, or facial details that would be reduced at 512. It also raises memory and processing requirements.
FLUX
FLUX LoRA workflows vary by trainer, quantization choice, latent caching, and hardware. Many center training around 1024-class dimensions and support buckets. Use the current documentation and sample configuration for the exact trainer rather than copying an SDXL command blindly.
The model family gives you a baseline. Source quality and available memory determine whether that baseline is practical.
Judge Source Adequacy Before Resizing
Pixel dimensions are only a container. A 2048 by 2048 file can still contain little useful detail if it was enlarged from a thumbnail, blurred by focus, or heavily compressed. Inspect at 100 percent.
Keep a source when:
- The target concept is sharp enough to identify.
- Important edges and textures are real rather than upscaler guesses.
- Compression blocks do not cover defining features.
- The concept occupies enough of the frame to survive resizing.
- The crop includes the features the LoRA needs to learn.
Reject or replace a source when the face is only a few dozen pixels tall, a product logo is unreadable, fabric texture is smeared, or the only available detail comes from artificial sharpening halos.
Upscaling can make dimensions pass a loader check, but it cannot recover evidence that was never captured. Use an upscaler only when you accept that it introduces a processing style. Do not use it as a routine way to turn weak thumbnails into training references.
Use Aspect-Ratio Buckets
Bucketing groups files with similar shapes so a batch can use one tensor size without distorting every image. OneTrainer and kohya-ss sd-scripts both document bucket-oriented dataset controls, although option names and algorithms differ.
Buckets help when a dataset includes portraits, landscapes, close crops, and full-body frames. They reduce destructive cropping and preserve composition variety. They do not eliminate resizing or cropping completely.
Review these trainer settings:
- Target training resolution.
- Whether aspect-ratio bucketing is enabled.
- Minimum and maximum bucket dimensions.
- Bucket step or dimension increment.
- Whether images are allowed to upscale.
- Crop behavior for images between bucket shapes.
Do not copy bucket limits without checking the current trainer. Model architectures often require dimensions divisible by a particular factor, and implementations can enforce different limits.
Control Crop Risk
Automatic bucket assignment may crop an edge after resizing. That is usually harmless for extra sky or background. It is damaging when it removes identity or construction details.
Preview likely crops for:
- The top of a hairstyle, hat, ears, or head accessory.
- Hands, shoes, tails, wings, and other features near the frame edge.
- Product handles, corners, labels, and overall silhouette.
- Garment hems, sleeves, closures, and surface details.
- Signature border treatment in a style dataset.
If a defining feature sits against an edge, make a deliberate crop from the original or choose another image. Do not add a repeated white or black border to solve every tight frame. If padding is required, vary the surrounding content naturally when possible and verify that the padding itself will not become a learned cue.
Balance Detail Against Memory
Higher resolution increases latent size, compute, cache storage, and memory demand. It can also reduce batch size. A run at 1024 with batch size one is not automatically better than a stable 768 or 512 workflow with suitable sources and useful checkpoint comparisons.
Increase resolution when:
- The model family is designed for it.
- Sources contain fine details worth retaining.
- The concept depends on material, typography-like markings, or small construction features.
- Hardware can complete the run reliably.
Reduce resolution when:
- Most sources do not contain real high-resolution detail.
- Memory errors force unstable workarounds.
- Training speed prevents meaningful checkpoint testing.
- The concept is dominated by broad shape and color rather than fine texture.
Resolution is one training parameter, not a quality score.
Avoid Mixed Preprocessing Policies
A dataset can contain mixed aspect ratios, but preprocessing should follow one documented policy. Problems appear when some images are stretched, some are padded, some are aggressively restored, and others remain natural.
Keep a small record of:
- Original dimensions.
- Final dimensions if files are exported in advance.
- Crop or padding applied.
- Upscaling or restoration applied.
- Intended target resolution.
This record makes a failed run diagnosable. It also lets you return to originals instead of repeatedly processing already processed files.
A Practical Size Selection Workflow
- Identify the exact base model architecture and trainer.
- Read the trainer's current resolution and bucket documentation.
- Measure the smallest meaningful feature the LoRA must learn.
- Inspect the lowest-quality sources at full size.
- Choose a model-appropriate target that does not depend on fake detail.
- Enable supported aspect-ratio bucketing.
- Preview representative portrait, landscape, square, and edge-tight images.
- Run a short test with cached previews or early checkpoints.
- Confirm that important features survive and proportions remain natural.
Do this before applying bulk resize operations.
Size Mistakes That Waste Runs
Making every file square
This can stretch anatomy and object geometry or crop away useful composition variety.
Assuming more pixels mean more information
Enlarged blur is still blur. Synthetic detail can become an unwanted texture prior.
Mixing tiny and excellent sources without review
Weak images can lower identity consistency even if the loader accepts them.
Ignoring the concept's size inside the frame
A 4K landscape where the subject is tiny may provide less subject evidence than a clean 768-pixel portrait.
Choosing resolution only from available VRAM
Hardware limits matter, but compatibility with the base model and source detail matters first. Use memory optimizations after choosing a defensible target.
Final Size Checklist
Before training, confirm:
- The target resolution fits the base model and current trainer.
- No image is stretched to fit a square.
- Bucketing is enabled and configured when mixed shapes are present.
- Tight images retain all defining features after the expected crop.
- Small sources contain real usable detail at the chosen scale.
- Upscaled files are exceptions and are marked for review.
- Landscape, portrait, and square examples all load correctly.
- A short test shows natural proportions and readable concept details.
Bottom Line
Pick the smallest model-appropriate resolution that preserves the details your LoRA must learn. Preserve aspect ratio, use buckets, and inspect crop behavior. Source adequacy matters more than the number printed in an image's properties.
Keep originals untouched. If a size choice fails, return to those originals and change one preprocessing decision at a time.
What to Do Next
Confirm the base model and trainer target resolution.
Follow prepare images for lora training for the detailed workflow and checks.
Inspect the smallest and most tightly cropped source images.
Follow lora training dataset guide for the detailed workflow and checks.
Preview bucket resizing before processing the full dataset.
Follow clean a lora training dataset for the detailed workflow and checks.
