Quick Start • Architecture • Training • Validation • Inference • Examples • Configuration • Troubleshooting. Use it to ground design choices in named patterns, trade-offs and examples.
SAM3-LoRA: Efficient Fine-Tuning with Low-Rank Adaptation
Snapshot 2026-08-04 16:17:00 UTC · version 1
Research document
SAM3-LoRA: Efficient Fine-Tuning with Low-Rank Adaptation
Quick Start • Architecture • Training • Validation • Inference • Examples • Configuration • Troubleshooting. Use it to ground design choices in named patterns, trade-offs and examples.
Editorial note: curated source snapshot published by Collider.club under the MIT License. Source attribution is preserved in the front matter.
Source snapshot
SAM3-LoRA: Efficient Fine-Tuning with Low-Rank Adaptation
Train SAM3 segmentation models with 99% fewer trainable parameters
Quick Start • Architecture • Training • Validation • Inference • Examples • Configuration • Troubleshooting
Overview
Fine-tune SAM3 (Segment Anything Model 3) on your own dataset using LoRA (Low-Rank Adaptation) — a parameter-efficient method that reduces trainable parameters from 100% to ~1% while maintaining performance. Train on a standard COCO-format dataset with text prompts taken from your category names, while preserving SAM3's open-vocabulary behavior: the fine-tuned model segments your target classes (e.g. crack, hole) and still returns nothing for unrelated prompts (e.g. car).
Recent Updates
2026-08-01 (v0.3.0):
- Guaranteed cross-class hard negatives — negatives are now two-tier: every dataset category absent from an image is always added as a must-return-nothing prompt (previously it was only randomly sampled, far too rare to separate confusable classes like
crackvswater ingress— all prompts ended up segmenting the same defect).num_negativesnow controls only the extra generic out-of-domain prompts sampled per image. Generic negatives that share a word with a dataset category (e.g.watervswater ingress) are removed automatically — training them to return nothing contradicts the related positive class. See Preserving prompt discrimination - Fixed baseline validation scoring ~0 mAP —
validate_sam3_lora.pyscored detections with rawpred_logitsonly, but SAM3's image model keeps the presence score separate (the officialPostProcessImagemultiplies it in). Without it, the zero-shot baseline floods every prompt with confident false detections and its mAP collapses to 0. Validation now uses the joint scoresigmoid(logit) × sigmoid(presence)by default (--no-presencerestores the old behavior). LoRA numbers also change slightly — rerun both sides for a fair comparison - New: pixel-level metrics (
--pixel-metrics, default on) — reports semantic IoU / precision / recall on the union of masks per prompt, alongside mAP/cgF1. Instance matching triple-penalizes a single GT crack predicted as several fragments (missed GT + each fragment a false positive); pixel metrics measure coverage regardless of fragmentation. See Fragmented predictions - New: proximity merging (
--merge-dilate N) —--mergeonly fused overlapping fragments (disjoint crack portions have mask IoU 0 and never merged). With a dilation radius, fragments whose dilated masks touch are joined into one instance before matching - New: box/score display options in evaluation scripts —
compare_lora_base.py/compare_lora_base_batch.pygain--boundingbox True/Falseand--score True/False(per-detection confidence labels, no boxes required);infer_sam.pygains--scoreso confidence shows independently of--boundingbox
2026-07-05 (v0.2.0):
- Fixed loss of prompt discrimination after fine-tuning — previously the model detected trained objects regardless of the text prompt (e.g. segmenting cracks when prompted
car). Root cause: positive-only training queries. See Preserving prompt discrimination - Automatic hard-negative text prompts — new
num_negativesconfig option; the trainers now add per-image prompts that must return zero detections, drawn from your other dataset categories plus a built-in generic pool. Newgeneric_negativesconfig option overrides the pool when a default concept (e.g.road) can genuinely appear in your images. Validation loss includes the negative prompts, so it tracks discrimination during training - Text encoder frozen by default (
apply_to_text_encoder: false) to preserve SAM3's text↔image alignment - Fixed
apply_to_mask_decoderhaving no effect (#28) — SAM3's mask module is namedsegmentation_head, which the component filter didn't match, so the flag was silently ignored (and the head's cross-attention was always adapted regardless of the setting). The flag now correctly controls LoRA on the segmentation head - New: merge LoRA into the base model (#30) —
merge_lora_weights.pyfolds a trained adapter into the base weights, producing a single checkpoint that loads into stock SAM3 with no LoRA code. See Merging LoRA weights - New: RefCOCO / RRSIS-D support (#26) —
convert_refcoco_to_coco.pyconverts referring-expression datasets to the trainers' COCO layout, with expression-level (--mode ref) or class-level (--mode category) prompts. See Prepare Your Data
2026-02-03:
- Fixed multi-class category assignment bug in training/validation
- Previously, images with multiple categories incorrectly assigned all objects to the mode (most frequent) category
- Now creates separate queries per category, mapping each object to its actual class
- Affected files:
train_sam3_lora_native.py,train_sam3_lora_with_categories.py,validate_sam3_lora.py
2026-01-31:
- Replaced
--no-boxeswith--boundingboxoption ininfer_sam.py - New
--boundingbox True/Falseflag for explicit bounding box control (default: False) - Updated README documentation and inference examples
2026-01-04:
- Added Multi-GPU training support using DistributedDataParallel (DDP)
- New
--deviceargument for easy GPU selection:--device 0 1 2 3 - Automatic torchrun launch when multiple GPUs specified
- Linear scaling of effective batch size across GPUs
Why Use This?
- ✅ Train on Consumer GPUs: 16GB VRAM instead of 80GB
- ✅ Tiny Checkpoints: 10-50MB LoRA weights vs 3GB full model
- ✅ Fast Iterations: Less memory = faster training
- ✅ Easy to Use: standard COCO dataset + YAML configs + simple CLI
- ✅ Keeps Prompt Discrimination: automatic hard-negative prompts, so the tuned model doesn't fire on unrelated words
- ✅ Production Ready: Complete train + inference pipeline
- ✅ Real Applications: Crack detection, defect inspection, and more
- ✅ Multi-GPU Support: Scale training across multiple GPUs with
--device 0 1 2 3
What is LoRA?
Instead of fine-tuning all model weights, LoRA injects small trainable matrices:
W' = W_frozen + B×A (where rank << model_dim)
Result: Only ~1% of parameters need training!
Architecture
SAM3-LoRA applies Low-Rank Adaptation to key components of the SAM3 architecture:
SAM3 Model Architecture with Full LoRA Adaptation
LoRA Adapters Applied To:
| Component | Description | Default | Notes |
|---|---|---|---|
| Vision Encoder (ViT) | Extracts visual features from input images | LoRA ✅ | High impact - primary feature learning |
| Text Encoder | Processes text prompts for guided segmentation | Frozen ❄️ | Keep frozen — adapting it erodes prompt discrimination |
| Geometry Encoder | Handles geometric prompts (boxes, points) | Frozen | Enable if using box/point prompts |
| DETR Encoder | Transformer encoder for object detection | LoRA ✅ | High impact - vision-text fusion |
| DETR Decoder | Transformer decoder for object queries | LoRA ✅ | High impact - object localization |
| Mask Decoder | Generates segmentation masks | Frozen | Enable for fine-grained mask quality (on in light_lora_config) |
Data Flow:
- Input: Image + Text/Geometric prompts
- Encoding: Multiple encoders process different modalities
- Transformation: DETR encoder-decoder refines representations
- Output: High-quality segmentation masks
LoRA Benefits:
- ✅ Only ~1% parameters trainable (frozen base + small adapters)
- ✅ Adapters can be swapped for different tasks
- ✅ Original model weights preserved
- ✅ Efficient storage (10-50MB vs 3GB full model)
Installation
Prerequisites
Before installing, you need to:
Request SAM3 Access on Hugging Face
- Go to facebook/sam3 on Hugging Face
- Click "Request Access" and accept the license terms
- Wait for approval (usually instant to a few hours)
Get Your Hugging Face Token
- Go to Hugging Face Settings > Tokens
- Create a new token or use existing one
- Copy the token (you'll need it in the next step)
Install
# Clone repository
git clone https://github.com/Sompote/SAM3_LoRA.git
cd SAM3_LoRA
# Install dependencies
pip install -e .
# Login to Hugging Face
hf auth login
# Paste your token when prompted
Alternative login method:
# Or set token as environment variable
export HF_TOKEN="your_token_here"
Requirements: Python 3.8+, PyTorch 2.0+, CUDA (optional), Hugging Face account with SAM3 access
Verification
Verify your setup is complete:
# Test Hugging Face login
huggingface-cli whoami
# Test SAM3 access (should not give access error)
python3 -c "from transformers import AutoModel; print('✓ SAM3 accessible')"
If you see errors, review the Troubleshooting section.
Quick Start
⚠️ Important: Make sure you've completed the Installation steps, including Hugging Face login, before proceeding.
Example Result: Train a model to detect concrete cracks with just ~1% trainable parameters!
Detection: "concrete crack" with 0.32 confidence • Precise segmentation mask
1. Prepare Your Data
Organize your dataset in COCO format with a single annotation file per split:
data/
├── train/ # Required
│ ├── img001.jpg
│ ├── img002.jpg
│ └── _annotations.coco.json
├── valid/ # Optional but recommended
│ ├── img001.jpg
│ ├── img002.jpg
│ └── _annotations.coco.json
└── test/ # Optional
├── img001.jpg
└── _annotations.coco.json
Note: Validation data (
data/valid/) is optional but strongly recommended for monitoring training progress and preventing overfitting.
COCO Annotation Format (_annotations.coco.json):
{
"images": [
{
"id": 0,
"file_name": "img001.jpg",
"height": 480,
"width": 640
}
],
"annotations": [
{
"id": 1,
"image_id": 0,
"category_id": 1,
"bbox": [x, y, width, height],
"area": 1234,
"segmentation": [[x1, y1, x2, y2, ...]],
"iscrowd": 0
}
],
"categories": [
{"id": 1, "name": "defect"}
]
}
Supported Segmentation Formats:
- Polygon:
"segmentation": [[x1, y1, x2, y2, ...]](list of polygons) - RLE:
"segmentation": {"counts": "...", "size": [h, w]}(run-length encoded)
Using RefCOCO-format datasets (RefCOCO, RRSIS-D, ...): referring-expression
datasets ship as instances.json + a refs(*).p pickle instead of per-split
COCO files. Convert them with:
python3 convert_refcoco_to_coco.py \
--refcoco-dir RRSIS-D/rrsisd \
--images-dir RRSIS-D/images/rrsisd/JPEGImages \
--output-dir data/rrsisd_coco \
--mode ref
--mode ref (default) turns each referring expression into a text prompt for
exactly the one instance it refers to — true referring segmentation; keep
num_negatives enabled so other expressions serve as hard negatives.
--mode category ignores the expressions and produces plain class-level
detection data from the original COCO categories. Use --symlink to link
images instead of copying.
2. Train Your Model
# Train with default config
python3 train_sam3_lora_native.py
# Or specify custom config
python3 train_sam3_lora_native.py --config configs/full_lora_config.yaml
Expected output:
Building SAM3 model...
Applying LoRA...
Applied LoRA to 64 modules
Trainable params: 11,796,480 (1.38%)
Loading training data from /workspace/data2...
Loaded COCO dataset: train split
Images: 778
Annotations: 1631
Categories: {0: 'CRACKS', 1: 'CRACKS', 2: 'JOINT', 3: 'LOCATION', 4: 'MARKING'}
Loading validation data from /workspace/data2...
Loaded COCO dataset: valid split
Images: 152
Annotations: 298
Found validation data: 152 images
Starting training for 100 epochs...
Training samples: 778, Validation samples: 152
Epoch 1: 100%|████████| 98/98 [07:47<00:00, loss=140]
Validation: 100%|████████| 19/19 [00:32<00:00, val_loss=23.7]
Epoch 1/100 - Train Loss: 156.234567, Val Loss: 17.032280
✓ New best model saved (val_loss: 17.032280)
Epoch 2: 100%|████████| 98/98 [07:24<00:00, loss=167]
Validation: 100%|████████| 19/19 [00:31<00:00, val_loss=20.1]
Epoch 2/100 - Train Loss: 142.891234, Val Loss: 15.641912
✓ New best model saved (val_loss: 15.641912)
...
Validation Strategy (mirrors how SAM3's own training/eval code is organized — see note below):
- During training: Only validation loss is computed (fast, no NMS or metrics)
- After training: Run
validate_sam3_lora.pyfor full metrics (mAP, cgF1) with NMS
This approach significantly speeds up training while still monitoring overfitting via validation loss.
Note on "mirrors SAM3": this is not a documented policy quoted from the SAM3 repo — it's a design choice in this repo that follows how SAM3's code is structured. Two concrete sources motivate it: (1) SAM3's trainer runs a periodic validation pass that accumulates loss/meters on a configurable cadence (
sam3/train/trainer.py, e.g.val_epoch_freq/Phase.VAL); and (2) full COCO mAP / cgF1 live in separate offline evaluators (sam3/eval/coco_eval_offline.py,sam3/eval/cgf1_eval.py) that are run as a distinct step — the COCO evaluator's own docstring notes category mAP requires predicting over every(image, class)pair, which is why it's kept out of the training loop. If you have an official SAM3 statement on eval cadence, a pointer is welcome so we can cite it directly.
3. Run Inference
# Basic inference (automatically uses best model)
python3 infer_sam.py \
--config configs/full_lora_config.yaml \
--image test_image.jpg \
--output predictions.png
# With text prompt for better accuracy
python3 infer_sam.py \
--config configs/full_lora_config.yaml \
--image test_image.jpg \
--prompt "yellow school bus" \
--output predictions.png
# Multiple prompts to detect different objects
python3 infer_sam.py \
--config configs/full_lora_config.yaml \
--image test_image.jpg \
--prompt "crack" "defect" "damage" \
--output predictions.png
Training
Basic Training
# Use default configuration (single GPU)
python3 train_sam3_lora_native.py
# Or specify custom config
python3 train_sam3_lora_native.py --config configs/full_lora_config.yaml
Preserving prompt discrimination (important)
A common issue after fine-tuning is that the model detects the trained objects
regardless of the text prompt — e.g. after training on crack/hole it still
segments cracks when you prompt "car". This happens because SAM3 is an
open-vocabulary detector, but the dataset only ever provides positive
prompts (every prompt has matching boxes). The presence head then decouples from
the text condition and learns to "always fire."
Two settings fix this:
Hard-negative text queries — prompts that must return nothing for an image. The dataloader adds them in two tiers:
- In-domain negatives (always added). Every dataset category absent from
an image becomes a zero-detection query on that image. These are the
confusable prompts — on a water-ingress image,
crackandconcrete spallingare told to return nothing, every single epoch. When these were only randomly sampled, the signal was too rare for the model to keep visually similar classes apart: after fine-tuning, all three prompts segmented the same defect. - Generic out-of-domain negatives (sampled).
num_negativesprompts per image drawn from a built-in pool (car,person,dog, …).2–4works well;0disables this tier only. The shipped configs set3.
training: num_negatives: 3 # generic tier only; absent dataset categories are always addedIf a concept from the default generic pool can genuinely appear in your images (e.g.
roadin pavement-crack photos — teaching "no road here" on a picture of a road is wrong supervision), replace the pool with objects that never appear in your data:training: generic_negatives: ["car", "person", "dog", "bicycle", "bottle"]Pool entries that share a word with one of your dataset categories are removed automatically (e.g.
wateris dropped when a category is namedwater ingress) — trainingwaterto return nothing on images full of water ingress would contradict the positive class.- In-domain negatives (always added). Every dataset category absent from
an image becomes a zero-detection query on that image. These are the
confusable prompts — on a water-ingress image,
Keep the text encoder frozen (
apply_to_text_encoder: false). Adapting the text encoder distorts SAM3's text↔image alignment — the very thing that separatescrackfromcar. Leave it frozen unless you have a strong reason not to.lora: apply_to_text_encoder: false
Also avoid over-training on small datasets (hundreds of epochs amplifies the collapse); prefer a lower learning rate and fewer epochs, and watch the validation loss, which now includes the negative prompts.
Do negatives generalize to prompts outside the pool?
Yes — the negative pool is not a blocklist. If you train with
generic_negatives: ["car", "person", "dog", "bicycle", "bottle"] and later
prompt plane, the model should still return nothing, even though plane was
never a training negative.
The negatives don't teach the model "don't fire on the word car" — they retrain
the presence head's general decision rule: fire only when the text embedding
actually matches the visual features. Each negative is one demonstration of
that rule; the model relearns the rule, not the word list. The frozen text
encoder is what makes this work: SAM3's pretrained encoder already places
plane near car/bicycle and far from crack, and freezing it
(apply_to_text_encoder: false) keeps that concept space intact — the
negatives just re-anchor the presence head to use it again. A more diverse pool
generalizes further, which is why the built-in one spans 20 varied concepts.
The caveat is near-synonyms of your target classes. After training on
crack/hole/deep crack, prompts like scratch, fracture, gap, or
line sit close to crack in the text encoder's space and may still trigger
detections. That's the model honestly saying "this concept looks like what I
was trained to find" — whether that's desired depends on your use case. If a
specific nearby word must return nothing, add exactly that word to
generic_negatives.
After retraining, spot-check three rings of prompts:
- Trained classes (
crack,hole) — should segment correctly. - Clearly unrelated words, including ones NOT in your pool (
plane,banana,keyboard) — should return nothing. This verifies generalization rather than memorization. - Near-synonyms (
scratch,gap) — check the behavior matches what you want; add problem words to the pool if not.
If ring 2 still fires, increase num_negatives, diversify the pool, or reduce
over-training (fewer epochs / lower LR) — don't try to enumerate every English
word.
Multi-GPU Training
Train on multiple GPUs using the --device argument. The script automatically handles distributed training setup.
# Single GPU (default - GPU 0)
python3 train_sam3_lora_native.py --config configs/full_lora_config.yaml
# Single GPU (specific GPU)
python3 train_sam3_lora_native.py --config configs/full_lora_config.yaml --device 1
# Multi-GPU (2 GPUs)
python3 train_sam3_lora_native.py --config configs/full_lora_config.yaml --device 0 1
# Multi-GPU (4 GPUs)
python3 train_sam3_lora_native.py --config configs/full_lora_config.yaml --device 0 1 2 3
# Multi-GPU (specific GPUs, e.g., 0, 2, 3)
python3 train_sam3_lora_native.py --config configs/full_lora_config.yaml --device 0 2 3
Multi-GPU Features:
- ✅ Automatic
torchrunlaunch when multiple GPUs specified - ✅ DistributedDataParallel (DDP) for efficient gradient synchronization
- ✅ DistributedSampler for proper data sharding
- ✅ Synchronized validation loss across all GPUs
- ✅ Model saving only on rank 0 (no file conflicts)
Effective Batch Size: With multi-GPU, your effective batch size scales linearly:
effective_batch_size = batch_size × num_gpus
| Config batch_size | GPUs | Effective Batch Size |
|---|---|---|
| 4 | 1 | 4 |
| 4 | 2 | 8 |
| 4 | 4 | 16 |
Expected Output (Multi-GPU):
Launching distributed training on GPUs: [0, 1]
Number of processes: 2
Multi-GPU training enabled with 2 GPUs
Building SAM3 model...
Applying LoRA...
Trainable params: 11,796,480 (1.38%)
Model wrapped with DistributedDataParallel
Effective batch size: 4 x 2 = 8
Starting training for 100 epochs...
Custom Configuration
Create a config file (e.g., configs/my_config.yaml):
lora:
rank: 16 # LoRA rank (higher = more capacity)
alpha: 32 # Scaling factor (typically 2×rank)
dropout: 0.1 # Dropout for regularization
target_modules: # Which layers to adapt
- "q_proj" # Query projection
- "k_proj" # Key projection
- "v_proj" # Value projection
- "fc1" # MLP layer 1
- "fc2" # MLP layer 2
# Which model components to apply LoRA to
apply_to_vision_encoder: true
apply_to_mask_decoder: true
apply_to_detr_encoder: false
apply_to_detr_decoder: false
training:
data_dir: "/path/to/data" # Root directory with train/valid/test folders
batch_size: 8 # Adjust based on GPU memory
num_epochs: 100 # Training epochs
learning_rate: 5e-5 # Learning rate (5e-5 recommended for SAM3 fine-tuning)
weight_decay: 0.01 # Weight decay
gradient_accumulation_steps: 8 # Effective batch = batch_size × accumulation
output:
output_dir: "outputs/my_model"
Important Notes:
- Category-aware prompts: The training automatically uses category names as text prompts (e.g., "crack", "joint") extracted from COCO annotations
- Each training image is prompted with its specific object categories (in lowercase)
- This approach improves performance by using task-specific vocabulary while leveraging SAM3's pre-trained text understanding
Then train:
python3 train_sam3_lora_native.py --config configs/my_config.yaml
Model Checkpointing
During training, two models are automatically saved:
best_lora_weights.pt: Best model based on validation loss (saved only when validation loss improves)last_lora_weights.pt: Model from the last epoch (saved after every epoch)
With validation data: Training monitors validation loss only (fast). Best model is saved when validation loss decreases.
Without validation data: Training continues normally but saves the last epoch as both files. You'll see:
⚠️ No validation data found - training without validation
...
ℹ️ No validation data - consider adding data/valid/ for better model selection
Merging LoRA weights into the base model
If you don't want to load the LoRA adapter separately at inference time, you can fold it into the base weights and get a single checkpoint that loads into the original SAM3 architecture — no LoRA code or config needed:
python3 merge_lora_weights.py \
--config configs/full_lora_config.yaml \
--lora-weights outputs/sam3_lora_full/best_lora_weights.pt \
--output sam3_merged.pt
--config must be the same config used for training (the adapter file
stores only the low-rank matrices, so the model is rebuilt with the same LoRA
structure before merging). --lora-weights defaults to
<output_dir>/best_lora_weights.pt from the config. Merging runs on CPU — no
GPU needed.
Load the result with stock SAM3 code:
from sam3.model_builder import build_sam3_image_model
model = build_sam3_image_model(load_from_HF=True, eval_mode=True,
bpe_path="sam3/assets/bpe_simple_vocab_16e6.txt.gz")
model.load_state_dict(torch.load("sam3_merged.pt", map_location="cpu"))
The merged model computes W' = W + B^T A^T · (alpha/rank) for every adapted
layer, so its predictions are numerically identical to base + adapter.
Trade-offs: the checkpoint is full model size (~3 GB) instead of a 10–50 MB
adapter, and the adaptation is baked in — you can no longer swap adapters on
one shared base model.
Validation
Overview
SAM3-LoRA uses a two-stage validation approach following SAM3's original design:
- During Training: Only validation loss is computed (fast, no expensive metrics)
- After Training: Run full evaluation with mAP, cgF1 metrics and NMS filtering
This approach significantly speeds up training while still monitoring overfitting via validation loss.
Quick Validation
After training completes, evaluate your model:
# Validate LoRA-adapted model (uses best model automatically)
python3 validate_sam3_lora.py \
--config configs/full_lora_config.yaml \
--weights outputs/sam3_lora_full/best_lora_weights.pt \
--val_data_dir /workspace/data2/valid
# Evaluate on test set
python3 validate_sam3_lora.py \
--config configs/full_lora_config.yaml \
--weights outputs/sam3_lora_full/best_lora_weights.pt \
--val_data_dir /workspace/data2/test
# Baseline: Validate with original SAM3 model (no LoRA) for comparison
python3 validate_sam3_lora.py \
--val_data_dir /workspace/data2/valid \
--use-base-model
Presence scoring (v0.3.0): detections are ranked by the joint score
sigmoid(pred_logit) × sigmoid(presence_logit), matching SAM3's officialPostProcessImage. This matters most for the--use-base-modelbaseline: the pretrained model's raw query logits are confident even for absent concepts, and scoring them alone floods the eval with false positives (baseline mAP reads as 0.0000). If your baseline scored ~0 with an older version, rerun — and rerun the LoRA side too, since the scoring change applies to both.--no-presencerestores the old raw-logit behavior.
Expected Output:
Running SAM3 LoRA Validation
Building SAM3 model...
Loading LoRA weights from outputs/sam3_lora_full/best_lora_weights.pt
Loaded COCO dataset: valid split
Images: 152
Annotations: 298
Processing: 100%|████████| 152/152 [02:15<00:00]
Validation Results:
================================================================================
Total predictions: 946 (after NMS from 1353 initial detections)
Total ground truth: 298
COCO Evaluation Metrics:
--------------------------------------------------------------------------------
mAP (IoU 0.50:0.95): 0.245
mAP@50 (IoU 0.50): 0.287
mAP@75 (IoU 0.75): 0.198
Category-agnostic F1 Scores:
--------------------------------------------------------------------------------
cgF1 (avg): 0.135
cgF1@50: 0.149
cgF1@75: 0.089
--------------------------------------------------------------------------------
Pixel-level metrics (union of masks per prompt; robust to
a single crack being predicted as multiple fragments):
Pixel IoU: 0.412
Pixel Precision: 0.573
Pixel Recall: 0.598
Mean per-prompt IoU: 0.387
================================================================================
Multi-class validation
The validator fully supports datasets with multiple categories per image. SAM3
is prompt-based: for each image it issues one text prompt per category, and the
model returns a separate set of predictions for every prompt. The validator
treats each (image, prompt) pair as its own evaluation unit and scores each
prompt only against the ground-truth objects of that category — so a crack
prediction is never matched against hole boxes, and prompts that should find
nothing are evaluated as empty units.
No extra flags are needed — multi-class COCO files (multiple entries under
categories) are handled automatically. The category names in your
_annotations.coco.json are used verbatim as the text prompts, so make sure
they read like natural concepts (e.g. "deep crack", not "class_2").
Note: earlier versions assumed a single prompt per image and would raise an index-out-of-range error in
create_coco_gt_from_dataseton multi-category datasets (the prediction batch dimension is the number of prompts, not images). This is fixed —git pullif you hit that error.
Validation Metrics Explained
| Metric | Description | Good Value | Excellent Value |
|---|---|---|---|
| mAP (0.50:0.95) | Mean Average Precision across IoU thresholds 0.5 to 0.95 | > 0.30 | > 0.50 |
| mAP@50 | Precision at IoU threshold 0.50 (looser) | > 0.40 | > 0.70 |
| mAP@75 | Precision at IoU threshold 0.75 (stricter) | > 0.25 | > 0.45 |
| cgF1 | Concept-level F1 (SAM3's primary metric) | > 0.25 | > 0.50 |
| cgF1@50 | cgF1 at IoU 0.50 | > 0.30 | > 0.60 |
| cgF1@75 | cgF1 at IoU 0.75 | > 0.15 | > 0.35 |
| Pixel IoU | Semantic IoU on the union of masks per prompt | > 0.40 | > 0.60 |
Understanding the Metrics:
- mAP: Standard COCO metric - higher is better, penalizes over/under-segmentation
- cgF1: SAM3's concept-level metric - balances precision and recall for concepts, not individual instances
- @50/@75: Different IoU thresholds (50% overlap vs 75% overlap)
- Pixel IoU / Precision / Recall: instance-free coverage metrics — see the next section
Fragmented predictions and pixel-level metrics (cracks)
Instance matching (mAP, cgF1) requires one prediction to cover one GT instance at IoU ≥ threshold. That triple-penalizes a common crack failure mode: the model predicts one GT crack as two disjoint portions. Each portion's IoU against the full GT is at best ~0.5 (intersection is only its own pixels, union is the whole crack), so typically the GT counts as missed and both portions count as false positives — even though together they cover the crack well. This is a large part of why mAP@75 is much lower than mAP@50 for thin structures.
Two tools address this:
Pixel-level metrics (
--pixel-metrics, on by default) — per (image, prompt) unit, all predicted masks are unioned and compared pixel-by-pixel against the union of GT masks. Fragmentation doesn't matter, only coverage does; this is the standard metric in the crack-segmentation literature. A large gap between Pixel IoU and mAP@50 tells you the model finds the cracks but splits them.Proximity merging (
--merge --merge-dilate N) — plain--mergeonly fuses overlapping fragments (disjoint portions have mask IoU 0 with each other and are never merged). With--merge-dilate N, masks are dilated by N pixels and fragments whose dilated masks touch are joined into a single instance (the output keeps the original undilated union), which then matches the GT as one prediction. N is in pixels at the 288×288 mask resolution;3–5works well for cracks (N=4 bridges gaps up to ~8 px at 288, roughly 28 px at 1008).
python3 validate_sam3_lora.py \
--config configs/full_lora_config.yaml \
--weights outputs/sam3_lora_full/best_lora_weights.pt \
--val_data_dir /workspace/data2/valid \
--merge --merge-dilate 4 \
--pixel-metrics True
Advanced Validation Options
1. Adjust Confidence Threshold:
# More conservative (fewer but higher confidence predictions)
python3 validate_sam3_lora.py \
--config configs/full_lora_config.yaml \
--weights outputs/sam3_lora_full/best_lora_weights.pt \
--val_data_dir /workspace/data2/valid \
--prob-threshold 0.5
# More permissive (more predictions, lower confidence)
python3 validate_sam3_lora.py \
--config configs/full_lora_config.yaml \
--weights outputs/sam3_lora_full/best_lora_weights.pt \
--val_data_dir /workspace/data2/valid \
--prob-threshold 0.2
2. Merge Overlapping Segments (for crack-like objects):
# Enable merging to reduce over-segmentation
python3 validate_sam3_lora.py \
--config configs/full_lora_config.yaml \
--weights outputs/sam3_lora_full/best_lora_weights.pt \
--val_data_dir /workspace/data2/valid \
--merge \
--merge-iou 0.15
# Aggressive merging for highly fragmented objects
python3 validate_sam3_lora.py \
--config configs/full_lora_config.yaml \
--weights outputs/sam3_lora_full/best_lora_weights.pt \
--val_data_dir /workspace/data2/valid \
--merge \
--merge-iou 0.05 \
--prob-threshold 0.5
3. Adjust NMS Settings:
# More aggressive NMS (fewer duplicate detections)
python3 validate_sam3_lora.py \
--config configs/full_lora_config.yaml \
--weights outputs/sam3_lora_full/best_lora_weights.pt \
--val_data_dir /workspace/data2/valid \
--nms-iou 0.5
# Less aggressive NMS (keep more overlapping segments)
python3 validate_sam3_lora.py \
--config configs/full_lora_config.yaml \
--weights outputs/sam3_lora_full/best_lora_weights.pt \
--val_data_dir /workspace/data2/valid \
--nms-iou 0.8
4. Baseline Comparison (Original SAM3 Model):
# Validate with original SAM3 model (no LoRA) for comparison
python3 validate_sam3_lora.py \
--val_data_dir /workspace/data2/valid \
--use-base-model
# This helps you understand the improvement from LoRA fine-tuning
# Compare against your LoRA model results to see performance gains
Validation Parameters Reference
| Parameter | Default | Description | When to Adjust |
|---|---|---|---|
--prob-threshold |
0.3 | Minimum confidence score | Lower if missing objects (0.2), higher if too many false positives (0.5) |
--nms-iou |
0.7 | NMS IoU threshold | Lower for fewer duplicates (0.5), higher to keep overlaps (0.8) |
--merge |
False | Enable segment merging | Use for crack-like or connected objects |
--merge-iou |
0.15 | IoU threshold for merging | Lower for aggressive merging (0.05), higher for conservative (0.25) |
--merge-dilate |
0 | Dilation radius (px at 288) to join disjoint fragments | 3-5 for cracks split into portions; requires --merge |
--pixel-metrics |
True | Report pixel-level IoU/precision/recall | Disable with False if you only want instance metrics |
--no-presence |
False | Score with raw logits (skip presence score) | Only to reproduce pre-v0.3.0 numbers |
--use-base-model |
False | Use original SAM3 (no LoRA) | For baseline comparison |
Interpreting Results
Scenario 1: Too Many Predictions
Total predictions: 1353
Total ground truth: 298
mAP@50: 0.29
Solution: Model is over-segmenting. Try:
- Increase
--prob-thresholdto 0.4-0.5 - Decrease
--nms-iouto 0.5-0.6 - Use
--mergewith--merge-iou 0.15
Scenario 2: Too Few Predictions
Total predictions: 150
Total ground truth: 298
mAP@50: 0.15
Solution: Model is under-detecting. Try:
- Decrease
--prob-thresholdto 0.2 - Train longer or with higher LoRA rank
Scenario 3: Good Quantity, Poor Quality
Total predictions: 310
Total ground truth: 298
mAP@50: 0.35 (low)
cgF1@50: 0.25 (low)
Solution: Detections are inaccurate. Need better training:
- Train for more epochs
- Use
configs/full_lora_config.yamlinstead of light config - Check data quality
Why Separate Evaluation?
Benefits:
- ⚡ 10x Faster Training: No expensive metric computation during training
- 📊 Better Monitoring: Validation loss is sufficient to detect overfitting
- 🎯 Accurate Metrics: Full evaluation with proper NMS and post-processing
- 🔧 Flexible Testing: Try different thresholds without retraining
Training Tips
Starting Out:
- Use
rank: 4orrank: 8for quick experiments - Set
num_epochs: 5for initial tests - Monitor that trainable params are ~0.5-2%
- Watch validation loss - it should decrease over epochs
Production Training:
- Increase to
rank: 16orrank: 32for better performance - Use
num_epochs: 20-50depending on dataset size - Enable more components (DETR encoder/decoder) if needed
- Use early stopping if validation loss stops improving
Troubleshooting:
- Loss too low (< 0.001): Model might be overfitting, reduce rank or add regularization
- Val loss > Train loss: Normal, indicates some overfitting
- Val loss increasing: Overfitting! Reduce rank, add dropout, or stop training
- Loss not decreasing: Increase learning rate or rank
- OOM errors: Reduce batch size or rank
- 63% trainable params: Bug! Should be ~1% - make sure base model is frozen
Inference
Run inference on new images using your trained LoRA model. The infer_sam.py script is based on official SAM3 patterns and supports multiple text prompts and NMS filtering for clean, non-overlapping detections.
Command Line
# Basic inference (automatically uses best model)
python3 infer_sam.py \
--config configs/full_lora_config.yaml \
--image path/to/image.jpg \
--output predictions.png
# With text prompt (recommended for better accuracy)
python3 infer_sam.py \
--config configs/full_lora_config.yaml \
--image path/to/image.jpg \
--prompt "yellow school bus" \
--output predictions.png
# Multiple prompts to detect different object types
python3 infer_sam.py \
--config configs/full_lora_config.yaml \
--image street_scene.jpg \
--prompt "car" "person" "bus" \
--output segmentation.png
# Use last epoch model instead
python3 infer_sam.py \
--config configs/full_lora_config.yaml \
--weights outputs/sam3_lora_full/last_lora_weights.pt \
--image path/to/image.jpg \
--prompt "person with red backpack" \
--output predictions.png
# With custom confidence threshold
python3 infer_sam.py \
--config configs/full_lora_config.yaml \
--image path/to/image.jpg \
--prompt "building" \
--threshold 0.3 \
--output predictions.png
# Adjust NMS to reduce overlapping boxes
python3 infer_sam.py \
--config configs/full_lora_config.yaml \
--image path/to/image.jpg \
--prompt "seal" \
--threshold 0.3 \
--nms-iou 0.3 \
--output clean_detections.png
# With bounding boxes
python3 infer_sam.py \
--config configs/full_lora_config.yaml \
--image path/to/image.jpg \
--prompt "crack" \
--boundingbox True \
--output with_boxes.png
NMS (Non-Maximum Suppression)
NMS removes overlapping bounding boxes to produce clean visualizations. Without NMS, you may see a grid-like pattern of many overlapping boxes.
# Default NMS IoU = 0.5 (good for most cases)
python3 infer_sam.py --config configs/full_lora_config.yaml --image test.jpg --prompt "object"
# More aggressive NMS (fewer boxes, less overlap)
python3 infer_sam.py --config configs/full_lora_config.yaml --image test.jpg --prompt "object" --nms-iou 0.3
# Less aggressive NMS (keep more overlapping detections)
python3 infer_sam.py --config configs/full_lora_config.yaml --image test.jpg --prompt "object" --nms-iou 0.7
NMS IoU Guidelines:
| Value | Effect | Use Case |
|---|---|---|
| 0.3 | Aggressive filtering | Single object per region, clean output |
| 0.5 | Balanced (default) | Most general use cases |
| 0.7 | Keep more boxes | Densely packed objects, overlapping instances |
Text Prompts
Text prompts help guide the model to segment specific objects more accurately. New feature: You can now use multiple prompts in a single command!
Single prompt examples:
"yellow school bus"- Specific color and object type"person wearing red hat"- Object with distinctive features"car"- Simple, clear object type"crack"- For defect detection"building with glass windows"- Object with distinguishing features
Multiple prompt examples:
# Detect different defect types
--prompt "crack" "spalling" "corrosion"
# Detect multiple objects in street scenes
--prompt "car" "person" "traffic sign"
Tips for better prompts:
- Be specific but concise
- Include distinctive colors or features when relevant
- Use natural language descriptions
- For multiple prompts, order from most to least important
- Match the vocabulary to your training data
Inference Parameters
| Parameter | Description | Example | Default |
|---|---|---|---|
--config |
Path to training config file | configs/full_lora_config.yaml |
Required |
--weights |
Path to LoRA weights (optional) | outputs/sam3_lora_full/best_lora_weights.pt |
Auto-detected |
--image |
Input image path | test_image.jpg |
Required |
--prompt |
One or more text prompts | "crack" or "crack" "defect" |
"object" |
--output |
Output visualization path | predictions.png |
output.png |
--threshold |
Confidence threshold (0.0-1.0) | 0.3 |
0.5 |
--nms-iou |
NMS IoU threshold (lower = fewer boxes) | 0.3 |
0.5 |
--resolution |
Input resolution | 1008 |
1008 |
--boundingbox |
Show bounding boxes (True/False) | True |
False |
--no-masks |
Don't show segmentation masks | - | False |
Python API
from infer_sam import SAM3LoRAInference
# Initialize inference engine with NMS
inferencer = SAM3LoRAInference(
config_path="configs/full_lora_config.yaml",
weights_path="outputs/sam3_lora_full/best_lora_weights.pt",
detection_threshold=0.5,
nms_iou_threshold=0.5 # Adjust for cleaner output (lower = fewer boxes)
)
# Run prediction with single text prompt
predictions = inferencer.predict(
image_path="image.jpg",
text_prompts=["yellow school bus"]
)
# Run prediction with multiple text prompts
predictions = inferencer.predict(
image_path="image.jpg",
text_prompts=["crack", "defect", "damage"]
)
# Visualize results
inferencer.visualize(
predictions,
output_path="output.png",
show_boxes=True,
show_masks=True
)
# Access predictions for each prompt (NMS already applied)
for idx, prompt in enumerate(["crack", "defect"]):
result = predictions[idx]
print(f"Prompt '{result['prompt']}':")
print(f" Detections: {result['num_detections']}")
if result['num_detections'] > 0:
print(f" Boxes: {result['boxes'].shape}") # [N, 4] in xyxy format
print(f" Scores: {result['scores'].shape}") # [N]
print(f" Masks: {result['masks'].shape}") # [N, H, W]
Configuration
LoRA Parameters
| Parameter | Description | Typical Values |
|---|---|---|
rank |
LoRA rank (bottleneck dimension) | 4, 8, 16, 32 |
alpha |
Scaling factor | 2×rank (e.g., 16 for rank=8) |
dropout |
Dropout probability | 0.0 - 0.1 |
target_modules |
Which layer types to adapt | q_proj, k_proj, v_proj, fc1, fc2 |
Component Flags
| Flag | Description | When to Enable |
|---|---|---|
apply_to_vision_encoder |
Vision backbone | Always (main feature extractor) |
apply_to_detr_encoder |
Object detection encoder (vision-text fusion) | Recommended |
apply_to_detr_decoder |
Object detection decoder (object queries) | Recommended |
apply_to_mask_decoder |
Segmentation head (mask generation) | Optional — adapts the head's prompt cross-attention |
apply_to_geometry_encoder |
Box/point prompt encoder | Only if training with geometric prompts |
apply_to_text_encoder |
Text understanding | Keep false — adapting it erodes prompt discrimination (see Preserving prompt discrimination) |
Preset Configurations
Minimal (Fastest, Lowest Memory)
lora:
rank: 4
alpha: 8
target_modules: ["q_proj", "v_proj"]
apply_to_vision_encoder: true
# All others: false
Balanced (Recommended)
lora:
rank: 16
alpha: 32
target_modules: ["q_proj", "k_proj", "v_proj", "fc1", "fc2"]
apply_to_vision_encoder: true
apply_to_mask_decoder: true
# Others: false
Maximum (Best Performance)
lora:
rank: 32
alpha: 64
target_modules: ["q_proj", "k_proj", "v_proj", "out_proj", "fc1", "fc2"]
apply_to_vision_encoder: true
apply_to_mask_decoder: true
apply_to_detr_encoder: true
apply_to_detr_decoder: true
Real-World Example: Concrete Crack Detection
SAM3-LoRA excels at detecting structural defects like cracks in concrete. Here's a real example:
Detection Results:
- Prompt: "concrete crack"
- Confidence: 0.32 (using threshold 0.3)
- Segmentation: Precise mask following the crack pattern
- Application: Infrastructure inspection, structural health monitoring
Run this example:
# Detect cracks in concrete structures
python3 infer_sam.py \
--config configs/full_lora_config.yaml \
--image path/to/concrete.jpg \
--prompt "concrete crack" \
--threshold 0.3 \
--output crack_detection.png
# Detect multiple defect types
python3 infer_sam.py \
--config configs/full_lora_config.yaml \
--image path/to/concrete.jpg \
--prompt "crack" "spalling" "corrosion" \
--threshold 0.3 \
--output defect_analysis.png
Use Cases:
- 🏗️ Civil engineering inspection
- 🌉 Bridge and infrastructure monitoring
- 🏢 Building maintenance
- 🛣️ Road surface analysis
- 🏭 Industrial facility assessment
Test Results: Road Damage Detection
We evaluated the fine-tuned SAM3-LoRA model on pothole detection, comparing it against the base SAM3 model without fine-tuning.
Validation Metrics Comparison
Validation performance: LoRA fine-tuned model vs Base SAM3 model
Key Findings:
- LoRA Model (Fine-tuned): Shows improved precision and better detection of multiple potholes
- Base Model: Tends to produce more false positives and misses some instances
- Dataset: Pothole detection on road surfaces (data3)
Visual Comparison
Side-by-side comparison: Ground Truth (Green) | LoRA Model (Red) | Base Model (Blue)
Observations from Visual Results:
| Image | Ground Truth | LoRA Model | Base Model | Analysis |
|---|---|---|---|---|
| img_0034 | 1 pothole | 1 detection ✓ | 5 detections ✗ | LoRA matches GT perfectly, Base has 4 false positives |
| img_0001 | 1 pothole | 1 detection ✓ | 1 detection ✓ | Both models perform well |
| img_0080 | 1 pothole | 2 detections ~ | 2 detections ~ | Both have 1 false positive |
| img_0070 | 1 pothole | 1 detection ✓ | 1 detection ✓ | Both models perform well |
| img_0060 | 4 potholes | 4 detections ✓ | 2 detections ✗ | LoRA finds all instances, Base misses 2 |
Summary:
- LoRA Model: 3/5 perfect matches, better recall on multi-instance images
- Base Model: 2/5 perfect matches, struggles with multiple instances and false positives
- Overall: Fine-tuning with LoRA significantly improves detection accuracy for domain-specific tasks
Training Details:
- Prompt: "pothole" (auto-detected from COCO category names)
- Architecture: Full LoRA adaptation (vision, text, DETR encoders/decoders)
- Dataset: Road damage images with COCO-format annotations
- Threshold: 0.5 confidence for both models
Examples
Example 1: Quick Test (5 Epochs)
# Create minimal config
cat > configs/quick_test.yaml << EOF
lora:
rank: 4
alpha: 8
dropout: 0.1
target_modules: ["q_proj", "v_proj"]
apply_to_vision_encoder: true
apply_to_mask_decoder: false
training:
batch_size: 1
num_epochs: 5
learning_rate: 1e-4
weight_decay: 0.01
output:
output_dir: "outputs/quick_test"
EOF
# Train
python3 train_sam3_lora_native.py --config configs/quick_test.yaml
# Inference with text prompt
python3 infer_sam.py \
--config configs/quick_test.yaml \
--weights outputs/quick_test/best_lora_weights.pt \
--image test.jpg \
--prompt "car" \
--output result.png
# Multiple prompts
python3 infer_sam.py \
--config configs/quick_test.yaml \
--image test.jpg \
--prompt "car" "person" "bus" \
--output result.png
Example 2: Production Training
# Create production config
cat > configs/production.yaml << EOF
lora:
rank: 32
alpha: 64
dropout: 0.1
target_modules: ["q_proj", "k_proj", "v_proj", "fc1", "fc2"]
apply_to_vision_encoder: true
apply_to_mask_decoder: true
apply_to_detr_encoder: true
apply_to_detr_decoder: true
training:
batch_size: 2
num_epochs: 50
learning_rate: 3e-5
weight_decay: 0.01
output:
output_dir: "outputs/production"
EOF
# Train (single GPU)
python3 train_sam3_lora_native.py --config configs/production.yaml
# Train (multi-GPU - 2 GPUs)
python3 train_sam3_lora_native.py --config configs/production.yaml --device 0 1
# Train (multi-GPU - 4 GPUs)
python3 train_sam3_lora_native.py --config configs/production.yaml --device 0 1 2 3
Example 3: Multi-GPU Training
# Quick 2-GPU training
python3 train_sam3_lora_native.py \
--config configs/full_lora_config.yaml \
--device 0 1
# 4-GPU training for large datasets
python3 train_sam3_lora_native.py \
--config configs/full_lora_config.yaml \
--device 0 1 2 3
# Use specific GPUs (e.g., skip GPU 1)
python3 train_sam3_lora_native.py \
--config configs/full_lora_config.yaml \
--device 0 2 3
# With custom master port (if default 29500 is in use)
python3 train_sam3_lora_native.py \
--config configs/full_lora_config.yaml \
--device 0 1 \
--master_port 29501
Tips for Multi-GPU Training:
- Effective batch size =
batch_size × num_gpus - Learning rate can be scaled:
lr × num_gpus(optional, try both) - Memory per GPU stays the same as single-GPU
- Training time scales roughly linearly with GPU count
Example 4: Programmatic Training
from train_sam3_lora_native import SAM3TrainerNative
# Create trainer
trainer = SAM3TrainerNative("configs/full_lora_config.yaml")
# Train
trainer.train()
# Weights saved to: outputs/sam3_lora_full/lora_weights.pt
Example 5: Batch Inference with Text Prompts
from infer_sam import SAM3LoRAInference
from pathlib import Path
# Initialize once
inferencer = SAM3LoRAInference(
config_path="configs/full_lora_config.yaml",
weights_path="outputs/sam3_lora_full/best_lora_weights.pt"
)
# Process multiple images with same prompt
image_dir = Path("test_images")
output_dir = Path("predictions")
output_dir.mkdir(exist_ok=True)
for img_path in image_dir.glob("*.jpg"):
predictions = inferencer.predict(
str(img_path),
text_prompts=["car"]
)
output_path = output_dir / f"{img_path.stem}_pred.png"
inferencer.visualize(
predictions,
str(output_path)
)
print(f"✓ Processed {img_path.name}")
# Process with multiple prompts per image
for img_path in image_dir.glob("*.jpg"):
# Detect multiple object types at once
predictions = inferencer.predict(
str(img_path),
text_prompts=["crack", "defect", "damage"]
)
output_path = output_dir / f"{img_path.stem}_multi.png"
inferencer.visualize(predictions, str(output_path))
# Print summary
for idx in range(3):
result = predictions[idx]
print(f" {result['prompt']}: {result['num_detections']} detections")
Advanced Usage
Apply LoRA to Custom Models
from lora_layers import LoRAConfig, apply_lora_to_model, count_parameters
import torch.nn as nn
# Your PyTorch model
model = YourModel()
# Configure LoRA
lora_config = LoRAConfig(
rank=8,
alpha=16,
dropout=0.1,
target_modules=["q_proj", "k_proj", "v_proj"],
apply_to_vision_encoder=True,
apply_to_text_encoder=False,
apply_to_geometry_encoder=False,
apply_to_detr_encoder=False,
apply_to_detr_decoder=False,
apply_to_mask_decoder=False,
)
# Apply LoRA (automatically freezes base model)
model = apply_lora_to_model(model, lora_config)
# Check trainable parameters
stats = count_parameters(model)
print(f"Trainable: {stats['trainable_parameters']:,} / {stats['total_parameters']:,}")
print(f"Percentage: {stats['trainable_percentage']:.2f}%")
# Train normally
optimizer = torch.optim.AdamW(
[p for p in model.parameters() if p.requires_grad],
lr=1e-4
)
Save and Load LoRA Weights
from lora_layers import save_lora_weights, load_lora_weights
# Save only LoRA parameters (small file!)
save_lora_weights(model, "my_lora_weights.pt")
# Load into new model
load_lora_weights(model, "my_lora_weights.pt")
Project Structure
SAM3_LoRA/
├── configs/ # Training configs (base, full, light, minimal, crack detection)
│ └── full_lora_config.yaml # Default training config
├── data/ # COCO format dataset
│ ├── train/
│ │ ├── img001.jpg # Training images
│ │ ├── img002.jpg
│ │ └── _annotations.coco.json # COCO annotations
│ ├── valid/
│ │ ├── img001.jpg # Validation images
│ │ ├── img002.jpg
│ │ └── _annotations.coco.json # COCO annotations
│ └── test/
│ ├── img001.jpg # Test images (optional)
│ └── _annotations.coco.json # COCO annotations
├── outputs/
│ └── sam3_lora_full/
│ ├── best_lora_weights.pt # Best model (lowest val loss)
│ └── last_lora_weights.pt # Last epoch model
├── sam3/ # SAM3 model library
├── lora_layers.py # LoRA implementation
├── train_sam3_lora_native.py # Training script (computes validation loss only)
├── train_sam3_lora_with_categories.py # Alternative trainer (category-focused)
├── validate_sam3_lora.py # Full evaluation script (mAP, cgF1, NMS)
├── validate_single_image.py # Single image validation with visualization
├── infer_sam.py # Inference script (recommended)
├── inference_lora.py # Legacy inference script
├── merge_lora_weights.py # Fold a trained adapter into the base weights
├── convert_refcoco_to_coco.py # RefCOCO/RRSIS-D -> COCO layout converter
├── README_INFERENCE.md # Detailed inference guide
└── README.md # This file
Troubleshooting
Common Issues
1. Hugging Face Authentication Error
Error: Access denied to facebook/sam3
Solution:
- Make sure you've requested access at https://huggingface.co/facebook/sam3
- Wait for approval (check your email)
- Run
huggingface-cli loginand paste your token - Or set:
export HF_TOKEN="your_token"
2. Import Errors
# Make sure package is installed
pip install -e .
3. CUDA Out of Memory
# Reduce batch size and rank in config
training:
batch_size: 1
lora:
rank: 4
4. Very Low Loss (< 0.001)
- Model may be overfitting
- Reduce LoRA rank
- Add more dropout
- Check if base model is properly frozen
5. Loss Not Decreasing
- Increase learning rate
- Increase LoRA rank
- Train for more epochs
- Check data quality
6. Wrong Number of Trainable Parameters
Expected: ~0.5-2% (for rank 4-16)
If you see 63%: Base model not frozen (bug fixed in latest version)
7. No Validation Data
⚠️ No validation data found - training without validation
Solution:
- Create
data/valid/directory with same structure asdata/train/ - Split your data: ~80% train, ~20% validation
- Training will work without validation but you won't see validation metrics
8. Annotation Format Errors
FileNotFoundError: COCO annotation file not found: /path/to/data/train/_annotations.coco.json
Solution:
- Ensure your data is in COCO format with
_annotations.coco.jsonin each split folder - Each split (train/valid/test) needs its own annotation file
- Images should be in the same directory as the annotation file
- Supported segmentation formats: polygon lists or RLE dictionaries
9. Want to See mAP/cgF1 During Training? Solution:
- Training only computes validation loss (fast; mirrors how SAM3's code splits loss-during-training from offline COCO/cgF1 evaluators — see the note in Quick Start)
- After training, run
validate_sam3_lora.pyfor full metrics with NMS - This approach significantly speeds up training while still monitoring overfitting
- Validation loss is sufficient to detect overfitting and select best model
10. Grid-Like Bounding Box Pattern in Inference
Problem: Visualization shows many overlapping boxes forming a grid pattern
Cause: Missing NMS (Non-Maximum Suppression) filtering. SAM3 uses 100+ object queries that produce many overlapping predictions.
Solution:
# Use lower NMS IoU threshold to remove overlapping boxes
python3 infer_sam.py \
--config configs/full_lora_config.yaml \
--image test.jpg \
--prompt "object" \
--nms-iou 0.3 \
--output clean_output.png
NMS IoU values:
0.3- Aggressive filtering (fewer boxes, cleaner output)0.5- Default, balanced0.7- Keep more overlapping detections
Performance Benchmarks
| Configuration | Trainable Params | Checkpoint Size | GPU Memory | Speed |
|---|---|---|---|---|
| Minimal (r=4) | ~0.2% | ~10 MB | 8 GB | Fast |
| Balanced (r=8) | ~0.5% | ~20 MB | 12 GB | Medium |
| Full (r=16) | ~1.0% | ~40 MB | 16 GB | Slower |
| Maximum (r=32) | ~2.0% | ~80 MB | 20 GB | Slowest |
Benchmarks on NVIDIA RTX 3090
Troubleshooting & Performance Optimization
Problem: Out of Memory (OOM) During Training
Symptoms:
Killed (exit code 137)
Training crashes after a few batches
Solutions (based on SAM3 original approach):
- Use Light LoRA Config (Recommended for GPUs with <24GB VRAM):
This HTML preview is truncated for page performance. The canonical Markdown file contains the complete snapshot.
Why MDRSS assigned this score
- evidence comes from multiple domains
- some evidence URLs look like primary-source hosts
Evidence (4)
concept:training-fine-tuningorg:collider-club Discussion 0
Sign in to join the discussion.