Bridge inspection still runs on the expert eye. An inspector looks at a photo and decides whether that dark patch is spalling, rust or just a wet spot, and two inspectors can label the same image differently. For our Software Engineering capstone at Bahçeşehir University, we wanted to see how far a foundation model could take this: give it a photo, get back a pixel mask for each damage type.
The result is DetectIQ. It is Meta's SAM3 fine-tuned with LoRA on the DACL10K dataset, served behind a FastAPI and React inspection app at detectiq.com.tr. This post covers the model side: the setup, a first run that overfit, the changes that fixed it, and an honest read of the final numbers.
The team was Tarık Deveci, Osman Yiğit Alver and Ezgi Nilsu Kiraz.
The task and the data
DACL10K is a bridge inspection dataset with about 10,000 images and 19 classes: 13 damage types (Crack, Spalling, Rust, Efflorescence, Cavity and others) and 6 bridge components (Bearing, Drainage, expansion joints, protective equipment and others). Annotations are polygons, so each image can carry several overlapping classes.
This is semantic segmentation, not detection. We want a mask per class, not a box.
We used 6,935 training images and 975 validation images. Each polygon file is rasterized into one binary mask per class, which gives a 19 by 288 by 288 target per image.
Why SAM3, and why LoRA
SAM3 accepts a text prompt alongside the image and returns a segmentation for what the text describes. That maps cleanly onto a multi-class problem: pass "Crack" to get the crack mask, pass "Rust" to get the rust mask. One model, 19 prompts, no per-class head to design.
The catch is size. facebook/sam3 has 842M parameters, and we trained on a single GPU. Full fine-tuning was not realistic, so we used LoRA through Hugging Face PEFT: freeze the base model and train small low-rank adapters on the attention projections only.
from peft import LoraConfig, get_peft_model
lora_config = LoraConfig(
r=8,
lora_alpha=16,
lora_dropout=0.1,
target_modules=["q_proj", "v_proj"],
)
model = get_peft_model(model, lora_config)
That leaves 2,138,112 trainable parameters, about 0.25% of the model. The saved adapter is around 8 MB, which matters later: the product loads the frozen base once and applies a file small enough to version like code.
Images go in at 1024 by 1024. SAM3's semantic head returns 288 by 288 logits, so ground truth masks are downsized to the same resolution.
Run 1: one prompt for everything, and it overfit
The first full run was deliberately simple. Every image got the same text prompt, "damage", and the target was the union of all annotated classes. The loss was plain binary cross-entropy, with batch size 1, learning rate 1e-4 and 10 epochs. It ran on an RTX 4090 laptop GPU and took about 21.5 hours.
Training looked great: train IoU climbed from 0.49 to 0.69 and train loss fell from 0.43 to 0.17. Validation told a different story.
| Epoch | Val loss | Val IoU |
|---|---|---|
| 1 | 0.365 | 0.549 |
| 2 | 0.352 | 0.568 |
| 4 | 0.366 | 0.580 |
| 5 | 0.370 | 0.571 |
| 10 | 0.432 | 0.561 |
Validation loss bottomed out at epoch 2 and climbed from there to 0.432, while validation IoU peaked at epoch 4 and then drifted down. The model was memorizing the training bridges, and the best checkpoint was the fourth one, not the last.
A single "damage" mask was also the wrong product. An inspector needs to know whether it is rust or a crack, not just that something is wrong.
What we changed
Four changes went into the second run. None of them are exotic; the point was to fix the specific failure we saw.
1. Class names as prompts. Each training step looks at which classes are present in the image, picks one at random, and uses its name as the text prompt with that class's mask as the target.
active = [i for i in range(19) if multi_class_mask[0, i].sum() > 0]
idx = random.choice(active)
prompt = DACL10K_CLASSES[idx] # e.g. "Spalling"
target = multi_class_mask[:, idx, :, :]
Over an epoch, the model sees every class paired with its own name, and it has to learn to separate them instead of lighting up anything that looks unusual.
2. BCE plus Dice. Many damage regions are thin or small. Pixel-wise BCE is dominated by the background, so a model can score a low loss while missing a crack entirely. Dice measures overlap directly, so we weighted the two equally.
loss = 0.5 * bce(logits, target) + 0.5 * dice(logits, target)
3. Augmentation. Random crop and resize, small rotations (up to 8 degrees), brightness and contrast jitter, light blur and noise, and perspective shifts, each applied with a set probability. Bridge photos vary a lot in angle and lighting, and run 1 had seen each training image in only one form.
4. Cosine schedule with warmup, and early stopping. The learning rate warms up over the first 10% of steps and then decays along a cosine curve to 5% of its peak. Early stopping watches validation loss with a patience of 2 epochs, and the best adapter is saved separately from the latest one.
We also moved to a Colab A100 (40 GB) and made the loop resume from mid-epoch checkpoints, since a long run on a notebook can be cut off at any point.
Run 2 results
| Epoch | Train loss | Val loss | Val IoU |
|---|---|---|---|
| 1 | 0.472 | 0.346 | 0.481 |
| 4 | 0.310 | 0.296 | 0.526 |
| 7 | 0.256 | 0.274 | 0.534 |
| 9 | 0.231 | 0.266 | 0.540 |
| 10 | 0.224 | 0.267 | 0.540 |
This time validation loss fell for nine straight epochs and then flattened next to the training curve. The overfitting from run 1 is gone.
The validation IoU in that table is averaged over every (image, class) pair. If you average per class instead, so a rare class counts as much as a common one, the mean IoU is 0.56 across the 19 classes.
These numbers are not directly comparable to run 1. Run 1 scored a single "damage" mask; run 2 scores each class on its own, which is a harder task.
Where the model is good, and where it is not
| Group | Classes (IoU) |
|---|---|
| Above 0.60 | PEquipment 0.81, Drainage 0.75, Hollowareas 0.74, EJoint 0.72, Graffiti 0.70, Bearing 0.69, ExposedRebars 0.62, JTape 0.60 |
| 0.35 to 0.60 | ACrack 0.58, Weathering 0.55, Restformwork 0.54, Rust 0.51, Wetspot 0.50, Efflorescence 0.48, Rockpocket 0.44, Spalling 0.44, Crack 0.41, WConccor 0.40 |
| Below 0.35 | Cavity 0.19 |
The pattern is clear. Components with a defined shape (protective equipment, drainage, joints, bearings) segment well. Diffuse or thin damage does worse. Crack sits at 0.41 even though it is the class inspectors care about most.
Two hypotheses we have not tested yet:
- Resolution. Masks are predicted at 288 by 288 from a 1024 input. A crack a few pixels wide in the original image can shrink to almost nothing, and IoU punishes thin shapes hard, since a one-pixel offset can halve the overlap.
- Prompt wording. We used DACL10K's class codes as prompts. "Crack" and "Rust" are real words, but "WConccor", "PEquipment" and "JTape" are not. A text encoder has no prior for them, so those classes must be learned almost from scratch. Plain-language prompts such as "protective equipment" or "joint tape" are a cheap experiment.
Cavity at 0.19 is the weakest class. Cavities are small, irregular and easy to confuse with shadows or rockpockets. More data and targeted augmentation are the obvious next steps there.
What the metric does not measure
One limitation deserves to be stated plainly. Validation only scores the classes that are actually annotated in each image. If an image contains rust and the model also paints a crack that is not there, that false crack is never counted.
So 0.56 describes how well the model outlines a damage type when that type is present. It does not describe how often the model invents damage. The product works around this at inference time by dropping any class that covers less than 1% of the image. The metric itself should also score absent classes, and that is the first thing we would change in the evaluation.
From adapter to product
The training repo produces one artifact, the LoRA adapter. The product side is a FastAPI backend and a React front end:
- The user uploads a photo. The API stores it and returns immediately, and analysis runs as a background task.
- The image is preprocessed once, then passed through the model 19 times, once per class prompt.
- Each logit map goes through a sigmoid and a 0.5 threshold. Classes under 1% coverage are dropped as noise.
- A colored mask with a legend is rendered, and coverage and the dominant class are saved for each photo.
The whole app talks to one detector interface. A mock detector with no ML dependencies lets the service boot and run its tests without a GPU, and switching to the real model is a single environment variable. That seam let the product and the model be built in parallel and meet at the end.
What I would do next
- Score absent classes in validation so false positives count.
- Try plain-language prompts for the coded class names.
- Predict at a higher mask resolution, or tile large images, for thin classes like Crack.
- Batch the 19 prompts into fewer forward passes to cut inference cost.
The model code and training notes are on GitHub, the product is live at detectiq.com.tr, and there is a shorter case study on my site.