research
Satellite Damage Became a Text Sequence
GeBDA asks a general vision-language model to emit building boxes and damage labels as one variable-length sequence.

Summary
GeBDA asks a general vision-language model to emit building boxes and damage labels as one variable-length sequence.
Instead of a dedicated detector, the preliminary system represents each building as coordinates followed by a damage class and predicts the full set autoregressively from before-and-after satellite images. The open Gemma-based implementation produced promising localization and grading results under the paper's prompt formulation. The abstract does not establish operational disaster-response readiness, and the authors characterize the implementation as preliminary.
Why it matters
GeBDA asks a general vision-language model to emit building boxes and damage labels as one variable-length sequence.
Limits and context
- The abstract does not establish operational disaster-response readiness, and the authors characterize the implementation as preliminary.
Key claims
GeBDA asks a general vision-language model to emit building boxes and damage labels as one variable-length sequence.
Qualification: The abstract does not establish operational disaster-response readiness, and the authors characterize the implementation as preliminary.
Evidence: source-2026-08-31-013
Sources
- arXiv preprint 2608.28567arXiv · primary research
Corrections
No corrections have been recorded for this story.