Semantic annotation
An ontology for off-road scene understanding.
The annotation taxonomy extends the foundation of the RELLIS-3D dataset with terrain and object classes that appear frequently in the Great Outdoors sequences, including gravel and mulch.
Open Ontology ImageLabel set
Twenty-two classes for terrain, objects, and void regions.
The classes cover vegetation, traversable and non-traversable terrain, structures, agents, vehicles, obstacles, and unlabeled areas for semantic segmentation research.
- Void
- Dirt
- Grass
- Trees
- Pole
- Water
- Sky
- Vehicle
- Object
- Asphalt
- Building
- Log
- Person
- Fence
- Bush
- Concrete
- Barrier
- Puddle
- Mud
- Rubble
- Mulch
- Gravel
- Snow
Caption annotation
Human-refined language labels for multimodal learning.
Captions are produced with a semi-automatic human-in-the-loop workflow designed to scale language annotation while keeping descriptions grounded in the visual data.
Representative frames
Dense image sequences are downsampled with CLIP-based visual features to select diverse frames across scene conditions.
Model drafts
Qwen2.5-72B generates initial short and long captions that summarize scenes, terrain, objects, and spatial context.
Human refinement
Annotators correct errors, reduce hallucinations, and standardize language so the final captions support vision-language and multimodal learning tasks.
Annotation examples
Representative RGB, LWIR, and language annotations.
This sample shows the paired visual annotations used in the release: semantic labels, RGB overlays, projected LWIR overlays, and human-refined captions.
Short caption
Snow-covered trail with tire tracks leads toward a garage, flanked by bare trees and a utility pole under a clear sky.
Long caption
The scene shows uneven, snow-covered off-road terrain with tire tracks curving toward a gabled building with garage doors. A utility pole and leafless trees border the area, while the open snow field remains navigable with obstacle avoidance around the structure and pole.
Statistics
Class distribution and annotation examples.
These visual summaries help researchers understand the semantic distribution before training or evaluating models.