Score
Generating geometry-aligned seeds and models that capture canopy structure to detect treetop locations and estimate canopy height while ensuring coverage of understory trees and minimizing manual input through automated prompts.
This work proposes a zero-shot segmentation framework to address the challenges of segmenting densely overlapping tree crowns in remote sensing imagery, where supervised methods suffer from high annotation costs and poor generalization. The approach innovatively adapts the topological flow field concept from cellular instance segmentation to crown delineation, integrating semantic priors with star-convex geometric constraints to mathematically separate touching crowns without any training. Built upon a Cellpose-SAM architecture, the method enforces instance decoupling through a vector convergence mechanism within the topological flow field. Experiments demonstrate that the framework achieves strong cross-sensor and cross-density generalization on the NEON and BAMFOREST datasets, efficiently producing high-quality tree crown instance labels.
Existing forest image analysis lacks a unified end-to-end framework capable of efficiently handling large-scale data across multiple tasks, regions, and sensors. This work proposes the first natively geospatial, cloud-optimized, and deeply machine learning–integrated interactive platform that combines pretrained vision models, automated annotation, and human-in-the-loop refinement mechanisms. It enables efficient processing within a single workflow—from standard aerial imagery to orthomosaic datasets exceeding 100 GB—while directly producing downstream-ready outputs. Validated on the PALMS dataset, the platform supports continuous iterative updates of models and annotations for novel scenarios, facilitating scalable and adaptive forest monitoring.
This study addresses the challenges of instance segmentation of broadleaf tree crowns, which suffer from high morphological variability and ambiguous crown boundaries, leading to limited accuracy and poor generalization. Leveraging a dataset of 18,507 high-quality manually annotated crown polygons, we develop a deep learning framework centered on Mask2Former with various backbone networks, trained and evaluated exclusively on UAV-derived RGB orthoimagery. Our approach achieves high-detail, generalizable individual tree crown segmentation across diverse ecological settings—including temperate forests in Japan and tropical rainforests in Borneo—demonstrating for the first time the critical role of large-scale, fine-grained annotations in enhancing model generalization. The trained model has been integrated into DF Scanner Pro software to support operational forest monitoring applications.
Existing public 3D forest point cloud datasets are limited in scale and sparse in annotations, hindering robust single-tree segmentation and impeding quantitative assessment of ecosystem functions such as forest carbon sequestration. To address this, we propose a synthetic data generation pipeline integrating high-fidelity game-engine rendering with physics-based LiDAR simulation, producing a large-scale, diverse, and pixel-accurately annotated 3D forest synthetic dataset. We systematically demonstrate— for the first time—that physical realism, scene diversity, and dataset scale constitute the three critical factors enabling successful synthetic-data-driven single-tree segmentation. Pretraining deep learning models on our synthetic data enables fine-tuning on less than 0.1 hectare of real-world plot data to achieve segmentation accuracy comparable to models trained on full-scale real datasets, thereby substantially reducing field data acquisition and annotation costs.
A lack of high-quality, multi-temporal, multispectral, and LiDAR-derived canopy height model (CHM) co-registered open datasets hinders sub-meter tree-height prediction. Method: We introduce PrediTree—the first open, sub-meter (0.5 m), multi-temporal, multispectral image–CHM paired dataset covering diverse forest ecosystems across France, comprising 3.14 million samples—and propose a U-Net-based encoder–decoder architecture with explicit temporal awareness, jointly modeling multi-temporal multispectral imagery and inter-image time-difference cues for end-to-end CHM regression. Contribution/Results: Trained on PrediTree, our model achieves a masked mean squared error of 11.78%, outperforming ResNet-50 by ~12% and a single-RGB baseline by ~30%. This work bridges critical gaps in both high-resolution forest structural dynamics modeling—providing the first large-scale, temporally aligned benchmark—and methodological capability for fine-grained, time-aware CHM prediction.
This work addresses the challenge of instance segmentation in forest LiDAR point clouds, where scarce annotations and overlapping tree crowns hinder accurate delineation. To overcome these limitations, the authors propose a promptable 3D tree instance segmentation model that integrates a Click-to-Query prompt encoder with a canopy height model (CHM)-guided initial prompting mechanism. The architecture further incorporates a state-space-based query decoder, local feature pooling, and a fusion strategy leveraging geometric and ecological priors. Requiring only a single user click or automatically generated treetop prompts, the method achieves efficient large-scale segmentation with minimal annotation dependency and linear time complexity. Evaluated across seven diverse forest regions and an independent test set, the model attains a 78.2% IoU with fewer parameters and faster inference than existing promptable approaches.
This study addresses the high cost and labor intensity of traditional forest plot surveys, which hinder scalability. The authors propose a fully automated method that reconstructs circular forest plots and extracts tree parameters using only a single pass of video captured by an ordinary smartphone. This approach uniquely integrates consumer-grade hardware with SLAM, pretrained vision models, and monocular depth estimation—eliminating the need for specialized equipment. By combining trunk instance segmentation, depth maps, and camera pose estimation, along with reference-length calibration, the system operates robustly from arbitrary starting points. Evaluated in both plantation and natural forests, the method achieves diameter-at-breast-height measurement errors of 1.51 cm (MARE 3.98%) and 2.30 cm (MARE 5.69%), respectively—matching the accuracy of conventional techniques while substantially reducing cost and operational complexity.
This study addresses the challenging problem of reconstructing high-fidelity 3D tree point clouds from a single orthophoto and a digital surface model (DSM) without access to real 3D annotations, species labels, or ground-based laser scans. The authors propose a neural reconstruction framework trained exclusively on procedurally generated synthetic tree data, leveraging geometric supervision together with differentiable shading and silhouette rendering losses to accurately recover fine structural details of trees in real-world scenes. As the first single-view tree reconstruction method to integrate differentiable rendering and synthetic data under fully unsupervised conditions, it outperforms existing approaches in terms of reconstruction quality, structural plausibility, and generalization capability, making it well-suited for applications such as interactive 3D digital mapping.
This study addresses the challenge of low segmentation accuracy for individual tree crowns in structurally dense tropical forests. To this end, the authors introduce SelvaMask, a high-quality benchmark dataset comprising over 8,800 manually annotated tree crowns across multiple tropical countries, and propose a modular detection–segmentation pipeline that integrates a vision foundation model (VFM) with a domain-specific detection prompter. By leveraging a detection prompting mechanism to effectively adapt the VFM to the target domain, the method substantially improves segmentation performance in dense forest regions. It achieves state-of-the-art results on SelvaMask, outperforming both zero-shot general-purpose models and fully supervised end-to-end approaches, while also demonstrating strong generalization across multiple external datasets from both tropical and temperate forests.
Existing 3D sensing methods struggle to accurately reconstruct the fine-grained structure of understory saplings and lack repeatable, georeferenced capabilities for long-term monitoring. This work proposes a three-tier representation framework integrating Neural Radiance Fields (NeRF), LiDAR SLAM, and GNSS, achieving— for the first time—an implicit 3D reconstruction with real-world scale and geospatial coordinates. Evaluated in Wytham Woods (UK) and Evo forest (Finland), the method captures stem height, branching architecture, and leaf-to-wood ratios of saplings 0.5–2 m tall with centimeter-level accuracy, significantly outperforming conventional terrestrial laser scanning (TLS). The approach establishes a novel, quantitative paradigm for repeatable, long-term ecological monitoring of forest understory dynamics.