Score
Design and implement bridge networks or embedding mappers that project kinematic or other physical sensor tensors into text-aligned semantic latent spaces, producing embeddings compatible with language-aligned models. These mappings are built and evaluated to enable semantic-level analyses such as chunk-level anomaly scoring without instance labels, cross-modal retrieval, and robust out-of-distribution transfer.
To address cross-sensor domain shift in bridge 3D semantic segmentation, this paper introduces the first bridge-specific 3D semantic segmentation benchmark for structural health monitoring, comprising high-precision LiDAR and photogrammetric scans from multiple countries, with fine-grained component-level annotations. We conduct cross-sensor generalization evaluation using three state-of-the-art models—PointPillars, KPConv, and SPVCNN—and quantitatively demonstrate, for the first time, that sensor-induced domain shift degrades mIoU by up to 11.4%. We further propose a reproducible empirical framework for domain shift assessment. Our key contributions are: (1) establishing the first dedicated 3D semantic segmentation dataset for bridge structures; (2) systematically characterizing the extent of sensor-induced domain shift; and (3) providing a standardized benchmark to advance intelligent bridge condition diagnosis and domain adaptation algorithm development.
研究探讨了嵌入模型在表达物理量度(如质量、距离、时间、体积)时的局限性,发现这些模型受表面字符串相似性影响较大,重新校准相似度也无法显著改善。
Existing data-driven multivariate time series (MTS) graph construction methods suffer from small-sample bias and fail to accurately capture spatiotemporal dependencies, thereby limiting graph neural network (GNN) performance. To address this, we propose a knowledge-enhanced graph construction framework: first, we explicitly model domain priors—such as physical laws—implicitly encoded in large language models (LLMs) as sensor-level knowledge linkage graphs; second, we design a differentiable graph alignment module to semantically fuse knowledge graphs with signal-driven graphs. Our approach enables transfer of generic knowledge to specific MTS scenarios via prompt engineering, knowledge graph extraction, and cross-graph alignment. Extensive experiments on diverse MTS downstream tasks—including classification and forecasting—demonstrate substantial improvements over state-of-the-art GNNs. Quantitative graph structure evaluation further confirms that knowledge injection significantly enhances both discriminability and interpretability of the learned graphs.
This work identifies systematic deficiencies in network foundation models (NFMs): underutilized representation space, poor alignment with domain-expert features, and weak robustness to protocol-level perturbations. To address this, we propose the first intrinsic-representation-oriented, three-dimensional evaluation framework—comprising geometric analysis (anisotropy quantification), alignment assessment (metric-space consistency), and causal sensitivity testing (contextual disentanglement capability)—and systematically evaluate four state-of-the-art NFMs across five real-world and controllable network datasets. We find pervasive representation degradation across mainstream models, attributable to the decoupling of training objectives from network semantics. Building on these diagnostics, we design a lightweight post-optimization strategy that preserves model architecture integrity while achieving up to a 0.35 improvement in F1-score. Our results underscore the critical role of representation diagnosis in enabling trustworthy, semantically grounded NFM evolution.
Large language models (LLMs) struggle to process raw motion sensor time-series data due to semantic sparsity, numerical input incompatibility, and computational constraints. To address this, we propose SensorLLM—a two-stage sensor-to-language alignment framework. Its core contributions are: (1) channel-specific special tokens coupled with auto-generated trend-oriented textual descriptions, enabling semantic encoding of multichannel, variable-length numeric sequences; and (2) an integrated pipeline combining textualized sequence representation, special token embedding, instruction tuning, and task-aware LoRA adaptation—enabling zero-shot human activity recognition (HAR). Evaluated across multiple benchmarks, SensorLLM achieves or surpasses state-of-the-art performance, demonstrating high accuracy, cross-device transferability, and strong generalization capability.
Numerical measurements inherently lack variable semantics, and existing methods relying on manual priors are prone to introducing bias or information loss. This work proposes CausalBridge, a framework that bridges this semantic gap through causal structure. By employing causal discovery to recover causal graphs solely from data, it eliminates human bias. Furthermore, it integrates structurally constrained semantic alignment, embedding-based solving, and large language models to infer the meanings of unnamed variables based on inter-variable dependencies. Experimental results demonstrate that the proposed method accurately recovers the semantics of both observed and latent variables across diverse scenarios. It significantly reduces naming costs while outperforming conventional association-based approaches in overall performance.
This work addresses the limitation of existing video anomaly detection methods, which predominantly rely on reactive multiple instance learning and struggle to proactively anticipate anomalies. The authors propose PULS, a continuous semantic world model that maps physical motion representations to a semantic embedding space via a KSD bridge and incorporates an ASP module to enhance the semantic separability of future states—enabling anomaly prediction without multiple instance learning. The study validates the “latent clarity” hypothesis, demonstrating that anticipated future representations exhibit greater semantic separability than current observations and reside on distinct submanifolds. Built upon V-JEPA 2, Qwen3-VL-Embedding-2B, and an L1-surprise gating mechanism, PULS achieves 0.8994 AUROC on UCF-Crime, 0.8162 zero-shot transfer performance on XD-Violence, an 8.9% early-prediction advantage at T−0.5s, and a zero-shot VQA accuracy of 44.5%, substantially outperforming observation-based baselines.
本文提出SOfA模型,通过引入可学习的标准化关节槽和语义关节嵌入解决不同传感器间骨架数据异质性问题,实现跨传感器统一骨架表示学习。
为解决多模态临床AI中输入对齐差及缺乏领域特定可解释表示的问题,本文提出基于运动学知识图谱和模板文本的ScoliDetect框架,通过双向交叉注意与潜在瓶颈聚合方法提高了青少年特发性脊柱侧弯筛查的准确性和可解释性。
This work addresses the scalability limitations of existing semantic-guided dimensionality reduction methods, which incur linearly growing computational costs due to per-sample invocation of large language models (LLMs). To overcome this, the authors propose a group-level semantic prototyping framework that shifts semantic computation from individual samples to user-defined groups. A single LLM call generates a structured group profile, which—combined with seed centroids—forms hybrid semantic prototypes. Efficient semantic alignment in the embedding space is then achieved through soft assignment, an abstention mechanism, and alignment-aware scaling updates. Evaluated on LitCovid (5K documents), the method reduces LLM invocations by over three orders of magnitude while matching the global alignment performance of per-sample approaches. Its multimodal applicability is further demonstrated on image data.