A JEPA Recipe for Tabular Foundation Models
研究解决了表格基础模型中JEPA架构的潜在项崩溃问题,通过调整价值头读取编码器字段及使用EMA差异的方法,提高了模型在分类和回归任务上的表现。
研究解决了表格基础模型中JEPA架构的潜在项崩溃问题,通过调整价值头读取编码器字段及使用EMA差异的方法,提高了模型在分类和回归任务上的表现。
This study addresses the challenge of reverse-identifying molecular targets and compounds from single-cell transcriptional perturbation responses. The authors propose the first multi-task Transformer-based retrieval model tailored for single-cell perturbation data, which jointly learns target prediction and molecular embedding within a fixed compound library. The model performs end-to-end inverse inference using differential expression profiles relative to cell-type-specific DMSO controls and incorporates a structure–transcriptome alignment constraint to enhance representational consistency. Evaluated on the Tahoe-100M dataset, the model achieves a target Recall@10 of 0.408 and a compound Hit@1 of 0.129, significantly outperforming baseline methods and demonstrating its effectiveness in retrieving known perturbation pairs.
This study addresses the challenge of multi-hop reasoning across tens of thousands of pages in nuclear regulatory document review, where evidence is highly dispersed. The authors propose a state-aware planning framework based on large language models (LLMs) that formulates reasoning as dynamic navigation within an unvectorized document tree. An agent progressively constructs and updates an internal dynamic knowledge graph through browsing, reading, and search actions until sufficient evidence is gathered. A novel auditable edge-reasoning module is introduced to enhance decision traceability without requiring offline indexing. Evaluated on the NuScale FSAR 200-question benchmark, the method achieves 81.5% accuracy and a RAGAS faithfulness score of 0.93, substantially outperforming existing approaches such as PageIndex (+38.0 percentage points), LightRAG, HippoRAG, and GraphRAG.
This work addresses the challenge of limited feature generalizability in respiratory sound classification caused by variations in recording quality and class imbalance. To this end, the authors propose QLung, a novel framework that introduces, for the first time, a no-reference audio quality assessment metric based on spectral entropy and root-mean-square energy. This metric dynamically adjusts the angular margin in a normalized angular classifier and is combined with a log-scaling strategy to enhance intra-class compactness and inter-class separability, thereby stabilizing training. Evaluated on the ICBHI dataset, QLung achieves a 2.46% improvement over the cross-entropy baseline and demonstrates state-of-the-art out-of-distribution generalization performance on the SPRSound dataset.
This study addresses the limitation of existing self-attention-based models for respiratory sound classification, which suffer from low-pass filtering effects that hinder effective modeling of localized abnormal acoustic features. To overcome this, the work introduces state space models (SSMs) to the task for the first time and proposes a spectrogram-aware regularization technique to better preserve mid- and high-frequency components. Additionally, a dual-axis Patch-Mix contrastive learning mechanism tailored for audio SSMs is designed to enhance feature discriminability. Evaluated on the ICBHI benchmark, the proposed method achieves a composite score of 64.48%, representing a 5% improvement over the Audio Spectrogram Transformer baseline, thereby validating the effectiveness of the proposed architecture and training strategies.
研究解决了表格基础模型中JEPA架构的潜在项崩溃问题,通过调整价值头读取编码器字段及使用EMA差异的方法,提高了模型在分类和回归任务上的表现。
This study addresses the challenge of reverse-identifying molecular targets and compounds from single-cell transcriptional perturbation responses. The authors propose the first multi-task Transformer-based retrieval model tailored for single-cell perturbation data, which jointly learns target prediction and molecular embedding within a fixed compound library. The model performs end-to-end inverse inference using differential expression profiles relative to cell-type-specific DMSO controls and incorporates a structure–transcriptome alignment constraint to enhance representational consistency. Evaluated on the Tahoe-100M dataset, the model achieves a target Recall@10 of 0.408 and a compound Hit@1 of 0.129, significantly outperforming baseline methods and demonstrating its effectiveness in retrieving known perturbation pairs.
This study addresses the challenge of multi-hop reasoning across tens of thousands of pages in nuclear regulatory document review, where evidence is highly dispersed. The authors propose a state-aware planning framework based on large language models (LLMs) that formulates reasoning as dynamic navigation within an unvectorized document tree. An agent progressively constructs and updates an internal dynamic knowledge graph through browsing, reading, and search actions until sufficient evidence is gathered. A novel auditable edge-reasoning module is introduced to enhance decision traceability without requiring offline indexing. Evaluated on the NuScale FSAR 200-question benchmark, the method achieves 81.5% accuracy and a RAGAS faithfulness score of 0.93, substantially outperforming existing approaches such as PageIndex (+38.0 percentage points), LightRAG, HippoRAG, and GraphRAG.
This work addresses the challenge of limited feature generalizability in respiratory sound classification caused by variations in recording quality and class imbalance. To this end, the authors propose QLung, a novel framework that introduces, for the first time, a no-reference audio quality assessment metric based on spectral entropy and root-mean-square energy. This metric dynamically adjusts the angular margin in a normalized angular classifier and is combined with a log-scaling strategy to enhance intra-class compactness and inter-class separability, thereby stabilizing training. Evaluated on the ICBHI dataset, QLung achieves a 2.46% improvement over the cross-entropy baseline and demonstrates state-of-the-art out-of-distribution generalization performance on the SPRSound dataset.
This study addresses the limitation of existing self-attention-based models for respiratory sound classification, which suffer from low-pass filtering effects that hinder effective modeling of localized abnormal acoustic features. To overcome this, the work introduces state space models (SSMs) to the task for the first time and proposes a spectrogram-aware regularization technique to better preserve mid- and high-frequency components. Additionally, a dual-axis Patch-Mix contrastive learning mechanism tailored for audio SSMs is designed to enhance feature discriminability. Evaluated on the ICBHI benchmark, the proposed method achieves a composite score of 64.48%, representing a 5% improvement over the Audio Spectrogram Transformer baseline, thereby validating the effectiveness of the proposed architecture and training strategies.