Omni2Web: Benchmarking Audiovisual Website Development
研究通过构建Omni2Web基准,利用音频和视觉信息解决网页编辑请求中指示代词指代不明的问题,评估多种模型在直接编辑、意图恢复及指令实用性上的表现。
研究通过构建Omni2Web基准,利用音频和视觉信息解决网页编辑请求中指示代词指代不明的问题,评估多种模型在直接编辑、意图恢复及指令实用性上的表现。
研究提出LAWA架构,通过紧凑的潜在动作表示未来意图,解决WAM中未来观察生成导致的延迟问题,提高效率和泛化能力。
This study addresses the challenge of synthesizing structural MRI scans for Alzheimer’s disease (AD), where neurodegenerative changes are subtle, region-specific, and progressive. The authors propose the first application of the conditional diffusion model Med-DDPM to generate AD-specific 3D MRI images, using anatomical segmentation masks from the ADNI dataset as conditioning inputs to accurately capture AD-related pathological alterations. Experimental results demonstrate that a segmentation model trained solely on synthetic data achieves a Dice score of 0.6532, comparable to that obtained with real data (0.6513). Furthermore, training with a hybrid dataset combining real and synthetic images yields the best performance, improving the Dice score to 0.7244 and significantly enhancing model recall and clinical diagnostic potential.
This work addresses the inefficiency of conventional scheduling in PIM-based autoregressive LLM inference, where performance is bottlenecked by the DRAM row cycle time (nRC), rendering optimizations targeting nCCDAB ineffective for GEMV operations. To overcome this limitation, the authors propose RH+, a novel scheduling method that reengineers the address mapping strategy to co-locate consecutive MAC operations within the same DRAM row. By merely adjusting the access stride, RH+ substantially enhances row locality, circumvents the nRC constraint, and overturns the traditional host-centric interleaving paradigm. Evaluated on an HBM3 PIM architecture using cycle-accurate simulation across four LLM workloads, RH+ achieves 8–12× speedup, over 74% energy reduction, and up to 52× improvement in energy-delay product (EDP).
This work addresses a critical mismatch between existing diffusion model acceleration techniques—which rely on element-wise activation sparsity—and the column-granularity data processing inherent in modern hardware, leading to overestimated practical sparsity benefits. The study presents the first systematic characterization of output sparsity across seven diffusion models at the column level, identifying three distinct activation distribution patterns and uncovering the interplay between model architecture and memory layout optimization. Through hardware-aware column-level sparsity profiling, cycle-accurate GDDR6 simulation, multi-threshold accuracy evaluation, and cross-modal comparison, the authors demonstrate that memory stalls account for 84–89% of total execution cycles and that element-wise sparsity poorly predicts actual hardware gains. Their approach achieves up to 30.6% reduction in execution cycles on UNet+Transformer models, with a maximum latency decrease (MLD) of 50.8%.
研究通过构建Omni2Web基准,利用音频和视觉信息解决网页编辑请求中指示代词指代不明的问题,评估多种模型在直接编辑、意图恢复及指令实用性上的表现。
研究提出LAWA架构,通过紧凑的潜在动作表示未来意图,解决WAM中未来观察生成导致的延迟问题,提高效率和泛化能力。
This study addresses the challenge of synthesizing structural MRI scans for Alzheimer’s disease (AD), where neurodegenerative changes are subtle, region-specific, and progressive. The authors propose the first application of the conditional diffusion model Med-DDPM to generate AD-specific 3D MRI images, using anatomical segmentation masks from the ADNI dataset as conditioning inputs to accurately capture AD-related pathological alterations. Experimental results demonstrate that a segmentation model trained solely on synthetic data achieves a Dice score of 0.6532, comparable to that obtained with real data (0.6513). Furthermore, training with a hybrid dataset combining real and synthetic images yields the best performance, improving the Dice score to 0.7244 and significantly enhancing model recall and clinical diagnostic potential.
This work addresses the inefficiency of conventional scheduling in PIM-based autoregressive LLM inference, where performance is bottlenecked by the DRAM row cycle time (nRC), rendering optimizations targeting nCCDAB ineffective for GEMV operations. To overcome this limitation, the authors propose RH+, a novel scheduling method that reengineers the address mapping strategy to co-locate consecutive MAC operations within the same DRAM row. By merely adjusting the access stride, RH+ substantially enhances row locality, circumvents the nRC constraint, and overturns the traditional host-centric interleaving paradigm. Evaluated on an HBM3 PIM architecture using cycle-accurate simulation across four LLM workloads, RH+ achieves 8–12× speedup, over 74% energy reduction, and up to 52× improvement in energy-delay product (EDP).
This work addresses a critical mismatch between existing diffusion model acceleration techniques—which rely on element-wise activation sparsity—and the column-granularity data processing inherent in modern hardware, leading to overestimated practical sparsity benefits. The study presents the first systematic characterization of output sparsity across seven diffusion models at the column level, identifying three distinct activation distribution patterns and uncovering the interplay between model architecture and memory layout optimization. Through hardware-aware column-level sparsity profiling, cycle-accurate GDDR6 simulation, multi-threshold accuracy evaluation, and cross-modal comparison, the authors demonstrate that memory stalls account for 84–89% of total execution cycles and that element-wise sparsity poorly predicts actual hardware gains. Their approach achieves up to 30.6% reduction in execution cycles on UNet+Transformer models, with a maximum latency decrease (MLD) of 50.8%.