SkillNet: Create, Evaluate, and Connect AI Skills
为解决AI技能缺乏系统积累和转移的问题,提出SkillNet,一个创建、评估和组织AI技能的开放基础设施。
为解决AI技能缺乏系统积累和转移的问题,提出SkillNet,一个创建、评估和组织AI技能的开放基础设施。
People with visual impairments face significant challenges in independent navigation and spatial orientation. Method: This study introduces a 3D-printed tactile map and icon system designed to support on-site orientation and mobility training. Employing an iterative, user-centered design process, we conducted field-based human factors evaluations—including tactile recognition and spatial cognition assessments—in real-world public environments, marking the first empirical validation of 3D tactile maps in authentic settings. Contribution/Results: Results demonstrate that realistic 3D icons are accurately recognized without legends, significantly enhancing mental map construction. The system improves users’ spatial cognition, autonomous navigation proficiency, and sense of belonging. Furthermore, the study distills empirically grounded guidelines and reusable design principles for inclusive design, offering both methodological frameworks and technical pathways to advance accessible built environments.
This study investigates whether current AI-based assistive device research aligns with the authentic needs of blind and low-vision (BLV) individuals. Method: We conducted a systematic literature review of 646 papers and in-depth interviews with 24 BLV users, integrating bibliometric analysis, user-need prioritization, and Spearman rank correlation testing. Contribution/Results: Our analysis reveals, for the first time, only a weak correlation between prevalent academic task formulations—such as object detection and image captioning—and actual user preferences. Instead, the top five most frequently cited needs center on real-time scene understanding and natural language–based conversational interaction. Users strongly prefer head-mounted, lightweight, and minimally intrusive devices. These findings challenge the dominant vision-centric paradigm in assistive AI research and provide empirical grounding for a user-centered design shift—emphasizing contextual awareness, multimodal interaction, and ergonomic form factors—thereby informing more effective, human-centered assistive technology development.
This study critically examines AI’s dual impact in healthcare: its transformative potential in genomics and public health, alongside profound ethical and institutional risks—including privacy breaches, algorithmic bias, physician deskilling, and imbalanced human–machine decision authority. Moving beyond technocentric paradigms, it introduces two foundational conceptual contributions: the reconfiguration of care as “datafied caregiving” and the normative calibration of “machine recommendation weight,” both grounded in philosophy of technology and bioethics. Employing an interdisciplinary analytical framework integrating medical ethics, philosophy of science, health policy, and big-data governance, the study uncovers structurally embedded risks overlooked in prevailing discourse. Its key contribution lies in reframing global regulatory and ethics review frameworks to center transparency, redistribution of epistemic and decisional authority, and preservation of clinical agency as core evaluative criteria.
Existing VideoQA datasets lack fine-grained modeling of professional sports actions, hindering effective reasoning for descriptive, temporal, causal, and counterfactual questions. To address this, we introduce Sports-QA—the first video question answering benchmark tailored to professional sports scenarios—covering multiple sports disciplines and four categories of complex reasoning tasks. Methodologically, we propose the Auto-Focus Transformer (AFT), which employs an attention-driven dynamic focusing mechanism to adaptively model multi-scale temporal information and integrates joint video–language representation learning. Extensive experiments demonstrate that AFT achieves state-of-the-art performance on Sports-QA, substantially outperforming general-purpose VideoQA models. This work constitutes the first systematic validation of an architecture explicitly designed for fine-grained sports action understanding and dynamic logical reasoning, establishing a new foundation for domain-specific VideoQA research.
本文提出LADDER框架,通过图引导的并行解码加速GraphRAG,解决多跳推理效率问题,采用事件驱动自时钟检索和不完全查询图传播方法。
本文提出TRACE模型,通过预设的临床潜在空间和路径分配方法,解决了深度学习ECG诊断模型不可审计的问题。
为了解决多语言环境下大语言模型代理评估不足的问题,通过BabelFlow方法构建了BabelArena基准,涵盖23种语言,用于评估模型在多语言任务中的表现。
为解决文本发布中的身份泄露问题,提出了一种基于统计保证的校准框架CPA,以评估针对大语言模型增强攻击者的重新识别风险。
该研究通过基于图模式的方法构建紧凑的部分对称性打破约束,解决了图搜索问题中对称性打破的挑战,提高了精度和效率。