Institution profile

City College of New York

Academic institutionnorthamerica · us
Official website
Research library48linked papers
Opportunities0open roles
Selected work

Representative Papers

Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages

Aug 12, 2026

This study addresses the systemic neglect of low-resource languages in current AI infrastructure across data curation, tokenization, evaluation, and deployment, which exacerbates educational and linguistic inequities. Focusing on Bengali as a case study, the work integrates multilingual corpus analysis, tokenization efficiency benchmarks, internet penetration statistics, and modeling of educational resource accessibility to expose structural barriers: extreme training data scarcity (with an English-to-Bengali data ratio of 67:1), high tokenization overhead due to syllabic orthography, limited online content, and a pronounced rural–urban digital divide. The research reframes data scarcity not merely as a technical bottleneck but as a manifestation of structural injustice and advocates for an “offline-first” infrastructure design paradigm to advance linguistic equity and foster more inclusive AI development.

0 citationsRead paper

Integrated Multimodal AI System for Retrieval-Augmented Reasoning, Object Sensing, and Damage Analysis

Aug 09, 2026

This work addresses the challenges of reliable damage assessment under adverse lighting and weather conditions, where conventional vision-based methods often fail and language models are prone to hallucination due to insufficient grounding in domain-specific documentation. To overcome these limitations, the authors propose a unified multimodal AI system that integrates retrieval-augmented generation (RAG), knowledge graphs, thermal-infrared and visible-light imaging, and wireless signal sensing. A novel hybrid retrieval mechanism combining graph-structured and vector-based representations is introduced to enhance cross-document reasoning. Additionally, a vision-language model generates synthetic damage data to augment training. Experimental results demonstrate that dynamic retrieval significantly improves factual consistency, graph-based retrieval outperforms purely vector-based approaches, and multimodal fusion effectively mitigates the constraints of individual sensors, collectively enhancing damage classification accuracy.

0 citationsRead paper
Recent publications

Latest Papers

Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages

Aug 12, 2026

This study addresses the systemic neglect of low-resource languages in current AI infrastructure across data curation, tokenization, evaluation, and deployment, which exacerbates educational and linguistic inequities. Focusing on Bengali as a case study, the work integrates multilingual corpus analysis, tokenization efficiency benchmarks, internet penetration statistics, and modeling of educational resource accessibility to expose structural barriers: extreme training data scarcity (with an English-to-Bengali data ratio of 67:1), high tokenization overhead due to syllabic orthography, limited online content, and a pronounced rural–urban digital divide. The research reframes data scarcity not merely as a technical bottleneck but as a manifestation of structural injustice and advocates for an “offline-first” infrastructure design paradigm to advance linguistic equity and foster more inclusive AI development.

0 citationsRead paper

Integrated Multimodal AI System for Retrieval-Augmented Reasoning, Object Sensing, and Damage Analysis

Aug 09, 2026

This work addresses the challenges of reliable damage assessment under adverse lighting and weather conditions, where conventional vision-based methods often fail and language models are prone to hallucination due to insufficient grounding in domain-specific documentation. To overcome these limitations, the authors propose a unified multimodal AI system that integrates retrieval-augmented generation (RAG), knowledge graphs, thermal-infrared and visible-light imaging, and wireless signal sensing. A novel hybrid retrieval mechanism combining graph-structured and vector-based representations is introduced to enhance cross-document reasoning. Additionally, a vision-language model generates synthetic damage data to augment training. Experimental results demonstrate that dynamic retrieval significantly improves factual consistency, graph-based retrieval outperforms purely vector-based approaches, and multimodal fusion effectively mitigates the constraints of individual sensors, collectively enhancing damage classification accuracy.

0 citationsRead paper