Institution profile

University of Puerto Rico

Academic institutionnorthamerica · us
Official website
Research library8linked papers
Opportunities0open roles
Selected work

Representative Papers

Limiting the Shrinkage for the Exceptional by Objective Robust Bayesian Analysis: the "Clemente Problem"

Jun 11, 2025

This paper addresses the “Clemente problem” in statistical inference—excessive shrinkage of outliers (e.g., elite athletes)—arising from the joint use of squared-error loss and light-tailed conjugate priors. To resolve this, we propose a Cauchy–Beta2 joint heavy-tailed prior framework, yielding a closed-form, objective, and robust location prior that unifies empirical Bayes (EB) and full Bayes (FB) approaches. Our method integrates heavy-tailed priors (e.g., Cauchy, double exponential), robust loss functions, and finite-translation estimators. Theoretically and empirically, the proposed paradigm substantially mitigates outlier shrinkage bias and achieves lower mean squared error than James–Stein estimation. Crucially, robust modeling choices exert a far greater impact on inference than the distinction between EB and FB methodologies. This work establishes the first analytically tractable, fully objective, and hyperparameter-free robust Bayesian inference framework—offering a principled alternative to conventional shrinkage estimation.

3 citationsRead paper

Pre-registered tests of solid-state-physics-inspired LLM compression: a cluster-level negative result at small-language-model scale

Sep 28, 2026

This study investigates whether compression methods inspired by solid-state physics can be effectively applied to small-scale language models. Employing rigorous research paradigms—including preregistration, three-sigma gating, and ensemble frameworks—the authors empirically evaluate five physical mapping hypotheses utilizing techniques such as Wannier functions, tight-binding models, DMRG truncation, and Wilsonian renormalization group. The work reports four negative results, falsifying most hypotheses. Notably, it reveals a prediction reversal phenomenon at small scales, confirms the absence of performance gains from rank-surface methods, and demonstrates that attention mechanisms exhibit critical or glassy characteristics rather than nearest-neighbor insulator behavior. All data have been made publicly available, providing crucial counterintuitive evidence and methodological references for interdisciplinary model compression.

0 citationsRead paper

A Modern Multimodal Assistant on a 6 GB 2011 GPU: Stage-Validated, All-GPU CUDA Inference for Fermi

Jul 16, 2026

This work presents the first end-to-end fully on-GPU inference of the modern multimodal large model MiniCPM-V-4.6 on a Fermi-architecture GPU (Tesla C2075) with only 6 GB of memory. By leveraging hand-written CUDA GEMM kernels, 8-bit weight dequantization, an sm_20-compatible SigLIP2 vision encoder, chunked delta-rule recursive computation, and a zero-overhead fused window attention mechanism, the approach overcomes severe memory and compute constraints. Key contributions include the first demonstration of full-GPU multimodal inference on Fermi hardware, the counterintuitive finding that 4-bit quantization degrades decoding speed, mitigation of platform-specific floating-point indexing nondeterminism, and significant performance gains: image encoding completes in just 0.93 seconds, and 10k-token prefill throughput increases 17-fold to 361 tokens per second, with visual forward-pass errors below 1.4e⁻⁵.

0 citationsRead paper

Power Studies For Two-Sample and Goodness-of-Fit Methods For Multivariate Data

May 12, 2026

This study addresses the lack of systematic evaluation of statistical power among multivariate two-sample and goodness-of-fit tests, which hinders the selection of effective methods in practice. Through extensive Monte Carlo simulations, it presents the first comprehensive comparison of numerous nonparametric tests under bivariate settings—encompassing both continuous and discrete data—as well as high-dimensional continuous scenarios. Based on empirical findings, the paper proposes a small yet complementary ensemble of methods that collectively ensure high power against a wide range of alternative hypotheses. This ensemble demonstrates strong robustness and broad coverage, significantly outperforming any single test and offering practitioners a reliable, principled recommendation for real-world applications.

0 citationsRead paper
Recent publications

Latest Papers

Pre-registered tests of solid-state-physics-inspired LLM compression: a cluster-level negative result at small-language-model scale

Sep 28, 2026

This study investigates whether compression methods inspired by solid-state physics can be effectively applied to small-scale language models. Employing rigorous research paradigms—including preregistration, three-sigma gating, and ensemble frameworks—the authors empirically evaluate five physical mapping hypotheses utilizing techniques such as Wannier functions, tight-binding models, DMRG truncation, and Wilsonian renormalization group. The work reports four negative results, falsifying most hypotheses. Notably, it reveals a prediction reversal phenomenon at small scales, confirms the absence of performance gains from rank-surface methods, and demonstrates that attention mechanisms exhibit critical or glassy characteristics rather than nearest-neighbor insulator behavior. All data have been made publicly available, providing crucial counterintuitive evidence and methodological references for interdisciplinary model compression.

0 citationsRead paper

A Modern Multimodal Assistant on a 6 GB 2011 GPU: Stage-Validated, All-GPU CUDA Inference for Fermi

Jul 16, 2026

This work presents the first end-to-end fully on-GPU inference of the modern multimodal large model MiniCPM-V-4.6 on a Fermi-architecture GPU (Tesla C2075) with only 6 GB of memory. By leveraging hand-written CUDA GEMM kernels, 8-bit weight dequantization, an sm_20-compatible SigLIP2 vision encoder, chunked delta-rule recursive computation, and a zero-overhead fused window attention mechanism, the approach overcomes severe memory and compute constraints. Key contributions include the first demonstration of full-GPU multimodal inference on Fermi hardware, the counterintuitive finding that 4-bit quantization degrades decoding speed, mitigation of platform-specific floating-point indexing nondeterminism, and significant performance gains: image encoding completes in just 0.93 seconds, and 10k-token prefill throughput increases 17-fold to 361 tokens per second, with visual forward-pass errors below 1.4e⁻⁵.

0 citationsRead paper

Power Studies For Two-Sample and Goodness-of-Fit Methods For Multivariate Data

May 12, 2026

This study addresses the lack of systematic evaluation of statistical power among multivariate two-sample and goodness-of-fit tests, which hinders the selection of effective methods in practice. Through extensive Monte Carlo simulations, it presents the first comprehensive comparison of numerous nonparametric tests under bivariate settings—encompassing both continuous and discrete data—as well as high-dimensional continuous scenarios. Based on empirical findings, the paper proposes a small yet complementary ensemble of methods that collectively ensure high power against a wide range of alternative hypotheses. This ensemble demonstrates strong robustness and broad coverage, significantly outperforming any single test and offering practitioners a reliable, principled recommendation for real-world applications.

0 citationsRead paper

A Software-Defined Radio Testbed for Distributed LiDAR Point Cloud Sharing with IEEE 802.11p in V2V Networks

Sep 17, 2025

Existing V2V networks suffer from a disconnect between simulation and real-world measurement, and LiDAR point cloud collaborative perception lacks validation under realistic vehicular communication conditions. Method: This paper develops a lightweight IEEE 802.11p testbed based on the ADALM-Pluto software-defined radio (SDR), featuring modular design, ROS-Docker integration for cross-node point cloud acquisition, transmission, and fusion, and novel incorporation of IPFS/Filecoin for decentralized point cloud storage and sharing. Contribution/Results: Experimental evaluation demonstrates robust channel quality, end-to-end latency <120 ms, storage convergence, and scalability to ≥5 cooperative vehicles. The platform bridges the gap between network simulation and physical-layer experimentation, and—crucially—provides the first SDR-level empirical validation of real-time, decentralized-storage-enabled LiDAR collaborative perception over 802.11p. It establishes a reproducible, low-cost experimental foundation for edge intelligence in V2X systems.

0 citationsRead paper