Score
Designs and evaluates protocols, plans, and systems for acquiring empirical data — including surveys, experimental and field studies, controlled and targeted sampling, mobile/web/web‑scale, video, and wearable capture — by specifying sampling targets, collection procedures, and large‑scale acquisition pipelines. Builds and analyzes sampling and data‑acquisition strategies that set inclusion/exclusion criteria, targeted (donor) sampling schemes, processing workflows, and resource/compensation trade‑offs to meet desired data quality, volume, and ethical or operational constraints.
Existing mobile sensing data collection methods neglect subjective feedback (e.g., questionnaires, self-reports), leading to fragmented contextual understanding and inaccurate behavioral modeling. To address this, we propose a human-in-the-loop collaborative sensing framework built upon the iLog platform. Our approach introduces three key innovations: (1) a dual-dimensional “context–time” modeling paradigm; (2) a calendar-style real-time monitoring dashboard; and (3) a dynamic acquisition plan revision mechanism. Leveraging context-aware modeling, real-time visual analytics, an adaptive experimental workflow engine, and purposeful human–system interaction design, the framework enhances controllability for researchers, participants, and the system itself. Evaluated with 350 university students, our method significantly improves semantic richness, contextual completeness, and overall data quality—enabling more accurate behavioral modeling and fine-grained personalized analysis.
本文评估了网络安全性测量中的采样策略,指出概率抽样比Top N和混合抽样更能提供稳定无偏的估计结果。
Traditional structured surveys often fail to elicit deep, nuanced qualitative insights. Method: This work introduces a hybrid qualitative data collection paradigm by embedding theory-driven interview probes—descriptive, individual, clarifying, and explanatory—into an LLM-powered chatbot. We conduct the first systematic, three-phase (exploratory, requirements elicitation, evaluation) HCI study comparing probe efficacy using a split-plot experimental design and a multidimensional qualitative response quality framework. Contribution/Results: Probes significantly enhance response depth and informational richness; descriptive probes perform best in exploration, while explanatory probes excel in evaluation. User acceptance is high across phases. This work advances the theoretically grounded application of LLMs in human-centered research and provides a reusable methodological foundation for scalable, high-fidelity qualitative data collection.
Large smart-farming deployments generate continuous scientific data from spatially distributed sensors, including soil, humidity, temperature, crop-health, and pest-related measurements. In vast agricultural fields, however, an energy-constrained unmanned aerial vehicle (UAV) often cannot collect data from every sensor during each mission. Existing UAV-assisted collection methods typically optimize coverage, route length, data volume, or freshness, but they do not always distinguish between data that is merely available and data that is scientifically valuable. This poster introduces a utility-driven spatial sampling framework for UAV-assisted smart farming. The field is partitioned into grid cells, each sized according to the UAV ground coverage range. After an initial exploration phase, each cell receives a scientific utility score based on freshness, redundancy, anomaly likelihood, and model uncertainty. The UAV then selects and visits a subset of high-utility cells under battery and return-to-base constraints. The proposed framework reframes UAV-based collection as adaptive scientific data management rather than exhaustive sensing.
Academic researchers face significant challenges in collecting mobile screen data—including limited access due to proprietary platform restrictions, stringent commercial monopolies, and heightened privacy compliance requirements. Existing open-source frameworks predominantly focus on sensor data and lack robust, privacy-compliant, and flexible mechanisms for capturing screen content. Method: We propose Crepe, the first no-code Android screen data collection tool designed specifically for academic research. It introduces a novel graph-query-based UI structural representation to enable semantic identification and high-precision localization of screen elements. Crepe integrates declarative demonstration learning, on-device processing, and a permission sandbox to ensure informed consent and real-time user opt-out. Contribution/Results: Empirical evaluation across diverse applications demonstrates that Crepe achieves zero-configuration extraction of dynamic text and UI controls with high accuracy, effectively circumventing data monopolies while enabling privacy-preserving screen-content research.
本文介绍了CANDOR,一种用于参与者主导的敏感数字轨迹数据收集与治理的基础设施,解决了平台API访问限制和隐私等问题。
This work addresses the lack of quantitative evaluation in existing methods regarding how generated data affects downstream model performance, which hinders reliable synthetic data quality assurance. The authors propose a model-aware synthetic data generation framework that, for the first time, leverages acquisition functions from active learning as interpretable, model-centric reward signals. Integrating reinforcement learning–based generation, a rejection-sampling alternative strategy, and generalization techniques across models and resource scales, the framework guides language models to produce data with higher information content and greater task impact. Experiments on mathematical reasoning, medical question answering, and code generation demonstrate that student models trained on the generated data achieve performance gains of 2–7% and exhibit significantly improved robustness against catastrophic forgetting.
This study addresses the challenge of inefficient drill-hole sampling for ore grade estimation in geologically complex regions by proposing a novel information value assessment method that does not require conditional simulation. Built upon Gaussian processes and Bayesian posterior predictive distributions, the approach incorporates structural similarity metrics to enable adaptive spatial sampling within heterogeneous or partitioned geological domains—without relying on assumptions of stationarity or error independence. The strategy effectively targets areas of high spatial uncertainty and supports cost–benefit trade-offs and short- to long-term exploration decisions even in the absence of ground-truth references. Experimental results demonstrate that, compared to regular grid sampling, the proposed method substantially reduces uncertainty in grade estimation, thereby establishing a foundational methodology for future embodied intelligent robotic collaborative exploration.
This study addresses the high overhead and low efficiency of full-collection distributed tracing in microservices by proposing RADAR, an intelligent agent that dynamically optimizes sampling strategies within OpenTelemetry and Kubernetes environments. Its core innovation lies in pioneering the use of data information entropy as a reinforcement learning reward signal, enabling the agent to autonomously explore optimal sampling rules that balance resource consumption with observability quality. Experimental results demonstrate that this approach reduces network bandwidth by 97.4% and CPU overhead by 99.0%, while preserving 85.6% of rare trace patterns and increasing the average entropy of stored information by 25%.
This study presents the first systematic comparison of river sampling and snowball sampling in online surveys conducted in West Africa. Initial respondents were recruited via Facebook geotargeted advertisements (river sample), and subsequent participants were tracked through their social sharing (snowball sample). The analysis quantifies differences between the two samples in survey completion rates, demographic composition, and response behaviors. Findings indicate that the snowball sample exhibited higher completion rates and significantly greater representation of women and new users, though responses were shorter and completed more quickly. The results highlight snowball sampling’s distinct advantage in enhancing participation among marginalized groups, offering empirical evidence and methodological guidance for designing social media–assisted online surveys.