Score
Designs and carries out empirical evaluations of workplace monitoring systems by specifying test criteria, metrics, and measurement procedures and by building or using testbeds to collect system logs, sensor outputs, and user feedback. Analyzes technical performance (accuracy, reliability, security), operational effects (usability, worker behavior, productivity), and social/legal impacts (privacy, fairness, compliance), identifying bias, validity issues, and unintended consequences.
This study addresses the well-being–efficiency trade-offs and ethical risks arising from passive sensing technologies—such as wearables, environmental sensors, computer vision, and multimodal behavioral modeling—in workplace settings. It systematically reviews empirical evidence on their impacts on employee mental/physical health and work performance. Adopting a novel interdisciplinary synthesis of existing studies, the work proposes a human-centered evaluation framework and a phased ethical implementation roadmap. The analysis clarifies the boundary conditions of technological efficacy and identifies six critical challenges: privacy erosion, algorithmic bias, inadequate intervention personalization, among others. Findings provide a theoretical foundation, design principles, and actionable pathways for developing next-generation intelligent workplace systems that uphold autonomy and human dignity—shifting the paradigm from surveillance-based management toward human potential augmentation. (149 words)
研究探讨了工作场所监控在检测内部威胁时的合法性界限与隐私损害问题,提出改进透明度、培训及调整工具以更准确地识别相关行为和心理指标。
论文讨论了使用合规数据进行AI系统评估的问题,提出需明确能力、结果、暴露机会等要素的一致性以确保评估有效性。
This work addresses the problem of implementation drift in evolving distributed systems, where runtime behavior gradually deviates from the original design. To tackle this issue, the paper proposes a design conformance assessment method based on distributed tracing data. It introduces, for the first time in the domain of distributed systems, conformance checking techniques from process mining, leveraging runtime traces collected via the OpenTelemetry standard and automatically comparing them against behavioral models defined at design time to quantify their alignment. The key contribution lies in establishing persistent, monitorable conformance metrics that enable continuous, automated evaluation of deviations between system implementation and design. This approach is readily applicable to modern distributed systems widely adopting OpenTelemetry for observability.
Existing RTI policy monitoring tools suffer from inefficient information acquisition and inadequate dynamic tracking capabilities. This study proposes a methodology for developing an open-source web-based system dedicated to Research, Technology, and Innovation (RTI) policy monitoring. It employs role-driven requirements engineering to precisely elicit heterogeneous stakeholder needs and introduces a novel modular architectural paradigm centered on a user-configurable dashboard, with strict separation across presentation, service, and data layers. The approach integrates open-data interoperability standards and interactive visualization techniques. As key contributions, the work delivers a reusable RTI monitoring system architecture specification and a standardized dashboard requirements template. These artifacts were empirically validated through deployment in the Austrian RTI Monitor—a national-scale platform enabling cross-departmental, real-time indicator tracking and evidence-informed policy coordination—thereby substantially enhancing the timeliness, accessibility, and scalability of RTI policy monitoring.
This work addresses the lack of systematic evaluation benchmarks for large language models (LLMs) in security audit log investigation tasks by introducing AuditBench, the first audit log benchmark specifically designed for attack investigation. AuditBench encompasses over 50 real-world scenarios across Linux and Windows systems and focuses on four core tasks: alert classification, persistence mechanism identification, among others. Through multidimensional experiments, the study systematically evaluates the impact of model scale, log representation, prompt design, and fine-tuning strategies on performance and error patterns, while also analyzing the quality of LLM-generated explanations. The findings reveal the capability boundaries and characteristic failure modes of various models across different investigative tasks, providing empirical foundations for deploying and optimizing LLMs in security operations.
Large language models (LLMs) can still generate unsafe outputs in deployment, necessitating efficient real-time monitoring. This work proposes a lightweight online safety monitoring mechanism that integrates signals from an external verification model, threshold-based decision rules, and risk control theory to produce reliable alerts through calibrated thresholds. The approach features a simple architecture that avoids computationally intensive procedures yet achieves detection performance on par with state-of-the-art sequential hypothesis testing methods across mathematical reasoning and red-teaming benchmarks. By combining practical efficiency with theoretical guarantees, the proposed method offers a viable solution for real-world LLM safety monitoring.
This work addresses the lack of systematic evaluation benchmarks for monitoring the behavior of large language models (LLMs), which hinders reliable assessment of their detection capabilities and false alarm rates across diverse tasks and failure modes. We propose AutoMonitor-Bench, the first comprehensive benchmark encompassing question answering, code generation, and reasoning tasks, featuring 3,010 annotated pairs of违规 and benign samples. We introduce a dual-metric evaluation framework based on Miss Rate (MR) and False Alarm Rate (FAR). Fine-tuning Qwen3-4B-Instruction on 153,581 training samples and evaluating across 22 mainstream LLMs reveals a pervasive MR–FAR trade-off. While fine-tuning improves recognition of known violations, it struggles to generalize to implicit, unseen violations, highlighting the inherent tension between safety and utility in building reliable LLM monitors.
This study addresses interpretive discrepancies between monitored individuals and supervising authorities in electronic monitoring systems, where divergent standpoints lead to misjudgments of behavior and imbalanced interactions. Drawing on China’s community correction system, the research employs semi-structured interviews (with 26 supervisees and 12 supervisors), situational analysis, and a CSCW theoretical framework to uncover structural misalignments in data interpretation. Introducing the concept of “interpretive misalignment,” the work reconceptualizes continuous sensing as distributed interpretive labor and identifies five categories of behavioral responses stemming from asymmetries in data, context, and inference. Building on these findings, the study proposes design directions that enhance transparency and mutual negotiability in data-driven decision-making, offering novel perspectives on intelligibility, contestability, and accountability across system boundaries.
Detecting runtime control-flow anomalies in complex systems remains challenging due to “unknown unknowns”—unforeseen deviations beyond predefined specifications. Method: This paper proposes a software monitoring approach integrating large language models (LLMs) with conformance checking. The method leverages LLMs to automatically align design models with source code, generate semantically consistent instrumentation strategies, and construct interpretable, lightweight control-flow models from event logs. Contribution/Results: It is the first work to introduce LLM-driven design-code co-modeling into dynamic monitoring—replacing manual rule specification with end-to-end automated monitor synthesis. Evaluated on a railway traffic management case study, the approach achieves 84.78% control-flow coverage, 96.61% F1-score, and 93.52% AUC for anomaly detection, significantly enhancing system reliability and trustworthiness under unknown environmental conditions.
This study addresses the challenge of effectively monitoring early-stage agent systems, where structural flaws often obscure task-level errors. The authors propose a three-dimensional (quality, suitability, efficiency) and three-granularity (intra-run, inter-run, structural) monitoring and triaging framework tailored for low-maturity agent systems. They introduce a novel system maturity staging model based on the coefficient of variation and monitoring granularity, integrated with a severity classification adapted from FMEA to guide human review. The resulting transferable monitoring architecture supports document-driven, multi-stage workflows, enhanced by a synthetic testbed with controlled error injection. Experimental results demonstrate that structural defects significantly mask task-level signals; 97% of issues can be automatically traced, with only 2% requiring human intervention, and each granularity level precisely identifies its corresponding defect type (coefficients of variation: 0.02, 1.25, and 0.00, respectively).