An Explicit Ordinal Bound for System T Dialogue Trees
该研究通过计算自然数K(t)并证明h(D(t))<θ_K(t),直接证明了System T对话树的经典序数高度低于ε_0。
该研究通过计算自然数K(t)并证明h(D(t))<θ_K(t),直接证明了System T对话树的经典序数高度低于ε_0。
This work addresses the challenges of factual inconsistency, redundancy, and weak length control in multi-document abstractive summarization by proposing a modular generate-then-select framework. Under a sentence-level budget constraint, the approach first employs a pretrained model to generate candidate summaries and decomposes them into sentence units; it then applies a combinatorial optimization selector to jointly maximize relevance, factual faithfulness, and diversity while minimizing redundancy. By decoupling generation from selection, the framework flexibly integrates various selection strategies—such as MMR, ILP, and DPP-inspired objectives—without requiring fine-tuning of the generator. Experimental results demonstrate significant improvements in factual consistency and source fidelity across CNN/DailyMail, Multi-News, FaithBench, and TofuEval benchmarks. Human evaluations further confirm consistent superiority over baselines in coherence, relevance, clarity, and conciseness, with only a marginal drop in ROUGE scores.
This work addresses the challenge of deploying lightweight models for agricultural weed detection on resource-constrained edge devices, where performance gains through conventional model scaling are infeasible. To overcome this limitation, the authors propose a Model-Driven Data Correction (MDDC) framework that operates with a fixed YOLOv8n architecture. By conducting automated error analysis, detection failures are systematically categorized into four distinct types. This classification informs a structured iterative pipeline of training, data correction, and retraining, augmented by versioned data management to ensure traceable and systematic improvements in data quality. Notably, MDDC represents the first approach to integrate error-type taxonomy with iterative data refinement, effectively breaking through performance bottlenecks imposed by model constraints. Experiments across multiple weed datasets demonstrate consistent improvements, achieving mAP@0.5 gains of 5% to 25%.
Adversarial attacks leveraging hybrid real-and-synthetic speech pose a novel threat to speaker verification, exposing the failure of conventional binary spoofing detection paradigms in mixed-utterance scenarios. Method: We introduce HSAD—the first benchmark dataset for hybrid speech anti-spoofing—and identify critical overgeneralization and miscalibration issues in existing models. To address these, we propose a data-level adaptation strategy and a fine-grained, multi-class evaluation framework that jointly leverages spectrogram-based encoding and self-supervised waveform representations (MIT-AST, Wav2Vec 2.0, HuBERT) for precise authenticity discrimination. Contribution/Results: Our approach achieves 97.3% accuracy and 98.9% F1-score on HSAD, demonstrating the efficacy of dataset-specific adaptation. This work establishes a new paradigm and foundational infrastructure for robust audio anti-fraud systems.
Traditional course advising tools struggle to integrate individual student profiles—such as transcripts and major curricula—resulting in limited personalization. This paper proposes a transcript-aware intelligent course advisor system tailored for undergraduate computer science students, integrating academic records with degree requirements and supporting both CLI and web interfaces for course recommendation, grade tracking, and four-year academic planning. We introduce a context-aware mechanism and novel evaluation metrics—including PlanScore, PersonalScore, Lift, and Recall—to quantify recommendation quality. The system employs a locally deployed large language model (via Ollama) as its AI engine, tightly coupled with the institutional course management system. Evaluation across 25 representative advising scenarios demonstrates that full contextual integration significantly improves recommendation relevance, validating its critical value for academic advising.
该研究通过计算自然数K(t)并证明h(D(t))<θ_K(t),直接证明了System T对话树的经典序数高度低于ε_0。
This work addresses the challenges of factual inconsistency, redundancy, and weak length control in multi-document abstractive summarization by proposing a modular generate-then-select framework. Under a sentence-level budget constraint, the approach first employs a pretrained model to generate candidate summaries and decomposes them into sentence units; it then applies a combinatorial optimization selector to jointly maximize relevance, factual faithfulness, and diversity while minimizing redundancy. By decoupling generation from selection, the framework flexibly integrates various selection strategies—such as MMR, ILP, and DPP-inspired objectives—without requiring fine-tuning of the generator. Experimental results demonstrate significant improvements in factual consistency and source fidelity across CNN/DailyMail, Multi-News, FaithBench, and TofuEval benchmarks. Human evaluations further confirm consistent superiority over baselines in coherence, relevance, clarity, and conciseness, with only a marginal drop in ROUGE scores.
This work addresses the challenge of deploying lightweight models for agricultural weed detection on resource-constrained edge devices, where performance gains through conventional model scaling are infeasible. To overcome this limitation, the authors propose a Model-Driven Data Correction (MDDC) framework that operates with a fixed YOLOv8n architecture. By conducting automated error analysis, detection failures are systematically categorized into four distinct types. This classification informs a structured iterative pipeline of training, data correction, and retraining, augmented by versioned data management to ensure traceable and systematic improvements in data quality. Notably, MDDC represents the first approach to integrate error-type taxonomy with iterative data refinement, effectively breaking through performance bottlenecks imposed by model constraints. Experiments across multiple weed datasets demonstrate consistent improvements, achieving mAP@0.5 gains of 5% to 25%.
Adversarial attacks leveraging hybrid real-and-synthetic speech pose a novel threat to speaker verification, exposing the failure of conventional binary spoofing detection paradigms in mixed-utterance scenarios. Method: We introduce HSAD—the first benchmark dataset for hybrid speech anti-spoofing—and identify critical overgeneralization and miscalibration issues in existing models. To address these, we propose a data-level adaptation strategy and a fine-grained, multi-class evaluation framework that jointly leverages spectrogram-based encoding and self-supervised waveform representations (MIT-AST, Wav2Vec 2.0, HuBERT) for precise authenticity discrimination. Contribution/Results: Our approach achieves 97.3% accuracy and 98.9% F1-score on HSAD, demonstrating the efficacy of dataset-specific adaptation. This work establishes a new paradigm and foundational infrastructure for robust audio anti-fraud systems.
Traditional course advising tools struggle to integrate individual student profiles—such as transcripts and major curricula—resulting in limited personalization. This paper proposes a transcript-aware intelligent course advisor system tailored for undergraduate computer science students, integrating academic records with degree requirements and supporting both CLI and web interfaces for course recommendation, grade tracking, and four-year academic planning. We introduce a context-aware mechanism and novel evaluation metrics—including PlanScore, PersonalScore, Lift, and Recall—to quantify recommendation quality. The system employs a locally deployed large language model (via Ollama) as its AI engine, tightly coupled with the institutional course management system. Evaluation across 25 representative advising scenarios demonstrates that full contextual integration significantly improves recommendation relevance, validating its critical value for academic advising.