runtime surrogate substitution

Design, build, or analyze runtime mechanisms that detect compromised or degraded components and replace them on the fly with compatible surrogate implementations (models, services, or lightweight substitutes) via hot‑swapping, state transfer, and adaptation layers so the system continues operating without restart. Also design the surrogate-based recovery and data‑protection procedures (for example, pseudonymization of sensitive inputs), compatibility checks, and graceful‑degradation policies that minimize task performance loss during and after replacement.

runtimesurrogatesubstitution

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.73
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of traditional memory corruption defenses, which often rely on system termination or reboot and thus fail to meet the stringent availability requirements of safety-critical cyber-physical systems (CPS). To enable continuous operation under attack, the authors propose Chameleon, a novel framework that, for the first time, leverages memory-safe machine learning (ML) agents to dynamically replace compromised components at fine-grained module granularity, achieving seamless behavioral recovery. Built on LLVM, Chameleon integrates ML-based agent modeling, memory-safety isolation, and real-time attack response mechanisms. Evaluations on seven robotic vehicles demonstrate that the ML agents achieve an average R² of 0.96, effectively mitigating real-world memory corruption attacks while significantly outperforming existing approaches in task completion rates with low runtime overhead.

Attack ResilienceCyber-Physical SystemsMemory Corruption Attacks

This work addresses the high cost of ground-truth evaluation in chemical and materials design, where existing machine learning surrogate models often lack reliability guarantees. Departing from conventional reliance on prediction accuracy metrics such as R²—which can paradoxically increase the risk of worst-case selections—the study proposes “rank preservation” as a core criterion for surrogate validation. It formally introduces the concept of “selection tax” and derives its theoretical upper and lower bounds. A safety certification framework for surrogates is established through selection-aware auditing, rank correlation analysis, and multi-task ground-truth validation. Experiments demonstrate that the proposed audit statistics achieve Spearman correlations of 0.80–0.99 with actual search performance, substantially outperforming R² (as low as 0.33). Certified screening strategies based on this framework reduce evaluation costs by up to 25-fold.

experimental replacementmodel validationselection bias

This study investigates whether code agents can repair failing software while preserving underlying logic and rebuild industrial engines from open specifications, addressing the measurement error challenges inherent in executable verification. To this end, this work proposes ReviveBench, a benchmark that establishes an evaluation framework incorporating hidden validators, native environment calibration, identifier obfuscation, and consistency analysis. Furthermore, contamination control experiments and auditing mechanisms are introduced to expose measurement errors, alongside three checking strategies designed to calibrate evaluation standards. Empirical results demonstrate that the strongest models achieve full success on revival tasks, while several models meet reconstruction metrics. Additionally, the proposed framework identifies and rectifies 28 defects within existing validators, thereby enhancing the reliability of automated code generation evaluation.

benchmark evaluationcoding agentsengine reconstruction

This study addresses the issues of residual stale outputs and loss of valid state during task revision in LLM agents by proposing a versioned execution control plane. The proposed method transforms abort-and-restart operations into seamless version transitions, decouples resource scheduling from release permissions, and unifies fast-fail with selective retention mechanisms to support the inheritance of completed KV cache states. A vLLM-based system implementation encompasses GPU execution, hierarchical recovery, and multi-tenant serving. Experimental results demonstrate that this approach reduces the median time-to-first-token latency by 17.1% while eliminating obsolete outputs and effectively preserving execution progress across repeated revisions.

KV state inheritanceLLM agentsserving system

Latest Papers

What's happening recently
View more

This study addresses the absence of real-world repository benchmarks and execution security risks in using large language models (LLMs) to repair runtime errors. It introduces HealBench, the first benchmark for runtime crash repair over real codebases, alongside HealGuard, a trusted repair protection framework based on taint analysis. By integrating static and dynamic taint analysis with HealCore—a restricted Python subset—the proposed approach constructs an automated repair agent. Experimental results demonstrate that under optimal configurations, the method achieves a 38.11% program execution recovery rate and a 28.68% test pass rate while successfully identifying 17.4% of potentially unsafe repairs. These findings validate both the effectiveness and security of LLM-driven automated crash repair in real-world scenarios.

large language modelsreal-world repositoriesruntime error healing

This study addresses the vulnerability of LLM agents to task failures and unsafe operations caused by the reuse of invalid persistent states. To mitigate this, we propose a pre-action state diagnosis and repair framework that employs record-level counterfactual replanning to localize critical corrupted records, combined with typed evidence binding and reliability checks to repair states while maintaining lineage. Furthermore, independent verification is enforced prior to execution through read-only fact validation and intent clarification mechanisms. This approach enables continuous correction and decoupled state-action verification. Experimental results demonstrate that our method achieves 93.3% accuracy under corrupted states—substantially outperforming the 38.7% baseline—while ensuring zero unsafe actions and exhibiting strong cross-environment transferability.

coding agentsinvalid recordsLLM agents

This study addresses the challenges of insufficient test coverage, misleading textual similarity, and non-executable environments in code equivalence determination by proposing the FEAgent framework. This approach integrates typed program graph alignment with differential agent execution, employing branch-aware input generation and double-blind LLM prediction of observable behaviors to produce an evidence ledger with explicit uncertainty, thereby enabling auditable equivalence judgments that effectively bridge the gap between testing and formal verification. Evaluated on EquiBench and SWE-bench, the framework successfully identifies 216 mislabeled benchmark pairs and 94 defective patches, revealing behavioral divergences overlooked by existing unit tests.

code equivalencefunctional equivalencepatch validation

This study addresses the vulnerability of distributed malware detection systems to endpoint misclassification caused by server failures. To mitigate this, we construct a Promela model based on Bitdefender’s production architecture and formally verify its graceful degradation fallback chain mechanism using the SPIN model checker combined with Linear Temporal Logic (LTL). This work presents the first formal verification of the fault-handling layer within a production-grade security system, thereby bridging a significant research gap in the field. Experimental results demonstrate that the system is free from deadlocks and false positives while guaranteeing unique verdicts, ensuring that degradation behaviors remain orderly and controllable under failure conditions.

distributed malware detectionendpoint securityfault tolerance

This work addresses the fragmentation of methods, inconsistent evaluation standards, and insufficient coverage of embedded architectures in LLM-assisted source code recovery by presenting the first Systematization of Knowledge (SoK) study in this domain. We construct a fine-grained taxonomy and a controlled recovery pipeline, combined with multi-dimensional ablation studies, to conduct a comprehensive empirical evaluation on 45,000 samples spanning multiple architectures, optimization levels, and programming languages. Our findings elucidate the mechanisms through which critical design choices influence recovery performance. This study provides essential guidance for standardized evaluation practices and future research directions in the field.

Binary-to-SourceEmbedded ArchitecturesEvaluation Benchmark

Hot Scholars

ZT

Zhao Tan

Facebook
Optimization and Deep Learning
MS

Manon Stipulanti

ULiège
Combinatorics on wordsdiscrete mathematicsautomata theoryformal languages theory
JZ

Jingbo Zhu

Northeastern University, China
Machine TranslationLanguage ParsingNatural Language Processing