Score
Designs, implements, and analyzes programs, libraries, and systems that allow multiple computations to execute simultaneously or overlap in time, covering synchronization, communication, scheduling, and coordination among threads, processes, or tasks. This includes creating and verifying concurrent data structures and algorithms, choosing and implementing synchronization primitives and concurrency control, and diagnosing or preventing correctness (race conditions, deadlocks) and performance issues (scalability, contention).
Concurrent programming faces a fundamental tension between expressiveness and determinism; conventional shared-memory models suffer from schedule-dependent behavior and non-reproducible outputs due to destructive updates. Method: This paper introduces Clock-Synchronized Memory (CSM), the first shared-memory abstraction that guarantees deterministic semantics for general-purpose concurrent programs—extending beyond the restricted primitives (e.g., registers, signals) supported in traditional synchronous programming (SP). Grounded in the formal mathematical semantics of SP, we design CSM memory primitives, a clock-synchronization protocol, and rigorously defined access rules, ensuring full compatibility with existing SP compilation and verification toolchains. Results: Experiments demonstrate that CSM significantly enhances expressiveness, modularity, and code reusability of concurrent programs while preserving formal verifiability. It provides a theoretically sound and practically deployable deterministic concurrency infrastructure for safety-critical systems.
Static synchronization mechanisms in distributed systems impose severe scalability bottlenecks by enforcing strong consistency even in the absence of actual conflicts. This work introduces the “dynamic concurrency” paradigm—the first approach to perform fine-grained, runtime state–aware conflict detection: synchronization is triggered only when concurrent operations induce genuine dependency conflicts under the current data state. Methodologically, we design a state-aware, generic conflict predicate that integrates dynamic conflict detection with lightweight synchronization arbitration. Experimental evaluation shows that our approach significantly reduces redundant synchronization overhead, achieving 32%–68% higher throughput and 41% lower latency under typical distributed workloads, while preserving linearizability. The core contribution lies in elevating conflict detection from static, operation-level reasoning to dynamic, state-level reasoning—establishing a novel, efficient foundation for concurrency control in high-concurrency distributed systems.
To address the low execution efficiency of guard-based synchronization in shared-variable concurrent models, this paper proposes an efficient guard-atomic-action synchronization mechanism for object-oriented languages, wherein guard logic is deeply bound to objects to enable condition-driven atomic execution regions. Methodologically, it introduces the first integration of coroutine scheduling, OS thread pooling, object-granularity customized queue/stack memory management, dynamic guard-condition evaluation, and lazy wakeup. Its core contribution lies in overcoming the traditional loose coupling between guard synchronization and object models, achieving substantial reduction in synchronization overhead through synergistic runtime and language-semantic optimizations. Evaluation on the Lime experimental language demonstrates that the mechanism outperforms mainstream concurrent platforms—including C/Pthreads, Go, Erlang, Java, and Haskell—on synthetic benchmarks.
Existing user-space coroutine/fiber synchronization mechanisms implicitly assume kernel scheduling, introducing unnecessary latency on critical paths and limiting high-concurrency throughput. This paper proposes Combine-and-Exchange Scheduling (CES), a novel synchronization paradigm for purely user-space cooperative scheduling. CES eliminates cross-thread overhead by retaining critical sections on the same thread during lock contention, while dynamically redistributing parallelizable tasks to idle threads. Crucially, it co-designs user-space synchronization primitives with the scheduler to fully bypass kernel intervention. Experimental evaluation demonstrates that CES achieves up to 3× higher throughput on application-level benchmarks and up to 8× speedup on microbenchmarks—significantly outperforming state-of-the-art user-space synchronization approaches.
Communication races in distributed Go programs—caused by violations of the happens-before relation—lead to erroneous message reception, including premature, missing, or partial message delivery. Method: We propose the first static verification framework for a Go subset that extends the happens-before ordering to both buffered and unbuffered channels, integrating formal channel semantics with abstractions of distributed execution traces. Contribution/Results: Our approach enables precise, sound static detection of communication races and fully verifies communication-race-freedom for representative message-driven distributed Go programs. By enforcing correct message ordering at compile time, it eliminates subtle runtime errors stemming from message reordering, thereby significantly enhancing the reliability of distributed systems. The framework is built upon rigorous modeling of Go’s concurrency primitives and supports automated, end-to-end verification without requiring program instrumentation or runtime monitoring.
This work addresses the problem of determining whether barrier placements in a control flow graph guarantee synchronization of all threads across every execution. To this end, it introduces the first formal definitions of convergent nodes, convergent edges, and well-synchronized programs, along with a novel set of branch-and-merge inference rules. By integrating path-sensitive information with thread divergence analysis, the paper presents a bidirectional, linear-time worklist algorithm that efficiently performs static analysis of synchronization regions. The proposed method significantly enhances the compiler’s ability to optimize barriers on warp-synchronous hardware, achieving both theoretical rigor and practical efficiency.
This work proposes λpitchfork, the first functional choreographic language supporting dynamic process spawning. Traditional concurrent programming requires separate programs for each participant, while existing choreographic approaches struggle to accommodate dynamic process creation. In contrast, λpitchfork enables runtime decisions about when, how, and with whom new processes are spawned, all while guaranteeing deadlock freedom. The language automatically generates distributed endpoint programs from a centralized choreographic specification. By integrating dynamic process spawning into functional choreographic programming, this approach rigorously combines theoretical soundness with practical expressiveness, effectively capturing real-world concurrency patterns such as load balancing and parallel divide-and-conquer. The correctness and practicality of λpitchfork are substantiated through formal verification and illustrative case studies.
This work addresses the challenges students face in understanding and debugging nondeterministic concurrency bugs—such as deadlocks and race conditions—when learning parallel programming. The authors propose ParaView, an educational tool that integrates execution trace visualization with large language model (LLM) analysis, uniquely combining program execution logs, visual representations of parallel behavior, and LLM-driven error explanations and repair suggestions for concurrent programming instruction. In an evaluation with 17 students, the use of ParaView led to significantly higher success rates in both debugging and implementation tasks. Most participants reported that ParaView effectively supported their learning, and the LLM accurately identified common concurrency errors and interpreted execution traces, though its repair suggestions remained limited in complex synchronization scenarios.