Score
Design, build, or analyze algorithms and tools that automatically deduce the static types of program expressions, function parameters, and module interfaces, producing inferred signatures and propagating type constraints across code. Use these inferred types to perform compatibility checks, detect mismatches, and generate or verify module and API type information.
This paper addresses the challenge of systematically detecting incorrect behaviors when functional programs interact with effectful library APIs. Methodologically, it introduces a trace-based type-theoretic framework that models API call sequences using symbolic regular expressions and performs compositional trace analysis via symbolic finite automata; it further designs a type inference algorithm for under-approximating erroneous behaviors of abstract data types (ADTs). Key contributions include: (i) the first trace-directed, systematic under-approximation technique for specifying incorrectness, (ii) support for modeling error propagation across function boundaries, and (iii) effective inference—within effectful library interaction contexts—of subsets of erroneous behaviors potentially triggered by ADT implementations. Empirical evaluation validates the feasibility and practicality of this compositional incorrectness analysis approach.
Traditional static analysis struggles to balance precision, reliability, and automation, limiting its practical applicability. This work proposes a novel parameterized static analysis approach that introduces user-provided local assumptions at selected program locations and incorporates them via a nondeterministic semantics, thereby constructing a mapping from sets of assumptions to analysis results. This formulation enables optimization-based search over large assumption spaces, overcoming conventional precision bottlenecks. The method’s effectiveness is demonstrated through experiments in two representative scenarios, significantly enhancing the adaptability and flexibility of static analysis in real-world applications.
To address the high manual annotation cost of pluggable type systems (e.g., NullAway) in legacy Java codebases, this paper proposes an automated type qualifier inference method. Our approach introduces NaP-AST—a lightweight program representation that explicitly encodes data-flow semantics as structural hints. We conduct the first systematic empirical comparison of graph transformation networks (GTNs), graph convolutional networks (GCNs), and large language models (LLMs) for this task, demonstrating that GTNs achieve superior performance. Evaluated on 12 open-source Java projects, our GTN-based method attains 0.89 recall and 0.60 precision, significantly reducing spurious type warnings. We further identify a performance inflection point at approximately 16K Java classes, beyond which model accuracy stabilizes. This work establishes a scalable, high-precision paradigm for static-analysis-driven type enhancement in industrial Java ecosystems.
Verifying coverage completeness of input generators in property-based testing remains challenging. Method: This paper proposes a static verification approach based on a “must-style” refinement type system, reformulating conventional “may-produce” type semantics into “must-produce” semantics. It formally defines full coverage for higher-order functions and inductive data types, enabling fully automated verification of generator completeness. Contribution/Results: To our knowledge, this is the first refinement type system provably guaranteeing generation of all inputs satisfying both type and constraint specifications. Experimental evaluation demonstrates substantial improvements in detecting coverage gaps across diverse complex generators, while significantly reducing manual verification effort.
Manual tuning of abstraction strategies in static program analysis is labor-intensive and struggles to balance precision and efficiency. Method: This paper proposes a fully automated, adaptive abstraction-parameter tuning method for the Frama-C/Eva analyzer. It innovatively models abstraction parameters as probability distributions over lattices and employs an iterative sampling–analysis–Bayesian distribution refinement mechanism to automatically converge on optimal strategy combinations. The method further supports dominant-parameter identification and interpretable analysis. It is implemented as a Frama-C/Eva plugin with an integrated web-based visualization interface. Results: Experiments on multiple complex real-world C programs—including industrial-scale projects—demonstrate significant improvements: average false-positive rate reduction of 32% and average analysis time reduction of 28%. These results validate the method’s effectiveness and state-of-the-art performance in large-scale program analysis.
This work addresses the significant runtime overhead commonly incurred by assertion checking in dynamically typed languages. It proposes a novel approach that, for the first time, systematically incorporates multi-calling-context information into a goal-directed, multi-variant abstract interpretation framework. By performing top-down inference of program properties under distinct calling contexts and selectively integrating the runtime semantics of assertions, the method substantially reduces redundant checks while preserving the ability to provide hints about unverified properties. An implementation in the Ciao system demonstrates that this technique markedly decreases the number of runtime checks and improves execution performance compared to existing approaches.
This work addresses the challenges of high-precision interprocedural static analysis in Python, which arise from its dynamic typing, dynamic dispatch, metaprogramming capabilities, and complex object model. To tackle these issues, we present PyFlow—the first general-purpose static analysis framework for Python based on the Interprocedural Finite Distributive Subset (IFDS) formulation. PyFlow leverages a multi-stage intermediate representation and parameterized abstract domains, enabling developers to specify only the data-flow semantics while automatically handling interprocedural hypergraph construction, fixed-point computation, and summary caching. Experimental evaluation demonstrates that PyFlow achieves the highest recall and F1 scores among nine state-of-the-art tools on both synthetic and real-world benchmarks, while maintaining precision comparable to advanced taint analysis engines—marking the first efficient and highly accurate application of IFDS to Python.
Traditional refinement type systems are difficult to adopt in mainstream languages due to their heavy annotation overhead, particularly when handling common properties such as integer ranges, which often require extensive manual annotations. This work proposes Ranger, a bidirectional type system for integer range refinements that integrates type inference with lightweight, flow-sensitive static analysis. Ranger supports imperative constructs—including variables and loops—while substantially reducing the annotation burden on users. Experimental evaluation using the Licorne language demonstrates that Ranger can concisely verify properties beyond the reach of standard type systems, such as index safety, and achieves greater annotation succinctness compared to both the Java Checker Framework and Liquid Java.
This work addresses runtime errors in Java applications caused by type mismatches between SQL and Java types in JDBC database access. By extending the Java compiler with the Checker Framework, the authors present the first static analysis technique capable of enforcing JDBC type safety across method boundaries. The approach requires no source code modifications and leverages optional annotations to enhance type inference, enabling static verification of Java type correctness during both PreparedStatement parameter setting and ResultSet value retrieval. A fallback checking mode is also provided for legacy systems. Experimental evaluation demonstrates that the technique effectively detects real-world type mismatch bugs, prevents runtime exceptions, and incurs acceptable compilation overhead, achieving a practical balance between soundness and usability.