🤖 AI Summary
This study addresses the challenge that opaque pointers cause type information loss, hindering indirect call analysis in matching function pointers with structure fields. We propose Facet, the first approach to reconstruct dispatching relationships on opaque intermediate representations without requiring end-to-end value-flow paths for correlating field identities. By decoupling field identification from function assignment, Facet optimizes call graph construction through a combination of static analysis, LLVM IR processing, and neuro-symbolic reasoning assisted by large language models. Experimental evaluations demonstrate that Facet reduces the target set size to 5.2 while achieving a recall of 0.99. Furthermore, it successfully uncovers 17 deep-seated vulnerabilities, including three that had remained latent for over a decade.
📝 Abstract
Resolving indirect calls is central to call-graph construction for C. Scalable type-based analyses such as MLTA use type information in LLVM IR to associate indirect calls with functions assigned to the corresponding structure fields. However, a single pointee type often misrepresents the memory a pointer addresses, and LLVM 17 removed pointee types in favor of opaque pointers. Therefore, field-sensitive analyses lose their matching key. Recovering the erased types restores the matching key but still misses the relation that the type encoded: which functions the program assigns to the field. We present Facet, to our knowledge the first analysis that reconstructs this dispatch relation over opaque IR. Facet identifies the structure field from which an indirect call loads its function pointer. It separately recovers the functions assigned to that field through initializers, stores, and aggregate copies. It then joins the two by field identity, without requiring an end-to-end value-flow path. Facet classifies proposed call-graph changes under distinct evidence rules for edge addition and removal and records the assumption behind each refinement. An LLM decides only the residual cases among symbolically bounded candidates. One analysis yields both a recall-preserving call graph and a refined call graph. On 14 C programs, Facet reduces the mean target-set size from 25.9 to 5.2 and raises observed recall from 0.79 to 0.99. Its recovered field identities agree with typed IR at 98.1% of jointly resolved sites. Applied to bug detection, the refined call graph found 17 deep bugs in C software from nginx to the Linux kernel, three of them latent for over a decade; 12 are confirmed.