Score
Design and implement algorithms, data representations, and software or hardware routines for algebraic structures and computations over finite fields and related integer domains, including finite-field arithmetic, construction of algebraic field structures and linear maps, and solving polynomial equations over fields. Build and optimize multiword/big‑integer and multiprecision operations — including parallel and GPU‑accelerated implementations, register layouts, type conversions, and operation‑count optimizations — to enable efficient finite‑field and multiprecision computations.
Verifying large-scale arithmetic circuits for wide-word operations often incurs prohibitive computational costs due to reliance on arbitrary-precision integer arithmetic, which scales poorly with word length. This work proposes a hybrid algebraic verification approach based on polynomial reasoning that integrates both linear and nonlinear rewriting strategies. Crucially, it introduces— for the first time—a parallel multimodal homomorphic image technique that performs algebraic reasoning simultaneously over multiple prime moduli, thereby entirely eliminating the need for large-integer computations. Implemented in the TalisMan2.0 tool, the method demonstrates significant performance advantages over existing verification schemes on multiplier benchmarks, offering both high efficiency and strong scalability.
This work addresses the time–space trade-off for polynomial computation under strict space constraints, overcoming the linear-space bottleneck of classical fast algorithms (e.g., FFT). Methodologically, it integrates fine-grained space complexity analysis, divide-and-conquer paradigms, sparse polynomial representations, and algebraic coding techniques. Key contributions include: (i) the first quasi-linear-time algorithm for sparse polynomial interpolation—paralleling FFT’s efficiency for dense polynomials—and a refined theoretical model for space complexity; (ii) sublinear-space implementations of fundamental operations—including multiplication, division, interpolation, and factorization—while preserving near-optimal time complexity. These advances enable more efficient polynomial arithmetic in memory-constrained environments, with direct implications for cryptography, error-correcting codes, and other domains reliant on structured polynomial computations.
This work addresses the lack of efficient support for high-precision integer division in the range of $2^{15}$ to $2^{18}$ bits on general-purpose GPUs. We propose a fully integer-based Newton–Raphson division algorithm that leverages shift-based inversion, prefix sums, and multi-precision multiplication to enable data-parallel acceleration on CUDA platforms, achieving the first efficient implementation for this precision range. A performance cost model centered on the number of multiplications guides our algorithmic optimizations. Experimental results demonstrate that the proposed method attains near-theoretical-optimal performance within the target precision and significantly outperforms the state-of-the-art CGBN library in handling medium-sized large integer divisions.
This work addresses longstanding circuit complexity bottlenecks for fundamental polynomial algebra problems—namely, computing greatest common divisors (GCDs), discriminants, resultants, Bézout coefficients, square-free factorizations, and inverses of Sylvester/Bézout matrices—in the AC⁰_F model. Prior to this work, only superpolynomial-size arithmetic circuits were known for these tasks. We present the first polynomial-size, constant-depth AC⁰_F arithmetic circuits for all these problems. Our method introduces a novel algorithmic paradigm that avoids explicit root access; instead, it implicitly handles root multiplicities and symmetric functions via structured matrix algebraic transformations and constant-depth evaluation of symmetric polynomials. Consequently, problems long believed “non-AC⁰-computable,” such as GCD computation, are now shown to reside in the class of polynomial-size, constant-depth circuits. The approach naturally extends to multivariate polynomials and multiple inputs, substantially enhancing parallelism and hardware feasibility in algebraic computation.
This study investigates the algebraicity and arithmetic properties of hypergeometric functions over the rational numbers, finite fields, and p-adic fields. Leveraging the SageMath computer algebra system, the work integrates techniques from algebraic number theory, finite field theory, and p-adic analysis to systematically implement, for the first time in an open-source framework, algorithms capable of determining algebraicity, computing valuations, and solving for minimal polynomials in positive characteristic. This implementation fills a critical gap in existing computational toolchains by enabling uniform arithmetic analysis of hypergeometric functions across multiple number-theoretic domains, thereby substantially enhancing SageMath’s capacity for algebraic manipulation of such functions.
This work investigates the efficient GPU implementation of high-order, block-level arithmetic operations for medium-sized large integers (ranging from 2¹⁵ to 2¹⁹ bits). Leveraging the functional language Futhark, the authors concisely express algorithms for addition, subtraction, multiplication, and division using high-level abstractions, while relying on the compiler to automatically map arrays to GPU registers for performance optimization. Through targeted compiler enhancements, the Futhark implementation achieves performance approaching that of hand-optimized C++/CUDA code and the CGBN library. These results demonstrate that high-level functional languages can effectively combine expressive power with computational efficiency in high-performance computing contexts.
This work addresses the long-standing absence of a truly quasi-linear time, output-sensitive algorithm for multiplying sparse polynomials with integer coefficients. By integrating modular black-box interpolation with sparse interpolation techniques, the authors present the first algorithm achieving rigorous quasi-linear bit complexity in this setting. Their approach refutes a prior claim of having resolved the problem and establishes output-sensitive quasi-linear bit complexity for integer-coefficient sparse polynomial multiplication. Moreover, over finite fields, the method further optimizes the bit complexity to be linear in the number of terms, the logarithm of the degree, and the logarithm of the field size.
This work addresses the inefficiency of solving structured sparse polynomial systems over finite fields by proposing an efficient resultant-based algorithm. Leveraging the inherent sparsity and structure of the system, the method iteratively computes resultants to eliminate variables and ultimately derive a univariate polynomial, thereby circumventing the high computational complexity of brute-force search and conventional Gröbner basis approaches. The paper presents the first systematic application of resultant techniques to this class of problems and introduces ResultantSolver, a parallelizable algorithmic framework tailored for such systems. Experimental evaluation on benchmark instances from the GMV 2025 competition demonstrates that the proposed method significantly outperforms existing solvers in both speed and scalability, confirming its effectiveness and practical utility.
This work reveals that Cloud TPUs exhibit severe performance disadvantages—up to 4,693–6,908× slower—than GPUs for finite-field cryptographic computations, primarily due to the absence of wide-integer ALUs and extremely low spatial utilization (only 6.25% in the M dimension) of their matrix compute units. To address this, the authors propose a “spatial collapse” model that reformulates low-degree polynomial arithmetic into matrix-based Number Theoretic Transforms (NTT), integrated with Montgomery reduction for efficient finite-field operations. The study provides the first quantitative characterization of TPUs’ structural limitations in exact-domain computation and introduces a reproducible measurement framework grounded in HLO-level post-hoc validation, effectively circumventing interference from XLA fusion optimizations. This approach establishes a new paradigm for heterogeneous cryptographic computing.