An Exploratory Framework for Future SETI Applications: Detecting Generative Reactivity via Language Models

📅 2025-06-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether large language models (LLMs) can unsupervisedly detect latent regularities in noisy, non-linguistic acoustic inputs—motivated by the challenge of recognizing communication intent from extraterrestrial intelligence (SETI) with unknown signaling conventions. Method: We propose the “Generative Reactivity” framework, which abandons conventional symbolic decoding assumptions and instead treats structural coherence in model outputs as evidence of underlying order in the input. We introduce the Semantic Induction Potential (SIP), a composite metric integrating entropy, syntactic coherence, compression gain, and repetition penalty to quantify response strength. Contribution/Results: Using zero-shot, cross-modal evaluation on GPT-2 small (117M), we observe statistically significant SIP increases for humpback whale song and nightingale vocalizations versus white noise (p < 0.01), while human speech elicits only moderate reactivity. These findings demonstrate that LLMs can autonomously perceive statistical structure in non-human acoustic signals without prior encoding knowledge—establishing a hypothesis-free paradigm for SETI signal detection.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsCognitive Modeling & Cognitive Systems: Neural Spike Coding

Application Category

Search and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
We present an exploratory framework to test whether noise-like input can induce structured responses in language models. Instead of assuming that extraterrestrial signals must be decoded, we evaluate whether inputs can trigger linguistic behavior in generative systems. This shifts the focus from decoding to viewing structured output as a sign of underlying regularity in the input. We tested GPT-2 small, a 117M-parameter model trained on English text, using four types of acoustic input: human speech, humpback whale vocalizations, Phylloscopus trochilus birdsong, and algorithmically generated white noise. All inputs were treated as noise-like, without any assumed symbolic encoding. To assess reactivity, we defined a composite score called Semantic Induction Potential (SIP), combining entropy, syntax coherence, compression gain, and repetition penalty. Results showed that whale and bird vocalizations had higher SIP scores than white noise, while human speech triggered only moderate responses. This suggests that language models may detect latent structure even in data without conventional semantics. We propose that this approach could complement traditional SETI methods, especially in cases where communicative intent is unknown. Generative reactivity may offer a different way to identify data worth closer attention.
Problem

Research questions and friction points this paper is trying to address.

Detecting structured responses in language models from noise-like inputs
Assessing linguistic behavior in generative systems without symbolic decoding
Evaluating latent structure detection in non-semantic data for SETI applications
Innovation

Methods, ideas, or system contributions that make the work stand out.

Using noise-like inputs to test language models
Defining Semantic Induction Potential (SIP) score
Detecting latent structure in unconventional data
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Po-Chieh Yu
Taiwan Astronomical Research Alliance (TARA), Taiwan
P
Po-Chieh Yu
Institute of Astronomy and Astrophysics, Academia Sinica, Taipei, 10617, Taiwan