Institution profile

Lossfunk

Industry research
Research library13linked papers
Opportunities0open roles
Selected work

Representative Papers

Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable

Sep 29, 2026

This study addresses the tendency of language models to erroneously defer to authoritative sources, a phenomenon fundamentally distinct from conventional user sycophancy. By employing causal intervention and activation direction fitting, this work disentangles the model’s response mechanisms toward verified sources and user inputs. It provides the first demonstration that source deference and user agreement are behaviorally non-interchangeable, proposing an independent evaluation framework accordingly. A key contribution is the identification of an “authority direction” representation that transfers across datasets such as Trivia and PIQA. Ablating this direction reduces erroneous compliance rates by 65–80 percentage points without compromising performance on benchmarks like MMLU-Pro.

0 citationsRead paper

Denoising Models Develop Human-Like Perceptual Illusion Representations Across Architectures

Jul 19, 2026

This study investigates whether denoising models genuinely encode human visual illusions and their underlying mechanisms within internal representations. By analyzing internal activations across multiple architectures—combined with feature visualization, channel ablation, psychophysical modeling (FLODOG), and parametric illusion-strength experiments—the work uncovers, for the first time, a perception-like “phantom” representation that is decoupled from model output. Specifically, certain channels in intermediate layers exhibit high sensitivity to brightness illusions: their activation magnitudes correlate strongly with human perceptual judgments (Spearman ρ ≥ 0.70) and vary monotonically with illusion strength, yet they exert no influence on the final reconstructed pixels. These findings provide causal evidence for human-like perceptual mechanisms embedded within deep neural networks.

0 citationsRead paper

Frontier Coding Agents Use Metaprogramming to Adapt to Unfamiliar Programming Languages

Jun 09, 2026

This study investigates the capability of state-of-the-art large language model agents to handle esoteric programming languages—such as Brainfuck and Befunge-98—where their performance remains unclear despite strong results in mainstream languages. The authors introduce a systematic evaluation pipeline encompassing file editing, local execution, and hidden testing to assess multiple leading agents. Findings reveal that top-performing models, including Claude Opus 4.6 and GPT-5.4 xhigh, predominantly rely on metaprogramming strategies—specifically, generating target-language code via intermediate Python scripts—rather than directly writing in unfamiliar languages; disabling this approach leads to a marked performance drop. Moreover, distilling these auxiliary programs into weaker models (e.g., Sonnet 4.6 and GPT-5.4 mini) substantially enhances their effectiveness, highlighting resource orchestration as a critical factor underlying performance disparities among agents.

0 citationsRead paper

Discovering Reinforcement Learning Interfaces with Large Language Models

May 05, 2026

This work addresses the heavy reliance on manual design in defining environment interfaces—specifically observation mappings and reward functions—in reinforcement learning, for which automated solutions are largely absent. The authors propose LIMEN, a framework that achieves, for the first time, the joint automatic discovery of both observation and reward functions. LIMEN leverages a large language model to guide an evolutionary algorithm that generates executable programs from raw simulation states, iteratively refining the entire interface using policy training feedback. Experiments demonstrate that, given only trajectory-level success signals, LIMEN successfully discovers effective interfaces in both discrete grid-world and continuous control tasks. In contrast, optimizing either component in isolation fails completely in at least one domain, underscoring the necessity and superiority of co-design and substantially reducing the engineering cost of interface specification.

0 citationsRead paper

See, Symbolize, Act: Grounding VLMs with Spatial Representations for Better Gameplay

Mar 12, 2026

This work addresses the challenge faced by vision-language models (VLMs) in translating perceptual inputs into executable actions within interactive environments. The authors propose a method that integrates raw visual frames with symbolic scene representations and present the first systematic evaluation of how symbolic information influences VLM-based action generation. Multimodal policy experiments are conducted across Atari, VizDoom, and AI2-THOR platforms. Results demonstrate that high-quality symbolic representations substantially enhance VLMs’ decision-making performance in gameplay. However, symbols extracted autonomously by the model are often compromised by its inherent limitations and environmental complexity, and noisy or inaccurate symbols can severely degrade action efficacy. The study identifies the reliability of symbol extraction as a critical bottleneck for achieving effective symbol grounding in embodied interactive tasks.

0 citationsRead paper
Recent publications

Latest Papers

Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable

Sep 29, 2026

This study addresses the tendency of language models to erroneously defer to authoritative sources, a phenomenon fundamentally distinct from conventional user sycophancy. By employing causal intervention and activation direction fitting, this work disentangles the model’s response mechanisms toward verified sources and user inputs. It provides the first demonstration that source deference and user agreement are behaviorally non-interchangeable, proposing an independent evaluation framework accordingly. A key contribution is the identification of an “authority direction” representation that transfers across datasets such as Trivia and PIQA. Ablating this direction reduces erroneous compliance rates by 65–80 percentage points without compromising performance on benchmarks like MMLU-Pro.

0 citationsRead paper

Denoising Models Develop Human-Like Perceptual Illusion Representations Across Architectures

Jul 19, 2026

This study investigates whether denoising models genuinely encode human visual illusions and their underlying mechanisms within internal representations. By analyzing internal activations across multiple architectures—combined with feature visualization, channel ablation, psychophysical modeling (FLODOG), and parametric illusion-strength experiments—the work uncovers, for the first time, a perception-like “phantom” representation that is decoupled from model output. Specifically, certain channels in intermediate layers exhibit high sensitivity to brightness illusions: their activation magnitudes correlate strongly with human perceptual judgments (Spearman ρ ≥ 0.70) and vary monotonically with illusion strength, yet they exert no influence on the final reconstructed pixels. These findings provide causal evidence for human-like perceptual mechanisms embedded within deep neural networks.

0 citationsRead paper

Frontier Coding Agents Use Metaprogramming to Adapt to Unfamiliar Programming Languages

Jun 09, 2026

This study investigates the capability of state-of-the-art large language model agents to handle esoteric programming languages—such as Brainfuck and Befunge-98—where their performance remains unclear despite strong results in mainstream languages. The authors introduce a systematic evaluation pipeline encompassing file editing, local execution, and hidden testing to assess multiple leading agents. Findings reveal that top-performing models, including Claude Opus 4.6 and GPT-5.4 xhigh, predominantly rely on metaprogramming strategies—specifically, generating target-language code via intermediate Python scripts—rather than directly writing in unfamiliar languages; disabling this approach leads to a marked performance drop. Moreover, distilling these auxiliary programs into weaker models (e.g., Sonnet 4.6 and GPT-5.4 mini) substantially enhances their effectiveness, highlighting resource orchestration as a critical factor underlying performance disparities among agents.

0 citationsRead paper

Discovering Reinforcement Learning Interfaces with Large Language Models

May 05, 2026

This work addresses the heavy reliance on manual design in defining environment interfaces—specifically observation mappings and reward functions—in reinforcement learning, for which automated solutions are largely absent. The authors propose LIMEN, a framework that achieves, for the first time, the joint automatic discovery of both observation and reward functions. LIMEN leverages a large language model to guide an evolutionary algorithm that generates executable programs from raw simulation states, iteratively refining the entire interface using policy training feedback. Experiments demonstrate that, given only trajectory-level success signals, LIMEN successfully discovers effective interfaces in both discrete grid-world and continuous control tasks. In contrast, optimizing either component in isolation fails completely in at least one domain, underscoring the necessity and superiority of co-design and substantially reducing the engineering cost of interface specification.

0 citationsRead paper

See, Symbolize, Act: Grounding VLMs with Spatial Representations for Better Gameplay

Mar 12, 2026

This work addresses the challenge faced by vision-language models (VLMs) in translating perceptual inputs into executable actions within interactive environments. The authors propose a method that integrates raw visual frames with symbolic scene representations and present the first systematic evaluation of how symbolic information influences VLM-based action generation. Multimodal policy experiments are conducted across Atari, VizDoom, and AI2-THOR platforms. Results demonstrate that high-quality symbolic representations substantially enhance VLMs’ decision-making performance in gameplay. However, symbols extracted autonomously by the model are often compromised by its inherent limitations and environmental complexity, and noisy or inaccurate symbols can severely degrade action efficacy. The study identifies the reliability of symbol extraction as a critical bottleneck for achieving effective symbol grounding in embodied interactive tasks.

0 citationsRead paper