🤖 AI Summary
Current affective image editing methods are constrained by predefined strategies, limiting their ability to adaptively convey target emotions in a content-aware manner. This work proposes EmoScope, a multi-agent framework that formulates emotional editing as an open-ended affordance discovery task. By performing emotion-conditioned affordance reasoning, EmoScope dynamically constructs an image-specific editable space. The approach leverages a semantic hierarchy—comprising anchors, variables, and context—to balance content consistency with emotional expressiveness, and incorporates a plan-level interaction mechanism to enable lightweight user refinement. In large-scale human evaluations covering eight emotion categories, 88.1% of participants preferred results generated by EmoScope, demonstrating the method’s emotional adaptability and effectiveness.
📝 Abstract
Emotional image editing requires more than applying affective filters or modifying predefined visual factors: an effective edit must identify what a particular image can afford for a target emotion. Existing affective image manipulation methods, including recent agentic variants, largely operate within bounded strategy spaces based on predefined factor taxonomies, knowledge libraries, or conventional editing templates, and therefore often miss image-specific, context-grounded strategies. We introduce EmoScope, a multi-agent framework that reframes the task from "how should I edit?" to "what can I edit?" EmoScope first discovers an image-specific editable space through emotion-conditioned affordance reasoning, then uses a semantic hierarchy of anchors, variables, and context to balance content consistency and emotional expressiveness before executing and verifying the edit. Because its plans are expressed as image-specific affordances rather than retrieved templates, EmoScope also exposes the editing strategy as an interactive surface for user refinement at the plan level. In a large-scale human evaluation covering all eight Mikels emotion categories, with 4,693 valid responses across 1,824 pairwise questions, participants preferred EmoScope over two competitive baselines by 88.1% on average. Attribution analysis further shows that EmoScope selects target-emotion-adaptive strategies rather than applying a uniform template. The same affordance-level plan also supports lightweight user refinement in an interactive pilot. Finally, we show that classifier-based metrics exhibit emotion-conditional blind spots toward non-stereotypical, context-grounded edits, and present a relative content-emotion preference-affinity landscape showing that EmoScope's advantage varies systematically across image-emotion combinations.