🤖 AI Summary
This work addresses the inconsistency in conditional prediction and error accumulation in joint source-channel coding (JSCC) for wireless video transmission, which arises from asymmetric context between encoder and decoder. To overcome this, we propose a deep JSCC framework based on asymmetric contextual conditional learning that operates without explicit channel simulation. The encoder and decoder independently learn conditional representations, while a feature propagation mechanism leverages temporal correlations to effectively suppress error propagation. Furthermore, an entropy model combined with a masking strategy enables content-adaptive variable-bandwidth transmission. Experimental results demonstrate that the proposed method significantly outperforms existing deep video transmission approaches, achieving higher reconstruction quality and improved transmission efficiency by simultaneously mitigating error accumulation and reducing the frequency of intra-frame insertions.
📝 Abstract
In this paper, we propose a high-efficiency deep joint source-channel coding (JSCC) method for video transmission based on conditional coding with asymmetric context. The conditional coding-based neural video compression requires to predict the encoding and decoding conditions from the same context which includes the same reconstructed frames. However in JSCC schemes which fall into pseudo-analog transmission, the encoder cannot infer the same reconstructed frames as the decoder even a pipeline of the simulated transmission is constructed at the encoder. In the proposed method, without such a pipeline, we guide and design neural networks to learn encoding and decoding conditions from asymmetric contexts. Additionally, we introduce feature propagation, which allows intermediate features to be independently propagated at the encoder and decoder and help to generate conditions, enabling the framework to greatly leverage temporal correlation while mitigating the problem of error accumulation. To further exploit the performance of the proposed transmission framework, we implement content-adaptive coding which achieves variable bandwidth transmission using entropy models and masking mechanisms. Experimental results demonstrate that our method outperforms existing deep video transmission frameworks in terms of performance and effectively mitigates the error accumulation. By mitigating the error accumulation, our schemes can reduce the frequency of inserting intra-frame coding modes, further enhancing performance.