🤖 AI Summary
To address scalability and ultra-low-latency transmission challenges in real-time audio-video conferencing systems on FPGAs, this paper proposes a fully hardware-coordinated architecture. The design implements M-JPEG video encoding, PCM audio sampling, and a lightweight UDP protocol stack entirely in SystemVerilog, and integrates them end-to-end on a Nexys4 DDR FPGA. A key innovation is the tight hardware-level coupling between audio-video encoding and network transmission, enabling stable 30 FPS video streaming with synchronized decoding. Modular verification employs Cocotb, synthesis and implementation are performed in Vivado, and host-side synchronization playback is achieved via Python-based parsing. Experimental results demonstrate an end-to-end latency under 120 ms and throughput sufficient for real-time communication. The system is rigorously validated through both simulation and hardware deployment, confirming high stability, scalability, and suitability for resource-constrained FPGA platforms.
📝 Abstract
We present a design for an extensible video conferencing stack implemented entirely in hardware on a Nexys4 DDR FPGA, which uses the M-JPEG codec to compress video and a UDP networking stack to communicate between the FPGA and the receiving computer. This networking stack accepts real-time updates from both the video codec and the audio controller, which means that video will be able to be streamed at 30 FPS from the FPGA to a computer. On the computer side, a Python script reads the Ethernet packets and decodes the packets into the video and the audio for real time playback. We evaluate this architecture using both functional, simulation-driven verification in Cocotb and by synthesizing SystemVerilog RTL code using Vivado for deployment on our Nexys4 DDR FPGA, where we evaluate both end-to-end latency and throughput of video transmission.