Score
Designs and implements the integration of a rendering engine into a data-processing pipeline, including hooking renderer inputs and outputs into pipeline stages and managing synchronization between renderer frames and sensor or data streams. Builds calibration and validation tooling to expose and tune rendering parameters to match scenes or reference data, and to compare renderer outputs against ground truth.
为了解决动作条件视频模型数据获取难题,本文提出基于虚幻引擎的两阶段合成数据生成管道,生成大规模动作条件多视角视频。
Existing data engineering pipelines exhibit unstable data quality, delayed responsiveness, and poor fault tolerance in dynamic data environments, often degrading or failing due to data distribution shifts. To address these challenges, this paper proposes a three-level evolutionary data pipeline framework—progressing from *optimization* to *self-awareness* to *self-adaptation*—integrating operator composition optimization, online parameter tuning, real-time state monitoring, and feedback control. The framework enables autonomous pipeline diagnosis, dynamic parameter adjustment, and closed-loop environmental response. Its core innovation lies in transforming conventional static pipelines into intelligent systems endowed with perception–decision–execution capabilities. Experimental evaluation demonstrates significant improvements: data quality stability increases markedly, with error fluctuation reduced by 42%, and environmental adaptability is substantially enhanced. The framework establishes a deployable, automation-ready paradigm for next-generation data engineering.
研究比较了五个库的七种输入路径,从RGB JPEG文件到CUDA float16批处理,评估图像增强管道的吞吐量和GPU内存使用,以优化数据预处理效率。
Modern frontend frameworks (e.g., React, Vue) employ declarative rendering, state synchronization, and component reuse—introducing substantial complexity in generating executable code from design mockups and limiting the performance of existing vision-language models (VLMs). Method: We propose a reflective agent workflow and introduce three novel image-code paired data synthesis strategies: evolutionary (enhancing diversity), waterfall-model (ensuring logical consistency), and incremental development (progressively increasing complexity). We establish a “vision-first” VLM training paradigm—“observe image before generating code”—and integrate self-contained code extraction, multi-stage rendering, and descriptive caption generation. Our approach is instantiated and evaluated on the Flame VLM. Contribution/Results: Experiments demonstrate significant improvements over baselines in React code generation. Pass@k evaluation confirms that vision-prior understanding substantially enhances both functional correctness and overall code quality.
This study addresses the software portability gap between model-driven development for cyber-physical systems and multicore execution platforms. To bridge this divide, we propose a workflow-preserving automated translation framework that converts Simulink models into OpenCL code. Specifically, the framework leverages GPU Coder to extract CUDA implementations and employs a custom pipeline to translate them into OpenCL host and device code, enabling cross-platform adaptation of syntax, APIs, and parameter packing. This approach transforms data-parallel models into multicore executables without requiring manual rewriting. As a key contribution, we validate the feasibility and effectiveness of the proposed retargeting pipeline by successfully deploying a Frenet trajectory planner on the Kalray MPPA multicore platform, demonstrating its practical applicability for high-performance embedded systems.
This study addresses the deployment challenges and hardware lock-in issues of Bird's Eye View (BEV) perception models caused by operator mismatches, particularly with sparse convolutions. To this end, we propose BEVPIPE, a framework that introduces a novel architecture decoupling dense subgraphs from sparse operators to eliminate CUDA dependencies. By integrating model partitioning, portable GPU computation APIs, custom operator extensions, and shared memory mechanisms, BEVPIPE enables efficient and portable deployment of multimodal BEV perception within standard inference runtimes. Experimental results demonstrate that the proposed framework achieves a 19.5× end-to-end speedup while retaining 98.5% of the reference mean Average Precision (mAP). Furthermore, its compatibility across diverse GPU backends is empirically validated, highlighting its practicality for flexible, high-performance deployment in autonomous driving applications.
This study addresses the lack of paired data and geometric evaluation benchmarks for video rendering-to-real translation by proposing the first end-to-end evaluation framework based on reconstructed digital twins. The framework establishes a benchmark comprising twelve scenes, providing editable 3D replicas, proxy renderings, and real target videos. By integrating neural reconstruction with visual geometry techniques, it automates data processing through static-dynamic decoupling and camera registration. Furthermore, combining multi-view verification with scene decomposition enables automated assessment of appearance fidelity and geometric-dynamic consistency. The project releases 1,496 frames of paired data with an 85.7% retention rate, achieving a depth si-RMSE of 0.2170 and an instance mIoU of 0.3673. These contributions significantly reduce manual modeling costs while effectively identifying causes of generation failures.
This study addresses the limitation of existing video benchmarks in evaluating models' capacity to follow fine-grained procedural world events. To this end, it constructs a novel benchmark grounded in replayable world records, generating videos through synchronized multi-view rendering and agent representations. Furthermore, the work proposes a vision-language model-based logic-rendering alignment metric that enables fine-grained consistency verification from procedural states to visual outputs. This approach effectively quantifies entity control, long-term memory, and interaction success rates, thereby establishing a rigorous evaluation standard for the visual fidelity of programmable world models.
研究通过分解浏览器架构,利用Web Workers、WebGL和WebAssembly优化DOM源粒子效果,旨在识别主要瓶颈并测试层优化是否能提升用户体验帧率。
This work addresses the evaluation gap between visual realism and verifiability in fine-grained image editing, as well as the subjective uncertainty inherent in human or VLM-based assessments. To this end, it introduces VeriEdit-Bench, the first deterministic evaluation benchmark grounded in structured asset source code. Methodologically, SVG source code is leveraged to control the editing process, generating precise targets and pixel-level masks that enable quantitative evaluation across four axes, including fidelity and preservation rate. This benchmark eliminates subjective bias by providing a fully reproducible scoring mechanism. Furthermore, it reveals significant performance disparities among existing models across multiple dimensions, cross-scenario ranking inversions, and distinct capability failure modes.