🤖 AI Summary
Real-time estimation of TV device reach—critical for millisecond-scale audience targeting in advertising—is hindered by traditional offline SQL-based computation, which incurs ~24-hour latency and fails to meet real-time requirements.
Method: We propose a real-time, multi-dimensional device reach prediction system. Its core innovation is an enhanced MinHash algorithm supporting hierarchical aggregation and set intersection, integrated with HyperLogLog (HLL) for constant-space, approximate cardinality estimation over billion-scale user datasets; further accelerated via SIMD vectorization.
Contribution/Results: The system achieves millisecond-level end-to-end latency, ≤5% relative error, and accuracy comparable to offline baselines. It currently powers real-time audience targeting for millions of advertising campaigns, delivering a 4× speedup over prior solutions while maintaining production-grade reliability and scalability.
📝 Abstract
Predicting the right number of TVs (Device Reach) in real-time based on a user-specified targeting attributes is imperative for running multi-million dollar ADs business. The traditional approach of SQL queries to join billions of records across multiple targeting dimensions is extremely slow. As a workaround, many applications will have an offline process to crunch these numbers and present the results after many hours. In our case, the solution was an offline process taking 24 hours to onboard a customer resulting in a potential loss of business. To solve this problem, we have built a new real-time prediction system using MinHash and HyperLogLog (HLL) data sketches to compute the device reach at runtime when a user makes a request. However, existing MinHash implementations do not solve the complex problem of multilevel aggregation and intersection. This work will show how we have solved this problem, in addition, we have improved MinHash algorithm to run 4 times faster using Single Instruction Multiple Data (SIMD) vectorized operations for high speed and accuracy with constant space to process billions of records. Finally, by experiments, we prove that the results are as accurate as traditional offline prediction system with an acceptable error rate of 5%.