Sunday, September 13, 2026

AI Finds Sports Highlights and Develops Evaluation Technology

Input
2026-09-13 08:00:00
Updated
2026-09-13 08:00:00
Overview of the sports highlight ranking algorithm. Provided by UNIST

[Financial News] A technology has been developed to properly evaluate the performance of artificial intelligence (AI) models that extract highlights from sports videos. The technology is expected to expand beyond the automatic production of sports highlights to applications such as summarizing films and dramas, organizing meeting records, and searching for key scenes in long surveillance videos.
A research team led by Professor Kim Tae-hwan at the Ulsan National Institute of Science and Technology (UNIST) announced on the 13th that it had created SVHighlights (Sport Video Highlights), a large-scale benchmark for evaluating AI models that extract highlights from sports videos.
A benchmark is a kind of common test used to compare which AI system produces more accurate answers. Most videos in existing benchmarks were approximately two to four minutes long. Creating reference answers requires people to watch each video from beginning to end and mark every highlight, meaning that longer videos demand astronomical amounts of time and money.
By contrast, the benchmark developed by the research team consists of 320 videos covering eight sports: soccer, baseball, basketball, volleyball, American football, ice hockey, rugby, and racing. It totals 640.18 hours. With an average length of two hours per video, the videos are 30 to 60 times longer than those in existing datasets and are similar in length to actual sports broadcasts.
To build this extensive benchmark, human involvement was limited to marking the start and end points of each game once per video and scanning automatically paired frames in grid images. Only 0.18% of the automatically paired frames were found to contain errors during this verification process.
The team was able to create the long and extensive dataset cost-effectively by using official sports highlight videos already available online. It also developed a matching algorithm that automatically identifies the timestamps—down to the minute and second—at which each highlight video originated from the original footage. The algorithm compares frames from the original and highlight videos at the pixel level using peak signal-to-noise ratio (PSNR) and pairs the most similar scenes. When designing the algorithm, the researchers considered not only visual similarity but also the chronological order of the scene relative to the previously identified scene.
The research team also developed an AI model called TF-SELECTOR, which combines existing scene-segmentation models, speech-recognition models, vision-language models, and large language models to effectively extract highlights from long videos. When tested on the SVHighlights benchmark, TF-SELECTOR outperformed existing models on most metrics. Compared with the next-best model, its HIT@1 score—measuring whether the scene selected as most important was an actual highlight—was 2.50 percentage points higher. Its HIT@K score—measuring how many actual highlights were found when selecting as many candidates as the number of actual highlights—was 4.04 percentage points higher, while its IoU, which indicates the overlap between predicted and reference segments, was 2.95 points higher.
Professor Kim Tae-hwan said, "By using highlights already produced by broadcasters, we replaced the reference-labeling process that previously required people to spend several hours. Because the data can continue to be expanded in the same way, it will help objectively evaluate long-video analysis models and develop models with better performance."
The research findings were accepted last month on the 9th at ACM KDD, an internationally prestigious conference in the field of data mining, held in Jeju Province.

[email protected] Yeon Ji-an Reporter