This full-text Seed edition was converted from the ACM version of record under supplied ACM publication authorization. Attribution, DOI, and source links are retained.

ABSTRACT

Recommender Systems have become essential tools in any modern video-sharing platform. Although, recommender systems have shown to be effective in generating personalized suggestions in video-sharing platforms, however, they suffer from the so-called New Item problem. New item problem, as part of Cold Start prob- lem, happens when a new item is added to the system catalogue and the recommender system has no or little data available for that new item. In such a case, the system may fail to meaningfully recommend the new item to the users.

In this paper, we propose a novel recommender system that is based on visual tags, i.e., tags that are automatically annotated to videos based on visual description of the videos. Such visual tags can be used in an extreme cold start situation, where neither any rating, nor any tag is available for the new video. The visual tags could also be used in the moderate cold start situation when the new video might have been annotated with few tags. This type of content features can be extracted automatically without any human involvement and have been shown to be very effective in representing the video content.

We have used a large dataset of videos and shown that automatically extracted visual tags can be incorporated into the cold start recommendation process and achieve superior results compared to the recommendation based on human-annotated tags.

CCS CONCEPTS

    Information systems → Recommender systems.

KEYWORDS

video recommendation, visual features, multimedia, visual tags, cold start, content-based filtering, tag recommendation

ACM Reference Format: Mehdi Elahi, Reza Hosseini, Mohammad H. Rimaz, Farshad B. Moghaddam, and Christoph Trattner. 2020. Visually-Aware Video Recommendation in the Cold Start. In Proceedings of the 31st ACM Conference on Hypertext and Social Media (HT ’20), July 13–15, 2020, Virtual Event, USA. ACM, New York, NY, USA, 5 pages. https://doi.org/10.1145/3372923.3404778

1 INTRODUCTION

One of the main challenges in Recommender Systems (RSs) is the New Item problem. This problem which is part of a bigger challenge called Cold Start problem happens when a new item is added to the item catalogue and no rating has been provided by the users to that item [13, 32]. In such a case, the RS may fail to effectively recommend that new item to the users.

One of the recommendation techniques that can remedy the cold start problem is Content-Based Filtering (CBF) which can exploit content data (e.g., item tag) in order to compute similarity among items and generate relevant recommendation based on content similarities [11, 12, 42]. In video domain, the content data can be represented by different features, described with the following hierarchical levels: high-level features, representing semantics illustrated by the concepts and events happening within a video. An example can be a plot of the film The Good, the Bad and the Ugly, showing three gunslingers who are competing to find a buried cache of gold during the American Civil War [9]. Mid-level features, representing syntactic features with the existing objects within a video and interactions of these objects with each other. An example is people, horses and guns in the same film. At the lowest level, low-level features typically representing stylistic aspects of videos defined by the aesthetic characteristics of the videos. This includes the design aspects which could picture the specific style of a video production. As an example, in the same movie predominant colors are yellow and brown [8].

Traditionally, content-based video recommendation has been focused on exploiting high-level and mid-level features [4, 22]. While they are effective in representing videos, however, they are expensive to acquire as they need human-annotation typically by a large network of users. Indeed, there are cases where such features are missing and causing an extreme cases of cold start problem.

In such cases, even the most complicated recommendation algorithms could be unable to generate recommendation of such new items. In video domain, this is a case where a video is uploaded to a video-sharing platform and none of the users has yet added any type of data (e.g., tags).

In this paper, we address this problem and propose a novel feature set called visual tags. Such features are automatically extracted and added to the video items. We build predictive models that can learn the correlation among visual features and the tags added to other videos. Such models are used to predict the tags for a new video item with no tags. Such visual tags are then being exploited in order to generate personalized recommendation for users. We have performed different experiments in order to evaluate the recommendation based on visual tags. We have considered two evaluation scenarios, i.e., moderate cold start and extreme cold start scenarios. The results have shown the effectiveness of the proposed features in both scenarios in comparison to recommendation based on human-annotated tags.

It is worth noting that, we focused on recommendation based on tags as prior studies have shown the superior performance of tags in comparison to other types of content-features (e.g., genre) [9]. Furthermore, using visual tags enables the system to include explanation when presenting recommended videos to users. Explanation may enhance transparency of the system and result in higher user satisfaction [37]. This is not very feasible with the pure (low-level) visual features.

2 BACKGROUND

This work is mainly related to two research fields, i.e., tag-based RSs and visually-aware RSs [5, 6, 22]. Several prior works have incorporated human-annotated tags into recommendation process

[1, 16, 17, 27, 29, 40, 41]. One of the prior works [21] integrated tagbased similarity within an extended Collaborative Filtering (CF) in order to improve the recommendation. In [14], the authors proposed a modified version of the SVD++ matrix-factorization model [20] by replacing the usage of implicit feedback with tagging information. This results in a substantial improvement of the RS performance. In [24] another matrix-factorization model was proposed which expands the item model with latent factors vectors associated to the features of the items. In [15], SVD++ is again extended with an approach similar to those described in [14, 24], in order to deal with a cross-domain recommendation scenario.

The usage of the low-level visual features has drawn minor attention in RS (e.g., in [8, 25, 31]). This is while this has been extensively investigated in the other fields such as computer vision [28, 35].

[2, 19] provide comprehensive surveys on the state-of-the-art techniques related to video content analysis and classification, and discuss a large number of low-level features (e.g. visual, textual, or auditory). Authors in [30] propose a framework for movie genre classification based only on visual features. Moreover, [36] proposes a deep learning approach to automatically detect the director of a movie based on low-level visual features. This work differs from the prior works as it proposes visual tags instead of pure visual features. The advantage of using visual tags is the explainability of recommendation based on visual tags in comparison to visual features.

3 PROPOSED METHOD

Through querying YouTube, we obtained a huge dataset of 13923 movie trailers [7, 26, 31] based on the titles available in the Movielens dataset [18]. Prior work showed a high similarity of visual

features extracted from movie trailers and their respective fulllength movies [8].

Our research methodology encompasses the following steps: Movie Segmentation: every movie is segmented into shots, i.e., se- quences of consecutive frames captured without interruption of the camera; Key-Frame Detection: within every shot the middle frame is selected as representative of the shot (key-frame); Feature Extraction: every key-frame is analyzed and the visual features are extracted; Feature Aggregation: the feature vectors are aggregated over the entire movie to form a feature vector descriptive of the whole movie; Prediction: aggregated visual feature vectors are used to train the prediction algorithms. The dataset is published online and available for download [1].

Movie Segmentation & Key-frame Detection. In order to segment movies into shots i.e. sequences of consecutive frames recorded without camera interference, we used a method based on Color Histogram Distance. This is due to the fact that transition between two shots of the video is typically very abrupt. By comparing the color histogram of every movie frame, the histogram intersection is computed to compare the activities. Lets denote ht and ht +1 as histograms of successive frames, then intersection is computed according to the following equation


Do you like what you are reading? Subscribe to receive updates.

Unsubscribe anytime