We present a novel multi-modal unspoken punctuation prediction system for the English language which combines acoustic and text features. We demonstrate for the first time, that by relying exclusively on synthetic data generated using a prosody-aware text-to-speech system, we can outperform a model trained with expensive human audio recordings on the unspoken punctuation prediction problem. Our model architecture is well suited for on-device use. This is achieved by leveraging hash-based embeddings of automatic speech recognition text output in conjunction with acoustic features as input to a quasi-recurrent neural network, keeping the model size small and latency low.
In the field of Information Retrieval, word embedding models have shown to be effective in several tasks. In this paper, we show how one of these neural embedding techniques can be adapted to the recommendation task. This adaptation only makes use of collaborative filtering information, and the results show that it is able to produce effective recommendations efficiently.
The evaluation of recommender systems is an area with unsolved questions at several levels. Choosing the appropriate evaluation metric is one of such important issues. Ranking accuracy is generally identified as a prerequisite for recommendation to be useful. Ranking metrics have been adapted for this purpose from the Information Retrieval field into the recommendation task. In this article, we undertake a principled analysis of the robustness and the discriminative power of different ranking metrics for the offline evaluation of recommender systems, drawing from previous studies in the information retrieval field. We measure the robustness to different sources of incompleteness that arise from the sparsity and popularity biases in recommendation. Among other results, we find that precision provides high robustness while normalized discounted cumulative gain offers the best discriminative power. In dealing with cold users, we also find that the geometric mean is more robust than the arithmetic mean as aggregation function over users.
In this paper we provide a brief introduction to a new version of the Coruna Corpus Tool. Currently available for Windows, macOS and Linux, the Coruna Corpus Tool is a corpus management tool that facilitates the retrieval of information from an indexed textual repository. Although it works like most concordance programs, its distinguishing feature is that it allows users to search for old or non-standard characters and tags in texts and metadata files, as well as to extract and export specific data for the purposes of research. With a new set of advanced search features and other recent improvements, researchers now have access to functionalities that significantly enhance the previous user experience.
In this paper we provide a brief introduction to a new version of the Coruña Corpus Tool. Currently available for Windows, macOS and Linux, the Coruña Corpus Tool is a corpus management tool that facilitates the retrieval of information from an indexed textual repository. Although it works like most concordance programs, its distinguishing feature is that it allows users to search for old or non-standard characters and tags in texts and metadata files, as well as to extract and export specific data f or the purposes o f r esearch. With a new s et o f advanced search features and other recent improvements, researchers now have access to functionalities that significantly enhance the previous user experience.
Information retrieval addresses the information needs of users by delivering relevant pieces of information but requires users to convey their information needs explicitly. In contrast, recommender systems offer personalized suggestions of items automatically. Ultimately, both fields help users cope with information overload by providing them with relevant items of information. This thesis aims to explore the connections between information retrieval and recommender systems. Our objective is to devise recommendation models inspired in information retrieval techniques. We begin by borrowing ideas from the information retrieval evaluation literature to analyze evaluation metrics in recommender systems [2]. Second, we study the applicability of pseudo-relevance feedback models to different recommendation tasks [1]. We investigate the conventional top-N recommendation task [5, 4, 6, 7], but we also explore the recently formulated user-item group formation problem [3] and propose a novel task based on the liquidation of long tail items [8]. Third, we exploit ad hoc retrieval models to compute neighborhoods in a collaborative filtering scenario [9, 10, 12]. Fourth, we explore the opposite direction by adapting an effective recommendation framework to pseudo-relevance feedback [13, 11]. Finally, we discuss the results and present our conclusions. In summary, this doctoral thesis adapts a series of information retrieval models to recommender systems. Our investigation shows that many retrieval models can be accommodated to deal with different recommendation tasks. Moreover, we find that taking the opposite path is also possible. Exhaustive experimentation confirms that the proposed models are competitive. Finally, we also perform a theoretical analysis of some models to explain their effectiveness. Advisors: Álvaro Barreiro and Javier Parapar. Committee members: Gabriella Pasi, Pablo Castells and Fidel Cacheda. The dissertation is available at: https://www.dc.fi.udc.es/~dvalcarce/thesis.pdf.
PRIN is a neural based recommendation method that allows the incorporation of item prior information into the recommendation process. In this work we study how the system behaves in terms of novelty and diversity under different configurations of item prior probability estimations. Our results show the versatility of the framework and how its behavior can be adapted to the desired properties, whether accuracy is preferred or diversity and novelty are the desired properties, or how a balance can be achieved with the proper selection of prior estimations.
Word embeddings techniques have attracted a lot of attention recently due to their effectiveness in different tasks. Inspired by the continuous bag-of-words model, we present prefs2vec, a novel embedding representation of users and items for memory-based recommender systems that rely solely on user–item preferences such as ratings. To improve the performance and prevent overfitting, we use a variant of dropout as regularization, which can leverage existent word2vec implementations. Additionally, we propose a procedure for incremental learning of embeddings that boosts the applicability of our proposal to production scenarios. The experiments show that prefs2vec with a standard memory-based recommender system outperforms all the state-of-the-art baselines in terms of ranking accuracy, diversity, and novelty.
Information retrieval addresses the information needs of users by delivering relevant pieces of information but requires users to convey their information needs explicitly. In contrast, recommender systems offer personalized suggestions of items automatically. Ultimately, both fields help users cope with information overload by providing them with relevant items of information. This thesis aims to explore the connections between information retrieval and recommender systems. Our objective is to devise recommendation models inspired in information retrieval techniques. We begin by borrowing ideas from the information retrieval evaluation literature to analyze evaluation metrics in recommender systems [2]. Second, we study the applicability of pseudo-relevance feedback models to different recommendation tasks [1]. We investigate the conventional top-N recommendation task [5, 4, 6, 7], but we also explore the recently formulated user-item group formation problem [3] and propose a novel task based on the liquidation of long tail items [8]. Third, we exploit ad hoc retrieval models to compute neighborhoods in a collaborative filtering scenario [9, 10, 12]. Fourth, we explore the opposite direction by adapting an effective recommendation framework to pseudo-relevance feedback [13, 11]. Finally, we discuss the results and present our conclusions. In summary, this doctoral thesis adapts a series of information retrieval models to recommender systems. Our investigation shows that many retrieval models can be accommodated to deal with different recommendation tasks. Moreover, we find that taking the opposite path is also possible. Exhaustive experimentation confirms that the proposed models are competitive. Finally, we also perform a theoretical analysis of some models to explain their effectiveness. Advisors : Álvaro Barreiro and Javier Parapar. Committee members : Gabriella Pasi, Pablo Castells and Fidel Cacheda. The dissertation is available at: https://www.dc.fi.udc.es/~dvalcarce/thesis.pdf.
In this paper, we present PRIN, a probabilistic collaborative filtering approach for top-N recommendation. Our proposal relies on continuous bag-of-words (CBOW) neural model. This fully connected feedforward network takes as input the item profile and produces as output the conditional probabilities of the users given the item. With that information, our model produces item recommendations through Bayesian inversion. The inversion requires the estimation of item priors. We propose different estimates based on centrality measures on a graph that models user-item interactions. An exhaustive evaluation of this proposal shows that our technique outperforms popular state-of-the-art baselines regarding ranking accuracy while showing good values of diversity and novelty.
galegoEste artigo presenta as actividades desenvolvidas polo grupo de innovacion educativa en Sistemas de Acceso a Informacion durante o curso 2017/2018. Este grupo, con docencia na Facultade de Informatica da Universidade da Coruna, realizou accions en tres linas de actuacion diferentes. A primeira delas, dirixida a mellora da calidade nos metodos de avaliacion, consiste no emprego dun protocolo para a deteccion de plaxios en practicas de programacion. A segunda actividade pretende mellorar a empregabilidade do alumnado e consiste en utilizar unha metodoloxia de aprendizaxe baseada en proxectos xunto cunha serie de ferramentas avanzadas para desenvolvemento software, permitindo recrear a actividade que deberan levar a cabo cando se incorporen ao mundo laboral. Por ultimo, e de cara a aumentar o conecemento das alternativas profesionais do alumnado, organizaronse unha serie de seminarios e charlas impartidas por profesionais dunha empresa internacional, unha empresa local multidisciplinar e un investigador da contorna academica. A experiencia obtida das diferentes actividades foi satisfactoria e enriquecedora tanto para o alumnado como para o profesorado, que xa baralla melloras de cara aos vindeiros cursos academicos. EnglishThis paper presents the activities performed by the educative innovation group in Information Access Systems during the academic year 2017/2018. This group, with teaching at the Faculty of Informatics of the University of A Coruna, carried out actions addressing three different topics. The first action was designed to improve the quality of the evaluation methods, and consisted in following a protocol for detecting plagiarism in programming exercises. The second activity aimed to improve the employability of the students and consisted in using a methodology based on project-based learning along with a series of advanced tools for software development, which recreated the activity that the students will carry out when they obtain their first job. Lastly, heading towards a better knowledge about the available professional alternatives, a series of seminars and talks were organized, which were performed by professionals from an international company, a local interdisciplinary company, and a researcher from an academic institution. The experience obtained from the different activities was satisfactory for both students and teachers, who are already considering improvements for the next academic years.
Information Retrieval is not any more exclusively about document ranking. Continuously new tasks are proposed on this and sibling fields. With this proliferation of tasks, it becomes crucial to have a cheap way of constructing test collections to evaluate the new developments. Building test collections is time and resource consuming: it requires time to obtain the documents, to define the user needs and it requires the assessors to judge a lot of documents. To reduce the latest, pooling strategies aim to decrease the assessment effort by presenting to the assessors a sample of documents in the corpus with the maximum number of relevant documents in it. In this paper, we propose the preliminary design of different techniques to easily and cheapily build high-quality test collections without the need of having participants systems.
Pseudo-relevance feedback (PRF) provides an automatic method for query expansion in Information Retrieval. These techniques find relevant expansion terms using the top retrieved documents with the original query. In this paper, we present an approach based on linear methods called LiMe that formulates the PRF task as a matrix factorization problem. LiMe learns an inter-term similarity matrix from the pseudo-relevant set and the query that uses for computing expansion terms. The experiments on five datasets show that LiMe outperforms state-of-the-art baselines in most cases.
•The definition of the User-Item Group Formation (UI-GF) problem.•Two formalizations of the UI-GF problem with different properties.•Experiments on five public Location-Based Social Network (LBSN) datasets.•Comprehensive comparisons of several algorithms for solving UI-GF.•Experiments showing the effectiveness and efficiency of the proposed methods.
The evaluation of Recommender Systems is still an open issue in the field. Despite its limitations, offline evaluation usually constitutes the first step in assessing recommendation methods due to its reduced costs and high reproducibility. Selecting the appropriate metric is a critical and ranking accuracy usually attracts the most attention nowadays. In this paper, we aim to shed light on the advantages of different ranking metrics which were previously used in Information Retrieval and are now used for assessing top-N recommenders. We propose methodologies for comparing the robustness and the discriminative power of different metrics. On the one hand, we study cut-offs and we find that deeper cut-offs offer greater robustness and discriminative power. On the other hand, we find that precision offers high robustness and Normalised Discounted Cumulative Gain provides the best discriminative power.
The research community has historically addressed the collaborative filtering task in several fashions. Although model-based approaches such as matrix factorisation attract substantial research efforts, neighbourhood-based recommender systems are effective and interpretable techniques. The performance of neighbour-based methods is strongly tied to the clustering strategies. In this paper, we show that there is room for improvement in this type of recommenders. For showing that, we build an oracle which yields approximately optimal neighbourhoods. We obtain ground truth neighbourhoods using the oracle and perform an analytical study of those to characterise them. As a result of our analysis, we propose to change the user profile size normalisation that cosine similarity employs in order to improve the neighbourhoods computed with k-NN algorithm. Additionally, we present a more appropriate oracle for current grouping strategies which leads us to include the IDF effect on the cosine formulation. An extensive experimentation on four datasets shows an increase in ranking accuracy, diversity and novelty using these cosine variants. This work shed light on the benefits of this type of analysis and paves the way for future research in the characterisation of good neighbourhoods for collaborative filtering.
The adaptation of Information Retrieval techniques for the item recommendation task has become a fertile research area. Previous works have established the correspondence between these two fields that allowed to adapt several retrieval techniques successfully. One line of study aims to model the item recommendation problem as a profile expansion task following the methods for query expansion in pseudo-relevance feedback. To solve the query expansion task in ad-hoc retrieval, several term association measures have been proposed in the past. In this paper, we adapt several of these measures to the top-N recommendation problem, specifically to the collaborative filtering scenario. Moreover, we perform experiments to study their effectiveness regarding accuracy, diversity and novelty. Our results show that some of the proposed measures can improve these aspects over well-known and commonly used recommendation similarity metrics (cosine similarity and Pearson's correlation coefficient).
Diversity and accuracy are frequently considered as two irreconcilable goals in the field of Recommender Systems. In this paper, we study different approaches to recommendation, based on collaborative filtering, which intend to improve both sides of this trade-off. We performed a battery of experiments measuring precision, diversity and novelty on different algorithms. We show that some of these approaches are able to improve the results in all the metrics with respect to classical collaborative filtering algorithms, proving to be both more accurate and more diverse. Moreover, we show how some of these techniques can be tuned easily to favour one side of this trade-off over the other, based on user desires or business objectives, by simply adjusting some of their parameters.
Relevance-Based Language Models are a formal probabilistic approach for explicitly introducing the concept of relevance in the Statistical Language Modelling framework. Recently, they have been determined as a very effective way of computing recommendations. When combining this new recommendation approach with Posterior Probabilistic Clustering for computing neighbourhoods, the item ranking is further improved, radically surpassing rating prediction recommendation techniques. Nevertheless, in the current landscape where the number of recommendation scenarios reaching the big data scale is increasing day after day, high figures of effectiveness are not enough. In this paper, we address one urging and common need of recommendation systems which is algorithm scalability. Particularly, we adapted those highly effective algorithms to the functional MapReduce paradigm, that has been previously proved as an adequate tool for enabling recommenders scalability. We evaluated the performance of our approach under realistic circumstances, showing a good scalability behaviour on the number of nodes in the MapReduce cluster. Additionally, as a result of being able to execute our algorithms distributively, we can show measures in a much bigger collection supporting the results presented on the seminal paper.
Query expansion is a successful approach for improving Information Retrieval effectiveness. This work focuses on pseudo-relevance feedback (PRF) which provides an automatic method for expanding queries without explicit user feedback. These techniques perform an initial retrieval with the original query and select expansion terms from the top retrieved documents. We propose two linear methods for pseudo-relevance feedback, one document-based and another term-based, that models the PRF task as a matrix decomposition problem. These factorizations involve the computation of an inter-document or inter-term similarity matrix which is used for expanding the original query. These decompositions can be computed by solving a least squares regression problem with regularization and a non-negativity constraint. We evaluate our proposals on five collections against state-of-the-art baselines. We found that the term-based formulation provides high figures of MAP, nDCG and robustness index whereas the document-based formulation provides very cheap computation at the cost of a slight decrease in effectiveness.
Álvaro Barreiro合作论文数IRLab, Computer Science Department, University of A Coruna, Spain17