Abstract Scientometric research often relies on large-scale bibliometric databases of academic journal articles. Long-term and longitudinal research can be affected if the composition of a database varies over time, and text processing research can be affected if the percentage of articles with abstracts changes. This article therefore assesses changes in the magnitude of the coverage of a major citation index, Scopus, over 121 years from 1900. The results show sustained exponential growth from 1900, except for dips during both world wars, and with increased growth after 2004. Over the same period, the percentage of articles with 500+ character abstracts increased from 1% to 95%. The number of different journals in Scopus also increased exponentially, but slowing down from 2010, with the number of articles per journal being approximately constant until 1980, then tripling due to megajournals and online-only publishing. The breadth of Scopus, in terms of the number of narrow fields with substantial numbers of articles, simultaneously increased from one field having 1,000 articles in 1945 to 308 fields in 2020. Scopus’s international character also radically changed from 68% of first authors from Germany and the United States in 1900 to just 17% in 2020, with China dominating (25%).
Finding new ways to help researchers and administrators understand academic fields is an important task for information scientists. Given the importance of interdisciplinary research, it is essential to be aware of disciplinary differences in aspects of scholarship, such as the significance of recent changes in a field. This paper identifies potential changes in 25 subject categories through a term comparison of words in article titles, keywords and abstracts in 1 year compared to the previous 4 years. The scholarly influence of new research issues is indirectly assessed with a citation analysis of articles matching each trending term. While topic‐related words dominate the top terms, style, national focus, and language changes are also evident. Thus, as reflected in Scopus, fields evolve along multiple dimensions. Moreover, while articles exploiting new issues are usually more cited in some fields, such as Organic Chemistry, they are usually less cited in others, including History. The possible causes of new issues being less cited include externally driven temporary factors, such as disease outbreaks, and internally driven temporary decisions, such as a deliberate emphasis on a single topic (e.g., through a journal special issue).
Ongoing problems attracting women into many Science, Technology, Engineering and Mathematics (STEM) subjects have many potential explanations. This article investigates whether the possible undercitation of women associates with lower proportions of, or increases in, women in a subject. It uses six million articles published in 1996–2012 across up to 331 fields in six mainly English-speaking countries: Australia, Canada, Ireland, New Zealand, the United Kingdom and the United States. The proportion of female first- and last-authored articles in each year was calculated and 4,968 regressions were run to detect first-author gender advantages in field normalized article citations. The proportion of female first authors in each field correlated highly between countries and the female first-author citation advantages derived from the regressions correlated moderately to strongly between countries, so both are relatively field specific. There was a weak tendency in the United States and New Zealand for female citation advantages to be stronger in fields with fewer women, after excluding small fields, but there was no other association evidence. There was no evidence of female citation advantages or disadvantages to be a cause or effect of changes in the proportions of women in a field for any country. Inappropriate uses of career-level citations are a likelier source of gender inequities.
Women's access to academic careers has been historically limited by discrimination and cultural constraints. Comprehensive information about gender inequality within disciplines is needed to understand the problem and target remedial action. India is the fifth largest research producer but has a low international index of gender inequality and so is an important case. This study assesses gender inequalities in Indian journal article publishing in 2017 for 186 research fields. It also seeks overall gender differences in interests across academia by comparing the terms used in 27,710 articles with an Indian male or female first author. The data show that there are at least 1.5 male first authors per female first author in each of 26 broad fields and 2.8 male first authors per female first author overall. Compared to the USA, India has a much lower share of female first authors but smaller variations in gender differences between broad fields. Dentistry, Economics and Maths are all more female in India, but Veterinary is much less female than in the USA. There is a tendency for males to research thing-oriented topics and for females to research helping people and some life science topics. More initiatives to promote gender equality in science are needed to address the overall imbalance, but care should be taken to avoid creating the larger between-field gender differences found in the USA.
Regression analyses and correlation tests associated with the paper: Female first author citation advantages do not reduce gender disparities in academia
This article assesses whether there are gaps in Wikipedia's coverage of academic information and whether there are non-obvious stylistic differences from academic journal articles that Wikipedia users and editors should be aware of. For this, it analyses terms in the titles of journal articles that are absent from all English Wikipedia page titles for each of 27 Scopus subject categories. The results show that English Wikipedia has lower coverage of issues of interest to non-English nations and there are gaps probably caused by a lack of willing subject specialist editors in some areas. There were also stylistic disciplinary differences in the results, with some fields using synonyms of "analysing" that were ignored in Wikipedia, and others using the present tense in titles to emphasise research outcomes. Since Wikipedia is broadly effective at covering academic research topics from all disciplines, it might be relied upon by non-specialists. Specialists should therefore check for coverage gaps within their areas for useful topics and librarians should caution users that important topics may be missing.
Biochemistry is a highly funded research area that is typified by large research teams and is important for many areas of the life sciences. This article investigates the citation impact and M endeley readership impact of biochemistry research from 2011 in the Web of Science according to the type of collaboration involved. Negative binomial regression models are used that incorporate, for the first time, the inclusion of specific countries within a team. The results show that, holding other factors constant, larger teams robustly associate with higher impact research, but including additional departments has no effect and adding extra institutions tends to reduce the impact of research. Although international collaboration is apparently not advantageous in general, collaboration with the U nited S tates, and perhaps also with some other countries, seems to increase impact. In contrast, collaborations with some other nations seems to decrease impact, although both findings could be due to factors such as differing national proportions of excellent researchers. As a methodological implication, simpler statistical models would find international collaboration to be generally beneficial and so it is important to take into account specific countries when examining collaboration.
Biochemistry is a highly funded research area that is typified by large research teams and is important for many areas of the life sciences. This article investigates the citation impact and Mendeley readership impact of biochemistry research from 2011 in the Web of Science according to the type of collaboration involved. Negative binomial regression models are used that incorporate, for the first time, the inclusion of specific countries within a team. The results show that, holding other factors constant, larger teams robustly associate with higher impact research, but including additional departments has no effect and adding extra institutions tends to reduce the impact of research. Although international collaboration is apparently not advantageous in general, collaboration with the United States, and perhaps also with some other countries, seems to increase impact. In contrast, collaborations with some other nations seems to decrease impact, although both findings could be due to factors such as differing national proportions of excellent researchers. As a methodological implication, simpler statistical models would find international collaboration to be generally beneficial and so it is important to take into account specific countries when examining collaboration.
Scientists and managers using citation-based indicators to help evaluate research cannot evaluate recent articles because of the time needed for citations to accrue. Reading occurs before citing, however, and so it makes sense to count readers rather than citations for recent publications. To assess this, Mendeley readers and citations were obtained for articles from 2004 to late 2014 in five broad categories (agriculture, business, decision science, pharmacy, and the social sciences) and 50 subcategories. In these areas, citation counts tended to increase with every extra year since publication, and readership counts tended to increase faster initially but then stabilize after about 5 years. The correlation between citations and readers was also higher for longer time periods, stabilizing after about 5 years. Although there were substantial differences between broad fields and smaller differences between subfields, the results confirm the value of Mendeley reader counts as early scientific impact indicators.
The importance of collaboration in research is widely accepted, as is the fact that articles with more authors tend to be more cited. Nevertheless, although previous studies have investigated whether the apparent advantage of collaboration varies by country, discipline, and number of co-authors, this study introduces a more fine-grained method to identify differences: the geometric Mean Normalized Citation Score (gMNCS). Based on comparisons between disciplines, years and countries for two million journal articles, the average citation impact of articles increases with the number of authors, even when international collaboration is excluded. This apparent advantage of collaboration varies substantially by discipline and country and changes a little over time. Against the trend, however, in Russia solo articles have more impact. Across the four broad disciplines examined, collaboration had by far the strongest association with impact in the arts and humanities. Although international comparisons are limited by the availability of systematic data for author country affiliations, the new indicator is the most precise yet and can give statistical evidence rather than estimates.
Scientists and managers using citation‐based indicators to help evaluate research cannot evaluate recent articles because of the time needed for citations to accrue. Reading occurs before citing, however, and so it makes sense to count readers rather than citations for recent publications. To assess this, Mendeley readers and citations were obtained for articles from 2004 to late 2014 in five broad categories (agriculture, business, decision science, pharmacy, and the social sciences) and 50 subcategories. In these areas, citation counts tended to increase with every extra year since publication, and readership counts tended to increase faster initially but then stabilize after about 5 years. The correlation between citations and readers was also higher for longer time periods, stabilizing after about 5 years. Although there were substantial differences between broad fields and smaller differences between subfields, the results confirm the value of Mendeley reader counts as early scientific impact indicators.
It is widely believed that collaboration is advantageous in science, for example, with collaboratively written articles tending to attract more citations than solo articles and strong arguments for the value of interdisciplinary collaboration. Nevertheless, it is not known whether the same is true for research that produces books. This article tests whether coauthored scholarly monographs attract more citations than solo monographs using books published before 2011 from 30 categories in the Web of Science. The results show that solo monographs numerically dominate collaborative monographs, but give no evidence of a citation advantage for collaboration on monographs. In contrast, for nearly all these subjects (28 out of 30) there was a citation advantage for collaboratively produced journal articles. As a result, research managers and funders should not incentivise collaborative research in book-based subjects or in research that aims to produce monographs, but should allow the researchers themselves to freely decide whether to collaborate or not. (C) 2013 Elsevier Ltd. All rights reserved.
Individuals and organisations producing information or knowledge for others sometimes need to be able to provide evidence of the value of their work in the same way that scientists may use journal impact factors and citations to indicate the value of their papers. There are many cases, however, when organisations are charged with producing reports but have no real way of measuring their impact, including when they are distributed free, do not attract academic citations and their sales cannot be tracked. Here, the web impact report (WIRe) is proposed as a novel solution for this problem. A WIRe consists of a range of web-derived statistics about the frequency and geographic location of online mentions of an organisation's reports. WIRe data is typically derived from commercial search engines. This article defines the component parts of a WIRe and describes how to collect and analyse the necessary data. The process is illustrated with a comparison of the web impact of the reports of a large UK organisation. Although a formal evaluation was not conducted, the results suggest that WIRes can indicate different levels of web impact between reports and can reveal the type of online impact that the reports have.
Many webometric studies have used hyperlinks to investigate links to or between specific collections of websites to estimate their impact or identify connectivity patterns. Whilst major commercial search engines have previously been used to identify hyperlinks for these purposes, their hyperlink search facilities have now been shut down. In response, a range of alternative sources of link data have been suggested, but all have limitations. This article introduces a new type of link that can be identified from commercial search engines, linked title mentions. These can be found by querying title mentions in a search engine and then removing those not associated with a relevant hyperlink. Results of a proof of concept test on 51 U.S. library and information science schools and four other sets of schools suggest that linked title mentions may tend to give better results than title mentions in some cases when used for site inlinks but may not always be an improvement on URL citations. For links between or co-inlinks to specified pairs of academic websites, linked title mentions do not generally provide an improvement over title mentions, but they do over URL citations in some cases. Linked title mentions may also be useful for sets of non-academic websites when the alternatives give too few or misleading results.
The rise of the social web and its uptake by scholars has led to the creation of altmetrics, which are social web metrics for academic publications. These new metrics can, in theory, be used in an evaluative role, to give early estimates of the impact of publications or to give estimates of non-traditional types of impact. They can also be used as an information seeking aid: to help draw a digital library user’s attention to papers that have attracted social web mentions. If altmetrics are to be trusted then they must be evaluated to see if the claims made about them are reasonable. Drawing upon previous citation analysis debates and web citation analysis research, this article discusses altmetric evaluation strategies, including correlation tests, content analyses, interviews and pragmatic analyses. It recommends that a range of methods are needed for altmetric evaluations, that the methods should focus on identifying the relative strengths of influences on altmetric creation, and that such evaluations should be prioritised in a logical order.
YouTube is one of the world's most popular websites and hosts numerous amateur and professional videos. Comments on these videos might be researched to give insights into audience reactions to important issues or particular videos. Yet, little is known about YouTube discussions in general: how frequent they are, who typically participates, and the role of sentiment. This article fills this gap through an analysis of large samples of text comments on YouTube videos. The results identify patterns and give some benchmarks against which future YouTube research into individual videos can be compared. For instance, the typical YouTube comment was mildly positive, was posted by a 29-year-old male, and contained 58 characters. About 23% of comments in the complete comment sets were replies to previous comments. There was no typical density of discussion on YouTube videos in the sense of the proportion of replies to other comments: videos with both few and many replies were common. The YouTube audience engaged with each other disproportionately when making negative comments, however; positive comments elicited few replies. The biggest trigger of discussion seemed to be religion, whereas the videos attracting the least discussion were predominantly from the Music, Comedy, and How to & Style categories. This suggests different audience uses for YouTube, from passive entertainment to active debating.
Webometric network analyses have been used to map the connectivity of groups of websites to identify clusters, important sites or overall structure. Such analyses have mainly been based upon hyperlink counts, the number of hyperlinks between a pair of websites, although some have used title mentions or URL citations instead. The ability to automatically gather hyperlink counts from Yahoo! ceased in April 2011 and the ability to manually gather such counts was due to cease by early 2012, creating a need for alternatives. This article assesses URL citations and title mentions as possible replacements for hyperlinks in both binary and weighted direct link and co-inlink network diagrams. It also assesses three different types of data for the network connections: hit count estimates, counts of matching URLs, and filtered counts of matching URLs. Results from analyses of U. S. library and information science departments and U. K. universities give evidence that metrics based upon URLs or titles can be appropriate replacements for metrics based upon hyperlinks for both binary and weighted networks, although filtered counts of matching URLs are necessary to give the best results for co-title mention and co-URL citation network diagrams.
In May 2011 the Bing Search API 2.0 had become the only major international web search engine data source available for automatic offline processing for webometric research. This article describes its key features, contrasting them with previous web search data sources, and discussing implications for webometric research. Overall, it seems that large-scale quantitative web research is possible with the Bing Search API 2.0, including query splitting, but that legal issues require the redesign of webometric software to ensure that all results obtained from Bing are displayed directly to the user.